Skip to content

PlaybookAgentic OSPart 5 of 7

Most common failure modes for agents during QA

I’m curious to hear what founders are running into while running agents through QA. For context, we don’t have specific QA agents running. Instead, we’ve set up infrastructure so every coding agent can QA its own work. Are there frequent pitfalls you run into while setting this up? Does your org do anything similar?

Here’s a summary of what we’ve learned while handling this for ourselves and our customers:

Agents need to go beyond a unit suite#

Agents hand in code that passes unit tests and fails in production much more often than humans do. Validating output means running the change against a live copy of the system it changes.

  • Failure mode we fixed: our CI “integration” tier ran on in-memory SQLite. It was fast and needed no external services, but Postgres-only behavior (partial indexes, JSONB operators, trigram search) was only tested once it hit a preview or staging. Plenty of our tests looked green while skipping that behavior entirely.

A comprehensive setup with three layers#

  1. An isolated running environment per change. Ours has two tiers. By default, every PR gets a frontend preview wired to shared staging. Some PRs get a full stack of their own: database branch, API, workers, and migrations run for that commit.
  2. Realistic data inside it. Most teams skip this step. An empty environment produces false positives.
  3. A written playbook for how to drive it, one per variant. Driving a browser and calling an API are different jobs.

The preflight#

Every playbook should start by asking whether its current environment can express the changes in the PR. That check belongs at the beginning, so you don’t find out it’s blocked after a 30-minute browser run. Ours (from the example above) checks whether each diff depends on backend behavior and whether this PR’s preview has its own backend. If not, it stops, posts the blocker, and recommends a new build with a full stack. It fails after one metadata call, not after a whole browser session.

The data ladder#

You should give each session the data it needs to validate its changes. That said, go with the cheapest and safest option first:

  1. Fixtures in unit tests
  2. A local seed
  3. A stamped dataset in the preview environment (ours refreshes daily, and every preview forks it copy-on-write)
  4. Read-only production analytics

Access matters as much as existence#

Double-check that authentication runs with credentials the agent can get without a human, a VPN or an interactive login.

  • Failure mode we fixed: our backend API-smoke playbook used to require the staging VPN and a cloud profile, so the cloud agents it was written for couldn’t run it.