Design an AI team that doesn’t trust its own output: separating spec, build and QA

Photo by Daniil Komov on Unsplash

The agent that writes the spec tends to be the one that vets it, so nothing gets a second set of eyes. I built two rosters against that: the virtual org that specs and builds my stack, and the agent team that mines evidence into this site. In both, no actor signs off on its own work: the spec author never reviews its own proposals, and the reviewer that does is locked out of the files it checks. That separation makes the output dependable enough to run, and it sits inside Reliable agent systems.

Note: By convention in this article, the Japanese first names mentioned are pseudonyms I gave to each of my AI agents to associate each role and easily refer to it. It’s used by me and my AI agent team as if they were real individuals and helps with the communication and greatly simplifies the work.

TL;DR

The second set of eyes is a separate write boundary, not a separate job title. A reviewer that can edit its own subject is not reviewing it, and only the right to write, removed, makes the split hold.

The design problem, and the answer it rules out

In that org, the spec author’s proposals are the system’s specs. Sovereignty and quality set the constraint: if that author also vetted its own output, those specs would have no second set of eyes.

The obvious answers fail the same way. One is to add a quality assurance (QA) role and call the second opinion solved. The other is to let the spec writer spawn reviewers over its own work. A critic that can read and edit those files is not an independent check; one sharing the same posture and working context returns the conclusion it was asked to audit. I rejected that self-review, even though the spec author already shared its siblings’ filesystem.

Split the work by write root and posture

The design decision was to partition the team by what a worker may touch and how it reasons, never by subject matter: write root and posture are the two axes.

The constraint that made this the right axis is a plain platform fact: a task card already accepts a per-card skill pin, its own model and provider overrides, and its own goal settings, so specialization arrives with the work. One profile can be a copywriter on one card and an auditor on the next.

I weighed two alternatives and rejected both on stated reasons. A function-split roster failed on redundancy: analyst, journalist, and documentalist repeat one read the corpus and write prose posture three times. A minimal three-role roster failed on boundaries: it puts the union of the filesystem, the web, and the chat toolsets on the highest-volume profile.

Six workers resulted, each holding a distinct write root. How each task gets routed by its shape is a separate decision .

How the separation runs

The stack runs on a fixed hand-off chain. I identify the task; Ren writes the spec and never authors the vault; I carry the brief to Takumi, who builds as the sole developer, in Claude Code; git push, then the virtual private server (VPS) cron and Izuna run. Ren does not spawn subagents into the sibling repositories despite sharing their filesystem, and Takumi’s independent senior-developer review of Ren’s proposals is the only real QA step in the chain.

That is the design: QA is an emergent function, not a role. It is what happens when the roles above work correctly together, and it has no seat of its own in the org chart.

The separation is enforced structurally, not by policy. Each profile’s HERMES_WRITE_SAFE_ROOT, set in its own environment file, is the write-root column of the roster. The auditor can read the whole vault, but a write to the stories or raw folders is hard-denied. A critic that can edit is a critic that can quietly fix a defect and report success. The same write root, read as an operating-system boundary against a hostile actor, is a different register .

The defect flag no one read

In production, a worker flagged a defect in its own graph’s output and nothing consumed the flag. Kura wrote it into the index prose as a sentence, and the graph finalized anyway. No card was opened and no profile was told.

Kura had correctly flagged that 4 of 8 new records duplicated pre-existing ones, and the root finalized regardless.

What fixed it is a small, fixed grammar called a FOLLOW-UP line. When a defect is in the current graph’s own output, its finder adds one line per defect to its card’s closing comment: FOLLOW-UP: <what is wrong> | owner: <profile> | records: <id, id>. Before finalizing a root task, the orchestrator turns each line into one correction card, assigned to the named owner and parent-linked to the root, and waits for those cards to close. If the same FOLLOW-UP recurs after its correction card ran, the root blocks and the orchestrator escalates to me.

Routing a defect a worker already flagged is not the same as catching one. Whether a model output was ungrounded, and should be stopped before it is reported, is its own guard.

What it costs to run

The escalation adds one orchestrator pass over every child card’s closing comment, once per root task. That is the whole added cost.

It needed no new runtime capability. The mechanism composes three things the platform already supported: the root’s assignee wakes when all its children complete, a child parent-linked to the root makes the root wait for it, and a worker opening a card and parent-linking it already worked.

The structural boundary carries its own maintenance rule. Keeping the auditor’s write guard intact means the miner cannot hold the terminal, since a git history read through it would void the guard; so the git log is pre-dumped into the raw sources folder, and the miner reads that dump as a file.

The fail-loud and review discipline on daily automation is a different register .

What the split does not cover

Two honest boundaries.

The write boundary does not cover every profile. Kura is the one profile whose restraint is policy, not structure, because index and cross-link maintenance touches every folder. Its write root is the entire vault, and its git commits are the thing nothing else audits: a commit that changes prose rather than structure or metadata is a defect nothing else will catch.

The second boundary is the cost of the first. QA as an emergent function means there is no separate reviewer to point at when the roles fail together, and the senior-developer review is only as independent as the brief that feeds it. I identify the task and carry the brief between the spec author and the builder, so a human sits in the chain by design, not as a fallback. That is where the trust comes from.

The always-run review pass on reader-facing copy is the same principle in a different place, and ranking AI dependencies by whether they survive a vendor cutting service is a different failure axis.

Two rosters, one rule: a second set of eyes has to be able to say no to the work, not just inspect it. The chain that runs my stack keeps a human at the hand-offs and an independent developer at the build, and both are load-bearing. The trade-offs behind the pattern are on the about page. If this is the shape of your problem, let’s talk.