A generation prompt carrying all the rules does not, by itself, produce reader-facing copy I would sign. I learned this by running prose through a live pipeline, and the answer was not a sharper prompt. It was a review pass that always runs. This is the ground behind Content at scale.
TL;DR
A first pass tops out around ~75-80% rule-compliant even with the rules and the negative examples written into the generation prompt, and a detect-and-fix layer only catches what someone already imagined. A mandatory review pass runs on every reader-facing stage.
How AI copy is held to quality today
Two practices carry the load, and both look reasonable on paper:
- A generation prompt with explicit rules and negative examples. The rules sit in the prompt, and the negative examples teach the model what to avoid.
- A detect-and-fix layer. The output is matched against a blocklist of known violations, and the model is asked again only when a match fires.
That is how I held quality first, and it is where most pipelines stop.
Where it breaks
The generation prompt does not get there on its own. Rule compliance on a first pass alone tops out around ~75-80%, even with the rules and the negative examples in the prompt, so poor text reaches readers. The detect-and-fix layer has its own ceiling: it catches anticipated violations, and anything outside the blocklist passes untouched.
Why the prompt cannot police itself
The failure is structural, not a tuning problem. The generation prompt is asked to both produce the output and police it, and a blocklist can only fire on a violation someone already imagined. Quality ends up enforced by exactly what the generator anticipated, which is the thing it cannot do reliably.
I watched this play out on a synthesis prompt I rewrote once. A fixed sentence-count iteration over-corrected into generic boilerplate and, in one test run, introduced a claim no source digest supported. The length discipline that held was the one enforced in code, outside the prompt.
What it costs at scale
The failure is per stage, so it compounds. Every reader-facing stage with no review pass leaks its own share of poor text, and un-caught output from an early stage feeds the next one. In the digest pipeline I run, the article-digest review runs before cluster synthesis reads the digests, then the synthesis revision and the title pass follow. A stage that skips its pass passes its own defects down the line.
The rule: a review pass on every reader-facing stage
The decision is a standing rule, not a preference: every reader-facing prose-generating stage must have a mandatory post-generation quality-review pass, always run and never conditional on a heuristic. Where there was a choice, the rule prefers always-run rules-based review over detect-then-fix, because it reviews everything rather than only what a blocklist anticipated. The rule carries one exemption: a prose stage whose only output is an internal debug string, never shown to the reader.
The reliable pattern is to ask a model, with clear rules, to review everything and correct what is wrong. In the digest pipeline that means the article-digest review, the synthesis revision, and the title pass. These are model-based, rules-driven passes run automatically; the human gate is a distinct actor, the person who signs off before anything ships . The goal is never unattended generation: the review pass, not the generator, is what makes the output publishable.
The objection: a better prompt should remove the need
The strongest counter is fair: reviewing everything is slower and costs more, and a better generation prompt should remove the need for it. It does not, on the record I have. First-pass compliance tops out around ~75-80% even with explicit rules and negative examples in the generation prompt, so the review pass, not the generation prompt by itself, is what makes the end result reliable. I budget it as load-bearing, not as a nicety.
The boundary stays visible. The rule buys reliability at a stated running cost, and it exempts only stages whose output no reader ever sees. Neighbouring quality mechanisms are their own patterns, not variants of this one: deterministic guards on grounded output , and two-stage rendering, where the model decides and the code renders .
The operator behind this method is on the about page.
If this is the shape of your problem, let’s talk.
