A free-text advice line reads the same whether the data supports it or not. That is the trap. A model writes with equal fluency from real forecast data and from nothing, so a false sentence reaches the user and nothing in its shape flags it. I run an automation that turns a weather forecast into a one-line recommendation, and I rejected the instinct to ship that line as written. The model proposes the line; code decides whether it may reach the user. That split is the core of Reliable agent systems.
The standard answer to hallucination is grounding: connecting the model to current data so its answers track verified sources. A vendor analysis of enterprise hallucination reports that this reduces hallucinations and supports decision-making, and that linking an answer back to its sources lets the user verify it. Grounding shapes what the model draws on. It does not check the one sentence that leaves. That gap is where a deterministic guard sits.
TL;DR
The guard is deterministic output validation. It drops the ungrounded line before it reaches the user, and it does not correct it. Two checks run in code, with no model in the loop: a specific-time arrival claim is dropped when the forecast shows no precipitation, and any precipitation mention is dropped when no hour carries real probability or intensity. What reaches the reader is grounded in the data.
The problem, and the answer it rules out
The advice line is free text, and nothing constrains it to the fields the forecast actually carries. It can state a rain-arrival time, or mention precipitation, when the data holds no such signal. The failure is silent because the sentence is fluent: a claim the data does not carry reads exactly like one it does.
The answer I ruled out was passing the advice through unfiltered. It is the default, and it is the trap, because it lets the hallucination reach the user with nothing standing between the model and the person reading it. The guard sits downstream instead. The model produces the line, and deterministic code decides whether it may be shown.
A guard on top of the JSON contract
Two decisions stack. The first is the substrate the guard runs inside: the two-stage LLM pattern, which I hold as non-negotiable. The model returns JSON only, and the code renders everything else. A call that returns prose, HTML, or Markdown instead of JSON is a regression, not a feature, because it puts presentation back into the model and out of the deterministic layer. That rendering contract is its own subject ; here it matters as the ground the guard stands on.
The second decision is the guard itself. Before the advice line is appended, validate_advice() runs. It holds two deterministic checks, both needing no model call.
Two alternatives were live, and I rejected each on a stated reason. Passing the advice through unfiltered fails on hallucination risk: it is the trap above, and it carries no check at all. Letting the model produce the final rendering fails as a regression: it undoes the two-stage pattern the guard is built on. Which tool carries the advice step at all is a routing decision of its own , not one this guard settles.
The two checks
The specific-time arrival check drops a line that says rain arrives at HH:00 when precipitation is 0.0 mm across the entire window. High probability alone does not justify an arrival claim, so the check reads the millimeter column, not the probability. The no-signal check drops any precipitation mention when no hour carries a probability above 15% or precipitation above 0.1 mm.
Both checks run on the advice string before it is appended, and both are model-free. A phrasing they leave alone is one that makes no arrival claim and no precipitation statement the data does not carry, such as the risk rises after HH:00 form. The line is dropped, never rewritten. When a check fires, the advice is removed; it is not corrected.
The guard is itself a code-level check, which raises a fair question: who checks the code that decides? Keeping what builds separate from what inspects is its own discipline, and this guard lives inside it rather than replacing it.
The line the data contradicted
The guard exists because of one production failure. The model wrote that rain arrives at HH:00 when precipitation was 0.0 mm throughout the window: a confident, fluent line, and false against the forecast it was meant to summarize. Both checks now drop that line silently, with a stderr warning and nothing else.
I found it because the unchecked advice could be read against the forecast it summarized, and the two did not match. The drop is silent by design, and it is a content check, not an operations check. Whether a pipeline fired, and whether it fired twice, is a different discipline ; this check removes an ungrounded claim from what a pipeline produced. The same weather automation, told in full as a diary, has its own place .
What it costs to run
The guard adds no service and no model call. It runs in the same pass that appends the advice, on data already in hand, so the added cost is the two string checks and nothing operational.
The substrate carries its own maintenance rule, and it is what keeps the pattern from eroding. Every new LLM prompt in the two-stage pattern must include an instruction to return JSON only, a JSON schema or an example structure, a null-safe fallback on error, and a note in the calling code marking the parse step. Neither the guard nor the pattern is metered, and I state no figure I did not measure.
What the guard does not solve
Three boundaries, drawn rather than glossed, because a guard that is oversold is a guard that will disappoint.
First, the guard drops a line; it does not correct or replace it. The reader gets less advice, not better advice, and a line that is arguably true but phrased as an arrival claim is dropped alongside the false one.
Second, the checks test the line against the data on two axes only: an arrival claim against precipitation in millimeters, and any precipitation mention against probability and intensity. A false claim outside those two axes passes untouched.
Third, the two-stage contract makes the output’s shape deterministic, not its content. Valid JSON can still be wrong, which is why a code-level check like this is a second layer and not a proof of correctness. The guard bounds what the model says, not the machine it runs on , and it is a floor on the advice line, not a step toward running unattended; it does not replace the judgement of the person who reads the result. A check that runs before a line reaches a person is the same impulse in a different place.
The operator behind this pattern, and the trade-offs I accepted to run it, are on the about page. If this is the shape of your problem, let’s talk.
