direct-response-agent
Simulation window · 64 inferences · declaration v14
Rendered 2026-08-11 18:54 UTC
Measured over 64 inference(s).
Conformance Built as declared
Whether the agent was built the way its declaration says, which is a different question from whether it behaved well, and fails to a different person.
Not checked
Said out loud rather than omitted. No findings and no answer look identical in the output, and only one of them is a clean bill of health.
declaration_contentNo published content hash to compare against, so nothing here confirms the inferences ran the declaration ai-gov published. Expected on a local report.
Guardrails
What each gate checked and how often it fired. A gate that never fires is not evidence of a clean population, only that nothing got past it.
| Instruction | Checked | Violations | Rate |
|---|---|---|---|
conform_to_output_schema | 64 | 2 | 3% |
Awaiting
Deduped by action, not by assessment. One labelling job unblocks every KPI that reads the evaluator.
Label some outputs — 18 of 40
staff_reply_decisionThis measurement is produced by a person, not by the platform. Until labels arrive there is nothing for the agent to be right or wrong about.
Owed by a reviewer.
Unblocks 2 KPI(s) reply_accepted_sharereply_decision_coverage
Health checks
A target says what good looks like, so these settle something.
Grounded In Thread
MissedPassed 55 of 64 measurement(s)
passrate
The evaluator behind this
Reply Accepted Share
AwaitingOver 18 measurement(s)
share of accepted
This measurement is produced by a person, not by the platform. Until labels arrive there is nothing for the agent to be right or wrong about.
Owed by a reviewer.
The evaluator behind this
Reply Decision Coverage
AwaitingCollected on 18 of 40 eligible inferences
coverage
This measurement is produced by a person, not by the platform. Until labels arrive there is nothing for the agent to be right or wrong about.
Owed by a reviewer.
The evaluator behind this
Guardrail Fire Rate
MetPassed 62 of 64 measurement(s)
passrate
The evaluator behind this
Length Requirement
MetPassed 62 of 64 measurement(s)
passrate
The evaluator behind this
Metrics
Measured, and nothing compares them to a target. They describe rather than settle.
Reply Similarity
No targetOver 64 measurement(s)
mean
The number is measured and nobody has decided yet what good looks like. Reported so it can be watched; it settles nothing until a target is declared.
The evaluator behind this
Nothing owed on 1
Inert on this report: nothing ages, nobody is paged, and no amount of waiting changes them.
| KPI | State | Why | Value |
|---|---|---|---|
label_stability | Raise the replicate count | This measurement needs the same input answered more than once and got one run of each. More runs of that shape do not help; the run needs more replicates per input. | — |