direct-response-agent

Simulation window · 64 inferences · declaration v14

declaration
a1b2c3d4-e5f6-4789-a012-3456789abcde
content hash
3cff3e86a506
window
2026-08-04 to 2026-08-11

Rendered 2026-08-11 18:54 UTC

Essentials passed 210 measurement(s) · tiers essential, always, analytics
Met 2Missed 1Awaiting 3No target 1

Measured over 64 inference(s).

Conformance Built as declared

Whether the agent was built the way its declaration says, which is a different question from whether it behaved well, and fails to a different person.

Not checked

Said out loud rather than omitted. No findings and no answer look identical in the output, and only one of them is a clean bill of health.

  • declaration_content No published content hash to compare against, so nothing here confirms the inferences ran the declaration ai-gov published. Expected on a local report.

Guardrails

What each gate checked and how often it fired. A gate that never fires is not evidence of a clean population, only that nothing got past it.

InstructionCheckedViolationsRate
conform_to_output_schema6423%

Awaiting

Deduped by action, not by assessment. One labelling job unblocks every KPI that reads the evaluator.

Health checks

A target says what good looks like, so these settle something.

Grounded In Thread

Missed
0.86target ≥ 0.95

Passed 55 of 64 measurement(s)

passrate

The evaluator behind this
key
groundedness
evaluator agent
evaluator-groundedness
scale
binary
producer
tier
always
results
64
pass rate
86%

Reply Accepted Share

Awaiting
≥ 0.40, not compared while awaiting

Over 18 measurement(s)

share of accepted

Label some outputs — 18 of 40

This measurement is produced by a person, not by the platform. Until labels arrive there is nothing for the agent to be right or wrong about.

Owed by a reviewer.

The evaluator behind this
key
staff_reply_decision
evaluator agent
scale
categorical
producer
external
tier
analytics
results
18

Reply Decision Coverage

Awaiting
0.45≥ 0.90, not compared while awaiting

Collected on 18 of 40 eligible inferences

coverage

Label some outputs — 18 of 40

This measurement is produced by a person, not by the platform. Until labels arrive there is nothing for the agent to be right or wrong about.

Owed by a reviewer.

The evaluator behind this
key
staff_reply_decision
evaluator agent
scale
categorical
producer
external
tier
analytics
results
18

Guardrail Fire Rate

Met
0.97target ≥ 0.95

Passed 62 of 64 measurement(s)

passrate

The evaluator behind this
key
output-conformance
evaluator agent
evaluator-output-conformance
scale
binary
producer
tier
essential
results
64
pass rate
97%

Length Requirement

Met
0.97target ≥ 0.90

Passed 62 of 64 measurement(s)

passrate

The evaluator behind this
key
response-length
evaluator agent
evaluator-length
scale
binary
producer
tier
always
results
64
pass rate
97%

Metrics

Measured, and nothing compares them to a target. They describe rather than settle.

Reply Similarity

No target
0.71no target

Over 64 measurement(s)

mean

No target set

The number is measured and nobody has decided yet what good looks like. Reported so it can be watched; it settles nothing until a target is declared.

The evaluator behind this
key
similarity-to-approved
evaluator agent
evaluator-similarity
scale
numeric
producer
tier
analytics
results
64
mean
0.71
p50
0.74
p90
0.88
Nothing owed on 1

Inert on this report: nothing ages, nobody is paged, and no amount of waiting changes them.

KPIStateWhyValue
label_stabilityRaise the replicate countThis measurement needs the same input answered more than once and got one run of each. More runs of that shape do not help; the run needs more replicates per input.