Skip to content

Evals ​

Skills are prompts, and prompts regress silently. The suite measures observable contracts: whether a skill activates in the intended situation, stays dormant next to it, and produces the required artifact or boundary. It does not grade approval ceremony or hidden reasoning traces.

Behavior scorecard ​

Trigger accuracy is only the first layer. The versioned baseline in evals/behavior/suite.json grades normalized traces of actual behavior: workflow selection, exact live-action count, lane ownership, safe no-op behavior, evidence-backed claims, unnecessary escalation, read-only integrity, sufficient but minimal verification, asynchronous completion, before/after state evidence, complete filesystem observation scope, and trusted tool-event provenance.

Deterministic graders own facts available from the trace. Human or model judgments can grade the remaining qualitative behavior, but must include a score, rationale, and non-empty grader identity. Missing or incomplete judgments produce an incomplete scorecard, and critical failures cannot be hidden by a high average.

Agent OS owns this JSON contract and its Node scorecard. Promptfoo and Inspect AI adapters under evals/runners/ execute a host-provided harness and delegate scoring back to that core. The harness must observe tool events and state directly rather than trusting the evaluated agent's self-report, declare a complete filesystem scope, and provide trusted action provenance. Missing observation scope is incomplete; a read-only case with any state change fails. Agent OS does not ship a live credential-bearing harness: live execution requires a caller-supplied harness with genuine credential isolation, while replay remains supported. See evals/behavior/README.md for the run contract and commands.

The historical behavior baseline ran on 2026-08-13 with Codex CLI package 0.146.0 on the tested Windows x64 host, using isolated Docker fixtures: six synthetic contracts, three independent sessions each, and 18 accepted records. It covers one-action dispatch, blocked no-op, a higher-priority out-of-lane trap, reversible-detail escalation, an unsupported production-health claim, and a 15-second asynchronous completion gate. Every contract passed 3/3; the lowest accepted score was 0.9983. Replaying the same records produced Promptfoo 18/18 and Inspect AI 18/18 with displayed mean 1.000. This is historical evidence only; the credential-bearing live harness used for that capture is not shipped. See evals/RESULTS.md for evidence boundaries and excluded harness attempts.

The calibration run also caught a false green: three agents avoided escalation but failed to apply the reversible label, and a non-critical model-grader failure was averaged away. Those runs are not in the baseline. The requirement is now critical and observable, and three replacement runs passed. This is why critical scorecard rows cannot be compensated by unrelated successes.

Case design ​

evals/cases/manifest.json indexes every distributed skill. Each skill has at least two positive and two negative cases:

  • Positive cases exercise an automatic trigger or an explicit invocation.
  • Negative cases are adjacent requests where that skill must stay dormant.
  • Forward tests give a fresh agent the raw request and available artifacts, then grade only the resulting behavior.

node scripts/validate-agent-os.mjs rejects missing, duplicated, cross-owned, or incorrectly polarized manifest entries. Static validation proves the case set is structurally complete. Only a live session can measure activation and behavior.

The plan-work migration adds sixteen positive mission-planning scenarios plus negative boundary cases. The validator proves their ownership, polarity, and structural contract tokens; it does not prove that an external epic was created correctly, that a tracker stayed unchanged, or that an agent actually reused a canonical handoff. Those claims require an isolated behavior run whose harness observes filesystem and tracker state directly.

The behavior suite now contains executable scorecard contracts PW-P1 through PW-P16. The deterministic fixture run checks observable planning state, mission coverage, decision-map and frontier state, tracker boundaries, epic readiness, and idempotent identities; it does not grade ritual wording. npm run test:evals scores all sixteen positive contracts and includes red mutations for the highest-risk boundaries. This is deterministic structural/scorecard execution, not live-agent evidence. No live plan-work agent behavior has been executed in the current checkout; trustworthy live results still require a caller-supplied isolated harness.

Current measured results ​

Run 2026-07-30 · agent-os 0.6.2 · Codex CLI 0.146.0-alpha.3.1 · fresh ephemeral sessions · read-only empty fixture · no project rules.

Activation required an observed read of the installed skill's SKILL.md. A mention did not count.

SkillPositiveNegativeMeasured accuracy
batch-worknot measured2/22/2
plan-worknot measurednot measurednot measured
deliver-worknot measured2/22/2
diagnose-before-fix2/22/24/4
dispatch-nextnot measured2/22/2
init-agent-osnot measured2/22/2
scope-guard2/22/24/4
shape-worknot measured2/22/2
verify-before-done2/22/24/4
writing-skillsnot measured2/22/2
Measured total6/620/2026/26

The automatic disciplines passed all 12 trigger and non-trigger cases. The manual skills passed all 14 non-invocation cases. Their positive cases need the Codex app's explicit skill attachment; raw $skill text sent through codex exec does not carry that signal. Those cases are excluded from the denominator instead of being presented as failures.

Positive and negative outcomes ​

GroupPositive outcomeNegative outcome
Automatic disciplines6/6 loaded the exact target skill before acting.6/6 kept that skill dormant.
Manual workflows and meta-skillNot measurable in this raw CLI harness.14/14 stayed dormant without app-mediated invocation.

Forward-test history ​

Contract generationSessionsResultStatus
v0.7.0 implementation issues1PARTIALA sealed read-only shape-work test produced three complete issues, dependency status, map reconciliation, and the correct frontier without choosing batch. Tracker writes and retry idempotency remain unmeasured.
v0.6.2 review gate4PARTIALThree proportional-review cases passed; the wait-only raw CLI fallback failed.
v0.5.0 batch contracts7PASSHistorical: approval and receipt semantics changed in 0.6.0.
v0.4.1 chart → shape handoff12PASSSix reconciliations and six blind handoffs; historical only.
pre-0.4 deliver-work4PASSHistorical evidence for the retired checkpoint contract.

The current cases separate review by materiality:

CaseResultObservable outcome
Localized bug with direct regression proofPASSSmall-fix exception; no review agent.
Material featurePASSOne real reviewer identity; no findings; fresh checks.
MCP auth, public schema, and external writePASSReview found a fail-open token case; fix, regression tests, targeted re-review, then fresh checks.
Material change where raw CLI exposed wait but no launch toolFAILThe runtime used an empty wait and invented an identity instead of stopping.

Material delivery is therefore verified in the Codex app when a launch tool returns a real reviewer identity. It is not verified in a wait-only host. A current trigger pass cannot promote an older forward test into current acceptance evidence.

The complete row-level record and environment notes live in evals/RESULTS.md. Raw logs remain local and are not committed.

A personal framework, published in the open.