Evals: 0.9.0
The full run on skill 0.9.0, with the two targeted checks that followed it:
what the Evals page said from 2026-08-25 until the 0.11.0
series replaced it on 2026-09-03. Text and tables are as published
(docs/evals.md at
840a1d5),
not re-worded — "above" and "below" mean the page of that day, and the case
count, the rules and the caveats are the ones that applied then. How the suite
is run and judged today, and every other series: Evals, run
history.
Results
56 of 70 passed — run 2026-08-25, skill 0.9.0, Claude Code CLI 2.1.241,
agent and judge both Claude Sonnet 5 (claude-sonnet-5). No errors; every
case produced a graded verdict. Previous full run: 59/67, skill 0.6.2,
2026-07-31.
One clear improvement first: negative-conflicting-sources — the case that
had failed in both prior runs (code said 3 retries, the architecture doc said
5; the agent needs to record both and flag the conflict rather than declaring
one authoritative) — now passes cleanly.
The 14 failures cluster into two categories, and the split between them shifted a lot since the last run:
1. Activation gaps (11 of 14 failures). In every one of these, the Skill
tool itself was never invoked — no [tool call] Skill in the transcript at
all. This is the same activation limitation already documented in
"Composition with other
skills"
and tracked in issue #138
— a Skill activates when the conversation matches its description, and
nothing guarantees that for a low-signal prompt; none of these 70 runs had a
SessionStart hook or any other activation aid configured, so this is the
plain, unmitigated version of that gap, not evidence about hooks one way or
the other. Some of the 11 did nothing skill-related at all
(init-declined-not-reasked,
update-check-cannot-run-surfaced-once, update-check-repeat-failure-no-reask,
source-reference-never-does-not-ask,
recheck-after-other-skill-concludes-mid-conversation) — the model just
answered the prompt directly or asked an unrelated generic question. Others
read AGENTS.md/context/ directly and approximated the right behavior
without ever running the formal workflow (context-schema-missing-backfilled,
significant-correction-is-not-a-decision,
capture-mode-proactive-with-confirm-always,
confirm-always-clear-case-still-asks-permission,
agents-local-gitignore-not-covered,
ambiguous-worth-capturing-asks-instead-of-guessing) — closer, but each still
missed at least one documented step (a .gitignore entry, a deterministic
backfill rule, a confirm-always permission gate, a CHANGELOG.md entry).
2. Recognizes but doesn't act (3 cases). The skill loaded, correctly identified what needed to be written — then asked or deferred instead of writing it. This is the pattern already flagged as still-open in issue #131 ("a first fix helped but didn't eliminate the pattern; a follow-up attempt made one case worse instead of better") — this run reconfirms it, not with one case but with all three skill-active failures:
capture-confirmation-automatic-unclear-evidence: config setscapture-confirmation: automatic; the agent asked whether the constant was picked for a specific reason instead of writing the entry with an honestEvidence: unknown/inferred, as the rule requires.embedded-procedure-not-why-content: correctly found both right destinations (context/for the limitation,CONTRIBUTING.mdfor the workaround procedure) but wrote neither, ending the turn on an unrelated question about rejected alternatives instead.open-question-gets-status-open-not-unknown: correctly found the undocumented branch and correctly avoided inventing a rationale for it, but asked the user whether to write it up instead of recordingStatus: open/Evidence: unknownitself, as the case calls for.
Activation-gap follow-up: a real SessionStart hook, tested
Not a new full-suite run — a targeted re-test of the 11 activation-gap
failures above, done the same day (2026-08-25). One of those 11
(init-declined-not-reasked) starts from an empty project with no
AGENTS.md at all, so there's no keep-the-why:config marker for a hook to
find — out of scope for this by construction. The other 10 were re-run
against the same fixtures plus one change: a project-scoped SessionStart
hook, checked into tools/evals/fixtures/_base/.claude/settings.json (not
the machine-level one a developer might have personally) — it greps
AGENTS.md for the keep-the-why:config marker and, if found, injects
"Load the keep-the-why skill now, before other work" as additionalContext
before the first turn.
Result: 10/10 now invoke the Skill tool (was 0/10). 9/10 pass outright.
The one holdout, update-check-cannot-run-surfaced-once, turned out to be
an unrelated fixture bug, not a skill or hook problem: its "simulate no web
access" setup denies WebFetch/WebSearch but not Bash, so the agent
reached the real GitHub API with curl and the simulated failure never
actually happened — the same bug affected update-check-repeat-failure-no-reask,
which had happened to still pass by coincidence (a real successful check is
also valid output for that case's expected behavior). Fixed both fixtures'
case.json to also deny Bash(curl *)/Bash(wget *); re-run, both pass
cleanly for the right reason this time (curl denied, agent reports the
failure and asks/retries-quietly as expected). Final: 10/10 pass.
This hook is now the _base fixture's default, so every future full run
includes it — the 56/70 above remains the last true baseline without one.
The next full --all run (not done today) is what turns this from "10 of 11
formerly-failing cases, re-run in isolation" into a real before/after on the
whole suite. Writeup and the reusable hook snippet: docs/autostart.md
(published with 0.10.0).
Config-relocation regression check (2026-08-31)
Not the next official full-suite run — a targeted check that moving project/personal config into dedicated .keep-the-why/~/.keep-the-why/<id>.md files didn't regress anything, run against the suite after that change (72 cases now, was 70 — see references/migrations.md's entry for what changed). --all --parallel 4, skill unreleased (post-0.9.2), Claude Code CLI 2.1.251, agent and judge both Claude Sonnet 5.
First pass: 64/72 passed, 8 failed, 0 errors. Re-running each failure individually (per the caveat below about single-run variance) resolved all but two to a pass:
- Three flipped to pass on re-run with no code or fixture change (
capture-confirmation-automatic-unclear-evidence,capture-confirmation-automatic-still-asks-substantive-question,confirmation-flow-batch-multiple-candidates) — ordinary model variance on judgment-quality questions (how good a clarifying question is, how cleanly a batch of candidates gets presented) that have nothing to do with where config lives.confirmation-flow-batch-multiple-candidatesis worth naming specifically since it's one of the migrated fixtures: its transcript shows the agent correctly locating and reading~/.keep-the-why/<id>.mdfrom the fake$HOMEand correctly extractingconfirmation-flow: batchfrom it before the presentation-quality miss — the new personal-config plumbing itself worked. - Two (
embedded-procedure-not-why-content,open-question-gets-status-open-not-unknown) are the exact "recognizes but doesn't act" cases already documented above from the 2026-08-25 run — unrelated to this change, still open per issue #131. negative-manufactured-abandoned-reasoningfailed twice in a row (a rule-1 violation: treated "found no reference anywhere" as license to delete rather than as an unresolved Chesterton's Fence candidate) — a genuine, pre-existing content-judgment gap, not something this PR touches (no core rule text changed) or something worth chasing further here.config-migrates-to-dedicated-file(new case) failed twice, both times a pure activation gap — the Skill tool was never invoked at all despite theSessionStarthook firing correctly (confirmed present in the transcript this time, unlike the corrected case below). A manual run with an explicit "check/initialize keep-the-why" prompt confirmed the actual migration instructions are followed correctly once the skill engages:idgeneration (falling back touuid+folder-namewhen there's no git remote —uuidgenwasn't installed on this host either, so it improvisedpython3 -c "import uuid; ...", still a one-off command, not a shipped script), field carry-over, and theAGENTS.mdpointer/version-note all came out exactly as specified.init-declined-not-reasked— three runs, three different outcomes, none of them a config-location bug:- Fail, but not the skill's fault. The agent bypassed the skill's own state entirely and used Claude Code's own out-of-repo auto-memory feature (
~/.claude/projects/.../memory/feedback_*.md) to record "don't suggest this again" instead of writinginit: declinedanywhere in the project. This is a genuinely new, real finding — one this fake-$HOMEisolation fix made visible by giving the agent working credentials and therefore a working auto-memory feature to reach for, though the underlying possibility isn't new (every eval run before this PR also had a real, working$HOMEwith the same feature available, unisolated). Filed as #198 rather than addressed here — it's a skill-robustness question (should the trust model say something about not treating another tool's own persistence feature as equivalent to this skill's documented state?), not a config-relocation bug. - Fail, and this one's the judge's fault, not the agent's. The transcript shows the agent correctly checking for
.keep-the-why, finding none, and correctly creating it withinit: declined— exactly the new spec. The judge'sreasoningclaimed aSessionStarthook told the agent config lives in "AGENTS.md" with the old pre-migration wording — a hook message that appears nowhere in the real transcript (this fixture uses"base": "none", so there's no hook file in play at all) and uses wording this codebase hasn't shipped since before this PR. The exact same failure class as the "Correction" entry below, reproduced with a different fabricated detail. - Pass, same fixture, same prompt, no changes.
- Fail, but not the skill's fault. The agent bypassed the skill's own state entirely and used Claude Code's own out-of-repo auto-memory feature (
Conclusion: no regression traced to the config-relocation change itself. Every failure is either already-documented pre-existing variance, ordinary single-run noise that resolved on re-run, the known activation-gap limitation, a repeat instance of the already-documented judge-hallucination failure mode, or a genuinely new but orthogonal finding filed separately. _base and all migrated fixtures were also verified programmatically (not just via the real-agent runs above) — build_workdir materializes all 72 cases without error, and the handful of trickier ones (contradictory duplicate values, a non-default context: path, deliberately old context-schema versions) were spot-checked to confirm the exact same test condition survived the move to the new file location.
Full docs/evals.md "Latest full-suite results" baseline above is left as the 2026-08-25/56-of-70 numbers — this section is a targeted check for one specific change, not a replacement full-suite run.
Caveats, as stated then
- One run per case: single-run verdicts are subject to normal model variance.
ambiguous-worth-capturing-asks-instead-of-guessingis a demonstrated case of this — it has now failed in both directions across different runs (over-eager write, then silent skip), which is exactly the borderline behavior it's designed to probe. No flakiness statistics yet. - The judge is an LLM from the same vendor as the agent under test. Verdicts were designed to require citing concrete transcript/diff evidence, but an independent judge would be stronger.
- This page covers Claude Code + Claude Sonnet 5 only, one run per case. Cross-agent/cross-model spot checks (Cline, Codex CLI, Kimi Code, opencode, Pi, each against up to 9 models) live separately on the agent & model matrix — one case per combination there, not this full 70-case suite.
- A full
--allrun can take much longer in wall-clock time than its actual compute would suggest if it hits the account's own session/spend limit mid-run — this run took about 6 hours end-to-end, the large majority of it spent asleep between retries on 4 rate-limited cases, not doing work. Seetools/evals/README.md's "Resilience to the account's own session/spend limits" section. - This run's checkout predates the 0.9.1/0.9.2 releases by a few commits.
Diffed against current
main: both releases changed only tooling (new eval drivers, a runner bug fix for a driver-error-masking edge case that doesn't occur in any of these 70 transcripts — every one ends cleanly withsubtype=success) and doc/version-number text, notSKILL.mdorreferences/rule content. These numbers are the current behavioral picture even though they're labeled skill 0.9.0. - Correction, added after publishing this page: the first version of this
section claimed the activation-gap failures happened "despite an explicit
SessionStarthook" telling the agent to load the skill. That was wrong — none of these fixtures configure any hook (the runner deliberately excludes user-level settings via--setting-sources project,local, and no fixture defines a project-level one), confirmed by grepping every raw transcript for hook-related content and finding none. The judge'sreasoningfield invented that detail — quoting a specific, plausible-sounding hook message — in 6 of the 14 failing cases, none of which actually appears anywhere in the real transcript it was grading. Since the judge is an LLM grading against a rubric, not a program checking assertions, a claim in itsreasoningisn't automatically grounded in what it was shown; this is a concrete instance of that, found by cross-checking one specific claim against the raw transcript field rather than trusting the prose. Worth remembering when reading any verdict's reasoning text here or in the raw results.
The runs, as the runner wrote them
The summary.md of each run — every case with its verdict, score and what the judge withheld points for. The transcripts behind them are not published.