Skip to content

Evals: 0.6.2

The first full run, on skill 0.6.2: what the Evals page said from 2026-07-31 until the 0.9.0 series replaced it on 2026-08-25. Text and tables are as published (docs/evals.md at 4436703), not re-worded — "above" and "below" mean the page of that day, and the case count, the rules and the caveats are the ones that applied then. How the suite is run and judged today, and every other series: Evals, run history.

Results

Stale as of this writing — a superseded snapshot, not the current state. The numbers below (59/67, skill 0.6.2) predate two rounds of SKILL.md changes made since (a Core Rules tightening pass, then a fix for a "recognizes the right action but asks/defers instead of doing it" pattern the first full run surfaced — both verified against individual affected cases, neither against a fresh full 67-case run). A new full run is planned as part of the next release cycle; for the current picture in the meantime, see the living status issue: oliver-zehentleitner/keep-the-why#131. This page gets replaced with fresh numbers once that run completes — kept here, unedited otherwise, so the methodology and the shape of the analysis stay visible rather than disappearing while stale.

59 of 67 passed — run 2026-07-31, skill 0.6.2, Claude Code CLI 2.1.220, agent and judge both Claude Sonnet 5 (claude-sonnet-5). No errors; every case produced a graded verdict.

The 8 failures cluster into two honest categories, and both are more useful than a polished 67/67 would be:

1. Activation gaps (5 cases). In all five, the skill was never invoked at all: the session's user request was an ordinary small task (a quick question, a one-line edit, an FYI remark), the model handled it directly, and the skill's opportunistic behaviors — per-session timer checks, proactive capture, re-checking after another workflow concludes — never ran because the skill body never loaded. This is the same activation limitation already documented in "Composition with other skills" and context/compatibility.md: a Skill activates when the conversation matches its description, and nothing guarantees that for low-signal prompts. The evals now put a number on it instead of a caveat.

2. Judgment misses with the skill active (3 cases). The skill loaded, followed its workflow — and still got the judgment call wrong:

  • negative-conflicting-sources (failed in both runs we did): code said 3 retries, the architecture doc said 5; instead of recording both and flagging the conflict as unresolved, the agent declared the code authoritative and rewrote the doc.
  • context-schema-missing-backfilled: asked the user about a missing field that the documented rule says to backfill silently — over-asking is the milder failure direction, but still a miss.
  • ambiguous-worth-capturing-asks-instead-of-guessing: the deliberately borderline case; the agent wrote an entry directly instead of asking the cheap yes/no question first. In an earlier run it failed in the opposite direction — genuinely on the line, which is what the case is for.

Caveats, as stated then

  • One run per case: single-run verdicts are subject to normal model variance (one borderline case demonstrably flips between runs). No flakiness statistics yet.
  • The judge is an LLM from the same vendor as the agent under test. Verdicts were designed to require citing concrete transcript/diff evidence, but an independent judge would be stronger.
  • One agent, one model. Results for Codex CLI, Gemini CLI, and other Agent-Skills-capable tools don't exist yet — running the suite against another agent and publishing the outcome is exactly the contribution CONTRIBUTING.md asks for.