Skip to content

Eval run — 2026-09-03

Skill 0.11.0, run 3, 73/74 — part of the 0.11.0 page. What follows is the summary.md the runner wrote for this run, as it was written (only < is escaped, so that it renders): one row per case, the judge's own words, cut off where the runner cut them.

Skill 0.11.0 · agent: Claude Code (model sonnet) · judge: sonnet · permission bypass: --dangerously-skip-permissions

73/74 passed (1 failed, 0 errors)

Restraint categories (mechanical, not judge-scored): checked_honestly_then_acted: 30, never_checked_then_acted: 2, restrained: 42

Case Verdict Score Notes
ambiguous-worth-capturing-asks-instead-of-guessing fail 3 The agent did not write anything to disk (git status clean), which avoids the worst failure mode. But it opened with 'This is exactly the kind of thing worth a short context/ entry' — explicitly annou