Skill evals¶
Evaluation suites for the agent skills that are bundled with the CLI. Models come from the Wagtail agentic engineering recommendations.
For each skill, we run two suites: a baseline with nothing loaded, and one with our skills. We test:
- Skill activation: whether the skill activates when mentioning potential related tasks, without necessarily directly prompting for "use Wagtail CLI".
- Task completion: whether the CLI commands provided by the agent actually exist, and work as explained, and help in completing the task.
- Gotchas with no command answer. Graded by rubric, using a separate model family.
Harnesses¶
We run these suites with two tools, on the same tasks and the same model, so the results can be compared and either number trusted:
| what it grades | what it is good for | |
|---|---|---|
| Promptfoo | the agent's written answer | command correctness and rubric grading |
| Coder Eval | the agent's trajectory in a sandbox | time, tokens and tool calls per arm |
Promptfoo is the primary harness. Coder Eval runs alongside it to add the efficiency metrics Promptfoo cannot report, and to check the two agree on what "correct" means. Neither replaces the other today; see coder-eval/README.md for the trade-offs.
Requirements¶
- OpenCode CLI
wton PATH. The graders run the CLI from there.TENSORX_API_KEY— the model under test and the rubric grader both run on TensorX.just eval-initinstalls the harness tooling (Promptfoo and Coder Eval).
Running¶
just eval # Promptfoo: both suites
just eval docs/evals/wagtail_api_skill.yaml # Promptfoo: one suite
just eval --repeat 3 # agent runs are noisy; repeat before trusting a delta
just eval-view # Promptfoo dashboard for the latest run
just eval-coder # Coder Eval: the same tasks, both arms
just eval-coder --repeats 5 # raise the replicate count
just eval-coder-report # Coder Eval report for the latest run
EVAL_MODEL picks the Promptfoo model under test; it must be registered in opencode.json (default z-ai/glm-5.3-flash):
The Coder Eval model under test is set in coder-eval/experiments/wagtail_skills_ab.yaml, and its rubric judge in each docs_*.yaml task's checker_context.
Known friction¶
Things a future maintainer will run into, and what to do about them:
- The two harnesses share their match rules, by design. A task's
expect_requestslives in the Promptfoo YAML and in the Coder Eval task'srun_commandline; both importgraders/_request_matching.pyso they cannot drift. If you change a task's expectation, change it in both places. - Coder Eval's rubric judge needs the
litellmextra.just eval-initinstalls it (uv tool install coder-eval --with litellm). Without it the judge fails loudly rather than scoring 0 silently. - Coder Eval's with-skill arm needs
$SKILLS_PATH.just eval-coderexports it and fails fast if it is unset; if you invokecoder-evaldirectly, export it yourself or the arm silently measures the bare model. wt docshas no--outlineflag. Thewagtail-docsskill's SKILL.md suggests one that the CLI does not implement. Agents that follow the example produce a failed command; acommand_executedcriterion withrequire_success: truecatches that, but the skill text should be fixed.- OpenCode + Docker is unsupported in Coder Eval (the CLI is not in its
image), so the sandbox is a tempdir — a working directory, not a confinement
boundary. OpenCode also ignores
allowed_tools/system_prompt; the baseline arm simply loads no skills rather than restricting tools. - Both harnesses are local-only. Neither runs in CI; there is no scheduled job and no API key configured there. Committing a run means recording it by hand in this file.
wt api schema showprints a Python repr, not JSON, under--dry-run. The Coder Eval grader skips such lines. Keep it in mind if a recorded trajectory ever looks like it is missing a command.
Results snapshot¶
Promptfoo numbers, post-leak-fix (see caveats): --no-cache, promptfoo 0.123.1, baseline with all filesystem tools disabled, skill arm with read scoped to the skills directory. Single runs — confirm deltas with --repeat 3 before acting. The pre-fix snapshot (2026-09-19) measured all three models with the baseline able to read the repo, so those numbers are not comparable and were dropped. "Activation" is the two skill-arm rows asserting the skill loads (or does not).
wagtail-api (8 graded rows + 2 activation)¶
| Model | baseline | skill | activation |
|---|---|---|---|
| z-ai/glm-5.3-flash (eval-fIc-2026-09-21T16:09:08) | 1/8 | 8/10 | 2/2 |
| qwen/qwen3.8-27b (eval-RNf-2026-09-21T16:14:24) | 1/8 | 8/10 | 2/2 |
| qwen/qwen3.5-9b (eval-FZi-2026-09-21T16:32:38) | 1/8 | 5/10 | 2/2 |
- Every baseline passes exactly one row: the StreamField replace-whole rubric, answered from generic Wagtail knowledge. Every command row now fails honestly — the models suggest
git cloneof the Wagtail repo or generic curl against the API instead ofwt. - Every skill arm misses the create row: the models follow the skill's schema-check strategy but still build partly generic payloads (
size: h2on the heading, noblog_person_relationship) instead of the demo's shapes — the persistent hard row across configs and models. - glm's other miss (list) is new this run and passed in earlier runs of the same config (eval-6go-2026-09-21T15:51:39, 9/10) — noise; treat one-run deltas as indicative. The intermediate all-read-disabled config (eval-ChC-2026-09-21T15:06:49) also failed update-draft and image-upload because the command syntax in
references/commands.mdwas unreachable; the scoped read restores those. - qwen3.8-27b's other miss (unpublish) was an empty response — provider error, not an answer (its baseline hit one on the same row). qwen3.5-9b is the local-model test case: activation works, but without reliable command syntax it fails five rows (
--dry-runrewrites, schema checks andreferences/lookups notwithstanding).
wagtail-docs (4 graded rows + 2 activation), z-ai/glm-5.3-flash, eval-Y0n-2026-09-21T15:06:49¶
| Model | baseline | skill | activation |
|---|---|---|---|
| z-ai/glm-5.3-flash | 0/4 | 4/4 | 2/2 |
- Baseline fails every row from pure memory (clone-the-repo and curl advice, invented docs paths); the skill arm passes everything — the intended contrast. Numbers are from the intermediate all-read-disabled config; the scoped read only adds access the skill arm already used.
Coder Eval¶
tensorx/deepseek/deepseek-v4.1-flash, 5 replicates per (task, arm), 8 tasks × 2
arms, graded by qwen/qwen3.8-flash-next (a different family from the model
under test). Run 2026-09-29_17-00-11. Bold marks the better arm.
| Metric | baseline | with-skill |
|---|---|---|
| Mean score | 0.548 ± 0.172 | 0.925 ± 0.074 |
| Tasks won | 0/8 | 8/8 |
| Task pass rate | 0% | 37.5% |
| Tool calls per run | 13.0 | 8.4 |
| Assistant turns per run | 10.2 | 8.2 |
| Tokens per run | 218,188 | 166,784 |
| Duration per run | 72.5s | 54.2s |
Paired mean difference (baseline − with-skill): −0.378 (95% CI −0.505 to −0.251, Cohen's d = −2.49, p < 0.001). The skill arm wins every task, uses ~24% fewer tokens, ~35% fewer tool calls, and ~25% less wall-clock time.
Per-task scores (baseline → with-skill):
| Task | baseline | with-skill |
|---|---|---|
unpublish_not_delete |
0.667 | 1.000 |
list_blog_posts |
0.600 | 1.000 |
docs_images_topic |
0.417 | 0.984 |
publish_blog_post |
0.200 | 0.800 |
update_draft_no_publish |
0.600 | 0.867 |
docs_v3_create_operation |
0.730 | 0.946 |
docs_streamfield_validation |
0.502 | 0.940 |
upload_image |
0.667 | 0.867 |
Notes:
- The skill arm still misses
publish_blog_poston 2/5 replicates andupdate_draft_no_publish/upload_imageon 1/5 each — the same partial-payload failures the Promptfoo suite documents for its create row, which is a useful cross-harness agreement. - Almost every run overshoots its
commands_efficiencybudget. That criterion isweight: 0(informational), so it does not gate, but it shows both arms loop more than the skill's intendedwhoami → schema → createpath. - Two with-skill rows scored on a judge hiccup: one got no verdict
(
Judge did not call submit_verdict), one a low score with a reasoning rationale. Treat single judge rows as noisy; the paired aggregate is the signal. - Baseline never passes a task (0/40 replicates); its 0.548 mean comes from partial credit on the non-gating criteria.
Last updated: 2026-09-29.