code-v16: early budget nudge + edit routing — 44.3%
The nudge works, the template didn't — text moves nothing this model does.
One control gene + one skill text, both P0 from the v15 analysis
v15 fixed the "ran out of steps unsubmitted" zeros with a 60%/85% budget reminder and a file-level verify gate; v16 moves the first reminder earlier and reroutes edits.
code-v15 --(control.budget_reminder.at_fractions [0.6,0.85] -> [0.25,0.6,0.85]
+ skills/coder: edit routing apply_patch-default -> edit_file-default
+ appended pytest "> /tmp/pt.log 2>&1 && tail" template + no-pipe rule)--> code-v16
Harness chain (harnesses/code-v16.yaml, parent code-v15). Lineage: v14 moved the system prompt into skills/coder (mechanical); v15 added the budget-reminder gene + file-level command_to_verify gate + max_iterations 50→80.
Flat score, changed composition — churn, not drift
Paired on 98 shared instances with the v15 baseline: 47×(0,0), 38×(1,1), net −2 on evaluated pairs — se≈5% at n≈100, same verdict shape as the v9 report.
5 newly fixed
astropy-8872 · psf-1766 / 1921 / 2317 · sympy-12419
The nudge-reachable class: repos that were reminder-adjacent in v15.
7 newly broken
django-15128 · seaborn-3069 · requests-1142 · pylint-6386 · sklearn-12973 / 13439 / 14496
3 of 7 are v15 fallback lucky passes the gate now blocks correctly (see 04).
Per-repo up
psf 29→57% · astropy 40→50% · sympy 30→40% · sphinx 50→56%
Per-repo down
sklearn 73→45% (n=11) · django 60→50% · pylint 22→11% · mwaskom 50→0% (n=2)
pydata / pallets / pytest-dev / matplotlib flat.
Three P0s, one clean win, one partial, one failure
The v16 delta carried three intended changes; their verdicts split exactly along the code-vs-text line.
Early reminder reaches its audience
The step-20 reminder fires while 43/94 trials are still zero-edit — it targets exactly the drift class. Downstream counters all improve: hit80 30→20, budget-exhausted 29→21, no_edit 9→7, and the pass group gets faster (duration −24%, first-edit 13→11) with identical outcomes. The fail group's first-edit is later (18.5→21): the nudge reminds, only judgment commits.
Edit routing shifts load, apply_patch heals
apply_patch calls 39→32 with errors 29→19 (success 26%→41%); edit_file 281→224 at ~97% success. The remaining apply_patch failure class is the `--- a/` prefix (5 rejections) — tolerance + recovery hint is the open fix.
Pipe template: 2% adoption, friction unchanged
The `/tmp/pt.log` template appears in 21/997 (2%) unique pytest commands; pipe rate 85%→87%; pipe-gate rejections 35→38; was-never-run 39→49; −k usage flat (253→257). The 5 newly-fixed trials still converged through pipe/file-less friction (requests-1921: 4 gate rejections then pass) — the gate burns rounds without blocking the capable, while the weak never adopt the template. Verdict: delete the template (done); move pipe friction to the tool/gate code layer. The no-pipe rule was kept out too — rejections prove it is not read.
Sharper teeth: every gold pass now goes through the gate
The v15→v16 changes tightened what a "pass" means — with intended and unintended costs.
Block-without-redirect loses winnable trials
3 of 7 regressions (django-15128, seaborn-3069, requests-1142) are v15's fallback lucky passes — the gate blocks them correctly, but the agent cannot convert to a compliant submit and burns to step 81. Correctness up, score flat: the missing half is redirect.
Rerun failures now catch real failures with detail
5× exit-1 with failing-test excerpts (vs 1 opaque case in v15), e.g. sklearn-12973 "1 failed, 29 passed … AssertionError" — true negatives the v15 gate would have waved through. The remaining hole: gate-passed-but-gold-fail 27→29 (pylint-6386, sklearn-14496: clean submits, wrong patches) — narrow commands exit 0 without covering gold tests.
What the harness can still buy vs what needs a stronger model
The v16 evidence splits the remaining gap into two budgets — order them: harness first, because it also de-noises the metric (v16 removed 11 lucky passes, so a stronger model swapped in now gets an honest benchmark).
Harness-fixable ≈ +3–6pp
Pipe compliance (auto-rewrite at tool or gate) · apply_patch `--- a/` tolerance + recovery hint · grep_search −i/output_mode alias tolerance (10 errors) · fuzzy tool-name match (still 1 shell-line hallucination) · pip-install waste (298 unique calls) · was-never-run redirect (49) · narrow-command FPs (declared-files ⊇ edited-files).
Needs-stronger-model ≈ +8–15pp (expensive)
Clean-submit wrong patches (the 29 fp — root-cause localization no exit-code gate can judge) · zero-edit drift (7 all-fail; 43/94 still unedited at step 20) · repo-complexity gradient (matplotlib 12–17%, pylint 11–22% vs pydata/sklearn 45–64%) · weak feedback loop (sklearn-12973: explicit failure excerpt, still unfixed in 27 steps) · long-context contract following (was-never-run ×49, plain-text submits ×16).
The score is flat — the behavior curve isn't
Ling stays in the 40s across the whole harness series; the same-harness deepseek-v4.1 swap (later runs) shows what the model dimension buys on top of v16.
What v16 established
"The nudge works because it meets the drift class where it lives — step 20, zero edits. The template fails because text cannot make a weak model compliant. Put friction in code, judgment in the model."
step-20 提醒是唯一被证明有效的赢:命中漂移人群、通过组提速 24%。管道模板 2% 的采纳率证明文本合规推不动这个模型——把摩擦下沉到代码层(自动重写 / gate 归一化),把判断留给模型。
Keep the nudge, delete the template, fix the code layer
- Pipe normalization in code — tool-side `| tail` auto-rewrite or gate auto-normalization; text-level attempts are retired (2% adoption).
- apply_patch `--- a/` tolerance + recovery hint — 5 rejections persist; edit_file-default routing stays.
- grep_search alias table — `-i` / `output_mode` tolerance, the observed next blocker (10 errors).
- Validation next run — pipe-gate rejections 38→~0; hit80 20→≤15; was-never-run 49→≤30; watch the 7-regression list (django-15128, seaborn-3069, requests-1142, pylint-6386, sklearn×3).