Benchmark home/code-v16 ling-flash
EXPERIMENT · EVAL100 · CODE-V16 · BUDGET NUDGE

code-v16: early budget nudge + edit routing — 44.3%

The nudge works, the template didn't — text moves nothing this model does.

Harness: code-v16 = code-v15 + reminder@25% + skill edit routing Model: Ling-3.0-flash Baseline: code-v15-ling 45.0 97/100 evaluated · 3 infra losses jobs/eval100-code-v16-ling-flash/
PASS RATE
43/97 · 0.443
−0.7pp vs v15 — 5 newly fixed / 7 newly broken (noise, se≈5%)
STEP-20 NUDGE
hit80 30→20
budget-exhausted 29→21 · no_edit 9→7 · reminder fired in 94/100 trials
PASS-GROUP SPEED
−24%
duration 309/296s → 236/210s (mean/med) · llm steps 51.6→42.2 · tools 50.6→41.4
PIPE TEMPLATE
2% adoption
21/997 commands · rejections 35→38 · was-never-run 39→49 — deleted post-analysis
Verdict: flat score, sharp composition change — the early nudge is the one clear win; the pipe template proves text-only compliance does not move this model. Friction must move to the code layer; the gate now vets everything and catches real failures with excerpts.
01 · SETUP

One control gene + one skill text, both P0 from the v15 analysis

v15 fixed the "ran out of steps unsubmitted" zeros with a 60%/85% budget reminder and a file-level verify gate; v16 moves the first reminder earlier and reroutes edits.

code-v15 --(control.budget_reminder.at_fractions [0.6,0.85] -> [0.25,0.6,0.85]
             + skills/coder: edit routing apply_patch-default -> edit_file-default
             + appended pytest "> /tmp/pt.log 2>&1 && tail" template + no-pipe rule)--> code-v16
Harness chain (harnesses/code-v16.yaml, parent code-v15). Lineage: v14 moved the system prompt into skills/coder (mechanical); v15 added the budget-reminder gene + file-level command_to_verify gate + max_iterations 50→80.
Tools
7 unchanged · submit_result since v13
max_iterations
80 (since v15)
Model
Ling-3.0-flash, fixed
Tokens
null in result.json — no $ comparison
3 infra losses excluded: 2× EnvironmentStartTimeoutError (matplotlib-24149/25775, no agent log) + 1× VerifierTimeoutError (sphinx-7985). Counting basis: tool.start authoritative; unique run_commands deduped by _id; rewards keyed by task prefix.
02 · HEADLINE

Flat score, changed composition — churn, not drift

Paired on 98 shared instances with the v15 baseline: 47×(0,0), 38×(1,1), net −2 on evaluated pairs — se≈5% at n≈100, same verdict shape as the v9 report.

5 newly fixed

astropy-8872 · psf-1766 / 1921 / 2317 · sympy-12419

The nudge-reachable class: repos that were reminder-adjacent in v15.

7 newly broken

django-15128 · seaborn-3069 · requests-1142 · pylint-6386 · sklearn-12973 / 13439 / 14496

3 of 7 are v15 fallback lucky passes the gate now blocks correctly (see 04).

Per-repo up

psf 29→57% · astropy 40→50% · sympy 30→40% · sphinx 50→56%

Per-repo down

sklearn 73→45% (n=11) · django 60→50% · pylint 22→11% · mwaskom 50→0% (n=2)

pydata / pallets / pytest-dev / matplotlib flat.

03 · MECHANISM

Three P0s, one clean win, one partial, one failure

The v16 delta carried three intended changes; their verdicts split exactly along the code-vs-text line.

P0-3 nudge
WORKS — pass group −24%
P0-1 routing
PARTIAL — 5× `--- a/` rejections persist
P0-2 template
FAILED — deleted post-analysis
Lesson
compliance must live in code, not prompts

Early reminder reaches its audience

The step-20 reminder fires while 43/94 trials are still zero-edit — it targets exactly the drift class. Downstream counters all improve: hit80 30→20, budget-exhausted 29→21, no_edit 9→7, and the pass group gets faster (duration −24%, first-edit 13→11) with identical outcomes. The fail group's first-edit is later (18.5→21): the nudge reminds, only judgment commits.

Edit routing shifts load, apply_patch heals

apply_patch calls 39→32 with errors 29→19 (success 26%→41%); edit_file 281→224 at ~97% success. The remaining apply_patch failure class is the `--- a/` prefix (5 rejections) — tolerance + recovery hint is the open fix.

Pipe template: 2% adoption, friction unchanged

The `/tmp/pt.log` template appears in 21/997 (2%) unique pytest commands; pipe rate 85%→87%; pipe-gate rejections 35→38; was-never-run 39→49; −k usage flat (253→257). The 5 newly-fixed trials still converged through pipe/file-less friction (requests-1921: 4 gate rejections then pass) — the gate burns rounds without blocking the capable, while the weak never adopt the template. Verdict: delete the template (done); move pipe friction to the tool/gate code layer. The no-pipe rule was kept out too — rejections prove it is not read.

04 · GATE SEMANTICS

Sharper teeth: every gold pass now goes through the gate

The v15→v16 changes tightened what a "pass" means — with intended and unintended costs.

Lucky passes
11 → 0
vpass events
61 → 72
Rerun exit-1 (real)
1 → 5, with excerpts
Gate FPs
27 → 29

Block-without-redirect loses winnable trials

3 of 7 regressions (django-15128, seaborn-3069, requests-1142) are v15's fallback lucky passes — the gate blocks them correctly, but the agent cannot convert to a compliant submit and burns to step 81. Correctness up, score flat: the missing half is redirect.

Rerun failures now catch real failures with detail

5× exit-1 with failing-test excerpts (vs 1 opaque case in v15), e.g. sklearn-12973 "1 failed, 29 passed … AssertionError" — true negatives the v15 gate would have waved through. The remaining hole: gate-passed-but-gold-fail 27→29 (pylint-6386, sklearn-14496: clean submits, wrong patches) — narrow commands exit 0 without covering gold tests.

05 · TWO DIMENSIONS

What the harness can still buy vs what needs a stronger model

The v16 evidence splits the remaining gap into two budgets — order them: harness first, because it also de-noises the metric (v16 removed 11 lucky passes, so a stronger model swapped in now gets an honest benchmark).

Harness-fixable ≈ +3–6pp

Pipe compliance (auto-rewrite at tool or gate) · apply_patch `--- a/` tolerance + recovery hint · grep_search −i/output_mode alias tolerance (10 errors) · fuzzy tool-name match (still 1 shell-line hallucination) · pip-install waste (298 unique calls) · was-never-run redirect (49) · narrow-command FPs (declared-files ⊇ edited-files).

Needs-stronger-model ≈ +8–15pp (expensive)

Clean-submit wrong patches (the 29 fp — root-cause localization no exit-code gate can judge) · zero-edit drift (7 all-fail; 43/94 still unedited at step 20) · repo-complexity gradient (matplotlib 12–17%, pylint 11–22% vs pydata/sklearn 45–64%) · weak feedback loop (sklearn-12973: explicit failure excerpt, still unfixed in 27 steps) · long-context contract following (was-never-run ×49, plain-text submits ×16).

Rough map from 44%: harness → ~47–50%, a stronger model on top → 55–60% band. (The later deepseek-v4.1 swap on this same harness would land 66–69% — above the map's band.)
06 · COMPARISON

The score is flat — the behavior curve isn't

Ling stays in the 40s across the whole harness series; the same-harness deepseek-v4.1 swap (later runs) shows what the model dimension buys on top of v16.

01020 304050 6070 code-v7 · 41/100 code-v7-2 · 44/100 · noise code-v8 · 41/100 code-v9 · 43/100 code-v9 (dsk) · 40/100 code-v11 (dsk) · 48/100 code-v12 (dsk) · 50/100 code-v16 ling run1 · 43/97 · 0.443 · THIS RUN code-v16 deepseek-v4.1 (run −6) · 66/100 code-v16 deepseek-v4.1 run2 · 69/100 41–44 plateau 41 44 41 43 40 48 50 43 66 69 v7v7-2v8 v9v9 (dsk)v11 (dsk) v12 (dsk)v16 (ling)v16 (dsk41)r2 (dsk41)
eval100 full runs · this run highlighted · deepseek-v4.1 bars are the later same-harness model swap (run −6: 66; run2: 69). Relative bar lengths are for comparison only — labeled values are authoritative.
07 · CONCLUSION

What v16 established

"The nudge works because it meets the drift class where it lives — step 20, zero edits. The template fails because text cannot make a weak model compliant. Put friction in code, judgment in the model."

step-20 提醒是唯一被证明有效的赢:命中漂移人群、通过组提速 24%。管道模板 2% 的采纳率证明文本合规推不动这个模型——把摩擦下沉到代码层(自动重写 / gate 归一化),把判断留给模型。

08 · NEXT

Keep the nudge, delete the template, fix the code layer

  1. Pipe normalization in code — tool-side `| tail` auto-rewrite or gate auto-normalization; text-level attempts are retired (2% adoption).
  2. apply_patch `--- a/` tolerance + recovery hint — 5 rejections persist; edit_file-default routing stays.
  3. grep_search alias table — `-i` / `output_mode` tolerance, the observed next blocker (10 errors).
  4. Validation next run — pipe-gate rejections 38→~0; hit80 20→≤15; was-never-run 49→≤30; watch the 7-regression list (django-15128, seaborn-3069, requests-1142, pylint-6386, sklearn×3).
09 · REPRODUCTION

Artifacts & lineage

Job
jobs/eval100-code-v16-ling-flash/
Baseline
jobs/eval100-code-v15-ling-flash/ (0.45)
Record
evaluation/records/eval100-code-v16-ling-flash__code-v16__dataset-eval100.json
Report
evaluation/analysis/eval100-code-v16-ling-flash.md
Standing constraint (all harness work): no task-specific logic in runtime code — no repo names, no per-task branches, no benchmark-only string matches in src/agent_runtime/. Audit 2026-09-24: clean (pytest/testbed appear only as description examples). Next refactor target: promote the pytest-shaped constants to per-harness verification genes.
← Previouseval100-code-v13-partial · submit_result, 83/100 judged Back to leaderboard ↑SunAgent Harness Benchmark home Next →run 2 · patch validation + tool catalog — flat, infra bug found