code-v16 + deepseek-v4.1-flash — 66/100
Same harness, same code, new model: +20.5pp — the largest jump in the series.
Everything fixed except the model
The run pairs against ling run 3 on the identical harness + working tree — only the model changes. Note the job dirname is a misnomer: the model was deepseek-v4.1-flash per operator statement + .env at run time (job artifacts don't record the model name — same gap as the v9 report).
baseline: jobs/eval100-code-v16-ling-flash-3/ Ling-3.0-flash 45/99 (45.5%)
swap: jobs/eval100-code-v16-ling-flash-6/ deepseek-v4.1-flash 66/100 (66.0%)
identical: code-v16 (genes 99aea98b) + patch validation + factual tool catalog
+ chunked writes + containment/parent-check working-tree changes
Wall 10:45→15:34 (~4.8h) · counting basis same as previous reports.
+20.5pp, and it is all fix correctness
Paired on 99 shared instances: 43×(1,1), 32×(0,0) — 22 fixed, 2 broken.
22 newly fixed, 8 repos
sympy ×7 (a full 30%→100% sweep) · django ×4 · astropy ×2 · sklearn ×2 · mwaskom ×2 · xarray / psf / pylint / pytest / sphinx singles
2 newly broken — both task-level, not model-level
django-12209 is now zero-edit-fail in 5 consecutive runs (v15, v16, run2, run3, deepseek) across two models — task pathology. django-13568 is the only genuine deepseek regression (1 trial).
Per-repo up
sympy 30→100% · mwaskom 0→100% · sklearn 73→91% · astropy / django / sphinx 50→70%
psf / pallets flat-high; rest flat-or-up.
The model-independent wall
matplotlib 12% in both models — broken out later as its own investigation, not a harness item.
Slower trials, correct tool use, near-zero format friction
Deepseek burns time in long tool executions, not extra rounds — ~45% more wall per trial than ling for +20pp.
apply_patch is alive here
35 calls, 29 succ / 6 err (83%) vs ling's ~0–40%. Format compliance is a model property: the parser improvements were necessary but not sufficient — ling never used them, deepseek does. write_file 32 vs 17 (repro scripts); find_files flat-low (14) — this model does not lean on native search (unlike deepseek-v4-0731 on code-v9).
Convergence discipline
no-edit 2 vs 8 · hit80 12 vs 23 · error-ends 1 vs 6 · vpass 88 vs 70. First-edit medians slower (15/22 vs 12/19.5): deepseek explores longer, then commits — only django-12209 / sphinx-9229 never edited.
85% piped, 1 rejection — compliance is contextual
The ling runs' pipe problem (85% piped, 35–38 rejections) looks completely different under a stronger model on the same gate.
Pipes for exploration, clean declares for submission
718/847 pytest commands (85%) contain pipes — an identical rate to ling — yet pipe-gate rejections are 1. Deepseek pipes for exploration but declares clean commands for submission: context-appropriate compliance that no prompt template ever produced (ling: 2% adoption). This confirms the run-2 verdict that prompt text cannot move this behavior class — a stronger model routes around the gate correctly on its own.
Remaining friction is substantive, not formal
was-never-run 24 (halved from 49) · rerun-caught real test failures 14 (up from 3–10, each with failing-test excerpts the model then iterates on) · budget exhausted 11 · file-less 3.
The structural hole: fp 30
Gate accuracy tp 58 / fp 30 / fn 8 — false positives persist at the same absolute level as ling (27–29). Narrow-but-green declarations remain the hole, independent of model strength; the lever is declared-files ⊇ edited-files, not more friction.
First production firings — and one red flag
Checks behave as designed
write_file parent check fired once (pytest-5787, /tmp/repro/…, trial failed on fix correctness): cost exactly 1 round, no loop — acceptable price, mechanism works. edit match errors ×4 (legit).
Containment red flag: /tmp in 62/100 trials
Deepseek touches /tmp as scratch (repro scripts, whl downloads) in 62/100 trials; ling: 11. Setting AGENT_RUNTIME_WORKSPACE to a repo root without a /tmp exception would break the stronger model's legitimate workflow in a majority of trials — the pending root-containment change must scope repo + tmp (or equivalent). Also: CHANGES/ChangeLog/setup.cfg writes failed 5× with CommandError exit 1 (small files, likely read-only/symlink targets) — surfaced cleanly by the raise-unification; look only if it recurs.
The two-dimensional map, empirically confirmed
The v16-ling report's rough map (44% → harness ~47–50% → stronger model 55–60%) is confirmed — with the model landing above the projected band.
Two dimensions, one confirmed map
"Harness work bought ~45%; the model swap bought the next ~20pp to 66%. Fix correctness (22 newly-fixed root causes, sympy sweep) dwarfs all plumbing effects — the fp mass needs a stronger fixer, not cleaner pipes."
harness 工程买到了 ~45%,换模型买到了接下来的 ~20pp。修复正确性(22 个新修根因、sympy 全扫)碾压所有管道类效应——fp 主体(30 个)需要更强的修复者,而不是更干净的管道。
Backlog reprioritized by this run
- Containment must handle /tmp before landing — 62/100 trials use it legitimately; repo-only scoping would break the stronger model.
- Alias table: moot for this model class (0 hits) but cheap — keep for weaker ones.
- Gate side: declared-files ⊇ edited-files (item 7) against the 30 fp — not more friction.
- Validation for next run — hold ≥60% with any model ≥ this class; was-never-run 24→≤15 via the auto-attach suggestion (item 6); matplotlib wall (12% across all four runs and both models) broken out as its own investigation.