Benchmark home/code-v16 deepseek-v4.1
EXPERIMENT · EVAL100 · CODE-V16 · MODEL SWAP

code-v16 + deepseek-v4.1-flash — 66/100

Same harness, same code, new model: +20.5pp — the largest jump in the series.

Harness: code-v16 + patch validation + factual catalog + chunked writes Model: deepseek/deepseek-v4.1-flash (OpenRouter) Baseline: ling run3 on identical code — 45/99 (45.5%) 100/100 evaluated · wall ~4.8h jobs/eval100-code-v16-ling-flash-6/ (dirname misnomer)
PASS RATE
66/100 · 0.66
+20.5pp vs same-harness ling · 22 fixed / 2 broken on 99 shared
FORMAT COMPLIANCE
83% apply_patch
29/35 calls succeed · zero grep-alias, zero unknown-tool, zero patch-format errors
PIPE PARADOX
85% piped → 1 rejection
pipes for exploration, declares clean for submission — vs ling's 35–38 rejections
GATE ACCURACY
tp 58 / fp 30 / fn 8
false positives flat vs ling (27–29) — narrow-but-green is model-independent
Verdict: the harness bought ~45%; the model swap bought the next ~20pp. Fix correctness (22 newly-fixed root causes, sympy 30→100%) dwarfs all plumbing effects — the fp mass (30) needs a stronger fixer, not cleaner pipes.
01 · SETUP

Everything fixed except the model

The run pairs against ling run 3 on the identical harness + working tree — only the model changes. Note the job dirname is a misnomer: the model was deepseek-v4.1-flash per operator statement + .env at run time (job artifacts don't record the model name — same gap as the v9 report).

baseline: jobs/eval100-code-v16-ling-flash-3/   Ling-3.0-flash        45/99  (45.5%)
swap:     jobs/eval100-code-v16-ling-flash-6/   deepseek-v4.1-flash   66/100 (66.0%)
          identical: code-v16 (genes 99aea98b) + patch validation + factual tool catalog
                     + chunked writes + containment/parent-check working-tree changes
Wall 10:45→15:34 (~4.8h) · counting basis same as previous reports.
Tools
7 · same as ling run
max_iterations
80
Model
deepseek/deepseek-v4.1-flash
Evaluated
100/100
02 · HEADLINE

+20.5pp, and it is all fix correctness

Paired on 99 shared instances: 43×(1,1), 32×(0,0) — 22 fixed, 2 broken.

22 newly fixed, 8 repos

sympy ×7 (a full 30%→100% sweep) · django ×4 · astropy ×2 · sklearn ×2 · mwaskom ×2 · xarray / psf / pylint / pytest / sphinx singles

2 newly broken — both task-level, not model-level

django-12209 is now zero-edit-fail in 5 consecutive runs (v15, v16, run2, run3, deepseek) across two models — task pathology. django-13568 is the only genuine deepseek regression (1 trial).

Per-repo up

sympy 30→100% · mwaskom 0→100% · sklearn 73→91% · astropy / django / sphinx 50→70%

psf / pallets flat-high; rest flat-or-up.

The model-independent wall

matplotlib 12% in both models — broken out later as its own investigation, not a harness item.

03 · BEHAVIOR PROFILE

Slower trials, correct tool use, near-zero format friction

Deepseek burns time in long tool executions, not extra rounds — ~45% more wall per trial than ling for +20pp.

Durations
pass 421s vs 255s · fail 630s vs 289s
LLM steps
DOWN: 40.9 / 50.4 vs 43.1 / 57.9
read_file
−34% (686 vs 1044), more passes
tool.error
12 (vs ling 15–51)

apply_patch is alive here

35 calls, 29 succ / 6 err (83%) vs ling's ~0–40%. Format compliance is a model property: the parser improvements were necessary but not sufficient — ling never used them, deepseek does. write_file 32 vs 17 (repro scripts); find_files flat-low (14) — this model does not lean on native search (unlike deepseek-v4-0731 on code-v9).

Convergence discipline

no-edit 2 vs 8 · hit80 12 vs 23 · error-ends 1 vs 6 · vpass 88 vs 70. First-edit medians slower (15/22 vs 12/19.5): deepseek explores longer, then commits — only django-12209 / sphinx-9229 never edited.

04 · THE PIPE PARADOX, RESOLVED

85% piped, 1 rejection — compliance is contextual

The ling runs' pipe problem (85% piped, 35–38 rejections) looks completely different under a stronger model on the same gate.

Pipes for exploration, clean declares for submission

718/847 pytest commands (85%) contain pipes — an identical rate to ling — yet pipe-gate rejections are 1. Deepseek pipes for exploration but declares clean commands for submission: context-appropriate compliance that no prompt template ever produced (ling: 2% adoption). This confirms the run-2 verdict that prompt text cannot move this behavior class — a stronger model routes around the gate correctly on its own.

Remaining friction is substantive, not formal

was-never-run 24 (halved from 49) · rerun-caught real test failures 14 (up from 3–10, each with failing-test excerpts the model then iterates on) · budget exhausted 11 · file-less 3.

The structural hole: fp 30

Gate accuracy tp 58 / fp 30 / fn 8 — false positives persist at the same absolute level as ling (27–29). Narrow-but-green declarations remain the hole, independent of model strength; the lever is declared-files ⊇ edited-files, not more friction.

05 · NEW CHECKS

First production firings — and one red flag

Checks behave as designed

write_file parent check fired once (pytest-5787, /tmp/repro/…, trial failed on fix correctness): cost exactly 1 round, no loop — acceptable price, mechanism works. edit match errors ×4 (legit).

Containment red flag: /tmp in 62/100 trials

Deepseek touches /tmp as scratch (repro scripts, whl downloads) in 62/100 trials; ling: 11. Setting AGENT_RUNTIME_WORKSPACE to a repo root without a /tmp exception would break the stronger model's legitimate workflow in a majority of trials — the pending root-containment change must scope repo + tmp (or equivalent). Also: CHANGES/ChangeLog/setup.cfg writes failed 5× with CommandError exit 1 (small files, likely read-only/symlink targets) — surfaced cleanly by the raise-unification; look only if it recurs.

06 · COMPARISON

The two-dimensional map, empirically confirmed

The v16-ling report's rough map (44% → harness ~47–50% → stronger model 55–60%) is confirmed — with the model landing above the projected band.

01020 304050 6070 code-v7 · 41/100 code-v7-2 · 44/100 · noise code-v8 · 41/100 code-v9 · 43/100 code-v9 (dsk) · 40/100 code-v11 (dsk) · 48/100 code-v12 (dsk) · 50/100 code-v16 ling run3 · 45/99 (baseline) code-v16 deepseek-v4.1 (run −6) · 66/100 · THIS RUN code-v16 deepseek-v4.1 run2 · 69/100 41–44 plateau 41 44 41 43 40 48 50 45.5 66 69 v7v7-2v8 v9v9 (dsk)v11 (dsk) v12 (dsk)v16 (ling)v16 (dsk41)r2 (dsk41)
eval100 full runs · this run highlighted. Harness genes + infra fixes carried Ling to ~45.5; the model swap added +20.5pp on top — above the 55–60% band the v16-ling report projected.
07 · CONCLUSION

Two dimensions, one confirmed map

"Harness work bought ~45%; the model swap bought the next ~20pp to 66%. Fix correctness (22 newly-fixed root causes, sympy sweep) dwarfs all plumbing effects — the fp mass needs a stronger fixer, not cleaner pipes."

harness 工程买到了 ~45%,换模型买到了接下来的 ~20pp。修复正确性(22 个新修根因、sympy 全扫)碾压所有管道类效应——fp 主体(30 个)需要更强的修复者,而不是更干净的管道。

08 · NEXT

Backlog reprioritized by this run

  1. Containment must handle /tmp before landing — 62/100 trials use it legitimately; repo-only scoping would break the stronger model.
  2. Alias table: moot for this model class (0 hits) but cheap — keep for weaker ones.
  3. Gate side: declared-files ⊇ edited-files (item 7) against the 30 fp — not more friction.
  4. Validation for next run — hold ≥60% with any model ≥ this class; was-never-run 24→≤15 via the auto-attach suggestion (item 6); matplotlib wall (12% across all four runs and both models) broken out as its own investigation.
09 · REPRODUCTION

Artifacts & record

Job
jobs/eval100-code-v16-ling-flash-6/
Baseline
jobs/eval100-code-v16-ling-flash-3/
Record
eval100-code-v16-ling-flash-6__code-v16__dataset-eval100.json
Report
evaluation/analysis/eval100-code-v16-deepseek-v4-1.md
The record's model field (deepseekv4.1-flash) confirms the operator statement; the job dirname "ling-flash-6" is a misnomer retained for artifact traceability. Record: 66 pass / 34 fail / 0 errored, rate 0.66 — matches the report exactly.
← Previousrun 3 · chunked write — 45.5%, best Ling Back to leaderboard ↑SunAgent Harness Benchmark home Next →run 2 · same model, drift 69/100 — noise rules adopted