Benchmark home/code-v16 ling-flash run 3
EXPERIMENT · EVAL100 · CODE-V16 · INFRA FIX VERIFIED

code-v16 run 3: chunked writing — 45.5%

Same harness, same model, zero new behavior — the infra fix alone lifted 43.0 → 45.5.

Code delta 5f99aa3: HarborWorkspace.write_file chunked Harness: code-v16 unchanged Model: Ling-3.0-flash Baseline: run 2 (43.0) 99/100 evaluated · 1 errored trial excluded jobs/eval100-code-v16-ling-flash-3/
PASS RATE
45/99 · 0.455
+2.5pp vs run 2 · 8 fixed / 5 broken — best v16 Ling run
WRITE-FAIL
10 → 0
zero E2BIG errors · apply_patch errors 21→0, usage 21→1 call (1/1 success)
TOOL ERRORS
37 → 15 (−60%)
error-status trials 10→6 · no-edit 11→8 · memory summaries 97→76
NEXT BLOCKER
grep aliases ×7
output_mode ×5 + -i ×2 — now the #1 tool-error category by itself
Verdict: the ceiling argument is confirmed — with large writes unblocked, the same harness+model gained +2.5pp from an infra fix alone. apply_patch's run-2 0% success is fully explained (9 infra + 12 correct rejections); no parser over-blocking occurred.
01 · SETUP

One infra fix, byte-identical below the threshold

Committed 5f99aa3 before the job: ≤64KB encoded writes stay the legacy single command (byte-identical); above, the write splits into 32KB pieces + decode-and-move. Harness still code-v16.

5f99aa3: HarborWorkspace.write_file chunked writing
  ≤64KB encoded  -> legacy single command (byte-identical)
  >64KB          -> 32KB pieces + decode-and-move
Backend-agnostic, zero task knowledge — the fix direction scoped in the run-2 report, now verified in production.
Harness
code-v16, unchanged
Model
Ling-3.0-flash
Evaluated
99/100 · 1 errored excluded
Counting
same as previous reports
02 · HEADLINE

+2.5pp, concentrated movers

Paired on 99 shared instances with run 2: 49×(0,0), 37×(1,1), 8 fixed / 5 broken.

8 newly fixed

django-12209 · requests-1724 · xarray-6938 · pylint-6386 · sklearn-11578 / 14053 / 25102 / 25973

xarray-6938 is the run-2 report's documented write-fail casualty; django-12209 was zero-edit since v15. Recovery, not drift.

5 newly broken

django-14007 · seaborn-3069 · xarray-6744 · sklearn-12973 / 13124

Per-repo up (concentrated)

sklearn 55→73% · psf 71→86% · pylint 11→22%

astropy / django / matplotlib / pallets / pydata / pytest-dev / sphinx / sympy counts exactly flat; mwaskom 50→0% (n=2 — seaborn-3069 ping-pongs pass→fail→pass→fail across the four runs).

What did NOT change

First-edit medians flat (11 / 19.5). The gain came from removing friction, not faster convergence — the signature of an infra fix, not a behavior change.

03 · MECHANISM

The fix verified, the remainder named

Write path clean

Zero "Argument list too long" errors (run 2: 10 trials; run 1: 12 hidden). apply_patch errors 21→0 with usage collapsing to a single successful call — effectively retired by the edit_file-default routing, and its former error mass did not migrate elsewhere.

Remaining error mass (15 trials)

output_mode ×5 + -i ×2 (item 3 alias table — now #1 by itself) · bare-grep hallucinations ×5 (variance; the factual catalog repairs, never prevents) · bash_execute ×1 · legit edit_file match errors ×2. Fail-group duration improved (370→289s mean) while pass-group held (255s).

Memory summaries 97→76: fewer error/retry loops to summarize. Consistent with friction removal.
04 · COMPARISON

Best Ling run of the series

01020 304050 6070 code-v7 · 41/100 code-v7-2 · 44/100 · noise code-v8 · 41/100 code-v9 · 43/100 code-v9 (dsk) · 40/100 code-v11 (dsk) · 48/100 code-v12 (dsk) · 50/100 code-v16 ling run3 · 45/99 · 0.455 · THIS RUN code-v16 deepseek-v4.1 (run −6) · 66/100 code-v16 deepseek-v4.1 run2 · 69/100 41–44 plateau 41 44 41 43 40 48 50 45.5 66 69 v7v7-2v8 v9v9 (dsk)v11 (dsk) v12 (dsk)v16 (ling)v16 (dsk41)r2 (dsk41)
eval100 full runs · this run highlighted (45.5, best Ling). The Ling cluster tops out ~45.5 on v16 — the model dimension is the next lever.
05 · CONCLUSION

Ceiling argument confirmed

"With large writes unblocked, the same harness+model went from 43.0% to 45.5% with zero new agent behavior. The next binding constraint is unambiguous: grep argument aliases — then pipe normalization."

分块写入修复后,同一 harness+模型、零新行为,43.0% → 45.5%。下一个约束点已经明确:grep 参数别名表,然后是管道归一化。工具错误总量已小到(15/100)需要权衡:继续修 harness 还是换更强的模型。

Tool-error mass is small enough (15/100 trials) that further harness work should be weighed against model-side gains — 29 prior false positives need a stronger fixer, not cleaner plumbing. (The deepseek-v4.1 swap on this exact codebase follows next.)
06 · NEXT

Alias table, then pipe normalization

  1. Item 3 alias table — output_mode / -i tolerance for grep_search; the #1 remaining tool-error category (7 of 15).
  2. Item 1 pipe normalization — tool-side auto-rewrite or gate auto-normalization.
  3. Validation next run — output_mode/-i errors 7→0; headline 45.5% held or better with the same concentrated-mover pattern (sklearn/psf/pylint) rather than broad drift.
07 · REPRODUCTION

Artifacts

Job
jobs/eval100-code-v16-ling-flash-3/
Baseline
jobs/eval100-code-v16-ling-flash-2/ (43.0)
Code
commit 5f99aa3
Report
evaluation/analysis/eval100-code-v16-ling-flash-3.md
No evaluation/records snapshot for this run; data as published in the analysis report. This run's 45/99 is also the baseline the deepseek-v4.1 model-swap report pairs against.
← Previousrun 2 · patch validation — infra bug exposed Back to leaderboard ↑SunAgent Harness Benchmark home Next →deepseek-v4.1 swap · 66/100 — +20.5pp on the same harness