code-v16 run 3: chunked writing — 45.5%
Same harness, same model, zero new behavior — the infra fix alone lifted 43.0 → 45.5.
One infra fix, byte-identical below the threshold
Committed 5f99aa3 before the job: ≤64KB encoded writes stay the legacy single command (byte-identical); above, the write splits into 32KB pieces + decode-and-move. Harness still code-v16.
5f99aa3: HarborWorkspace.write_file chunked writing ≤64KB encoded -> legacy single command (byte-identical) >64KB -> 32KB pieces + decode-and-moveBackend-agnostic, zero task knowledge — the fix direction scoped in the run-2 report, now verified in production.
+2.5pp, concentrated movers
Paired on 99 shared instances with run 2: 49×(0,0), 37×(1,1), 8 fixed / 5 broken.
8 newly fixed
django-12209 · requests-1724 · xarray-6938 · pylint-6386 · sklearn-11578 / 14053 / 25102 / 25973
xarray-6938 is the run-2 report's documented write-fail casualty; django-12209 was zero-edit since v15. Recovery, not drift.
5 newly broken
django-14007 · seaborn-3069 · xarray-6744 · sklearn-12973 / 13124
Per-repo up (concentrated)
sklearn 55→73% · psf 71→86% · pylint 11→22%
astropy / django / matplotlib / pallets / pydata / pytest-dev / sphinx / sympy counts exactly flat; mwaskom 50→0% (n=2 — seaborn-3069 ping-pongs pass→fail→pass→fail across the four runs).
What did NOT change
First-edit medians flat (11 / 19.5). The gain came from removing friction, not faster convergence — the signature of an infra fix, not a behavior change.
The fix verified, the remainder named
Write path clean
Zero "Argument list too long" errors (run 2: 10 trials; run 1: 12 hidden). apply_patch errors 21→0 with usage collapsing to a single successful call — effectively retired by the edit_file-default routing, and its former error mass did not migrate elsewhere.
Remaining error mass (15 trials)
output_mode ×5 + -i ×2 (item 3 alias table — now #1 by itself) · bare-grep hallucinations ×5 (variance; the factual catalog repairs, never prevents) · bash_execute ×1 · legit edit_file match errors ×2. Fail-group duration improved (370→289s mean) while pass-group held (255s).
Best Ling run of the series
Ceiling argument confirmed
"With large writes unblocked, the same harness+model went from 43.0% to 45.5% with zero new agent behavior. The next binding constraint is unambiguous: grep argument aliases — then pipe normalization."
分块写入修复后,同一 harness+模型、零新行为,43.0% → 45.5%。下一个约束点已经明确:grep 参数别名表,然后是管道归一化。工具错误总量已小到(15/100)需要权衡:继续修 harness 还是换更强的模型。
Alias table, then pipe normalization
- Item 3 alias table — output_mode / -i tolerance for grep_search; the #1 remaining tool-error category (7 of 15).
- Item 1 pipe normalization — tool-side auto-rewrite or gate auto-normalization.
- Validation next run — output_mode/-i errors 7→0; headline 45.5% held or better with the same concentrated-mover pattern (sklearn/psf/pylint) rather than broad drift.