code-v16 run 2: patch validation + factual tool catalog — 43.0%
Generic validation, precise hints — and an infra bug with a 10-trials-per-round price.
Two code-layer items, harness untouched
Committed as 17d2fd2 before the job; the skill is the template-removed version (per run 1's verdict). Item 2 = patch validation, item 4 = factual tool catalog.
17d2fd2:
patch.py unified-diff sniff · path legality precheck
read/write-failure raise unification (was: return {"error"})
create-requires-existing-parent
tool_dispatch.py unknown-tool -> factual catalog, no did-you-mean
Working-tree delta under test. Harness stays code-v16 (genes unchanged); skill = template-removed version.
Flat score, symmetric churn
Paired on 98 shared instances with run 1: 48×(0,0), 37×(1,1) — perfect symmetry around zero.
6 newly fixed
seaborn-3069 · requests-1142 · sklearn-12973 / 13124 / 13439 / 14496
Includes three of run 1's seven regressions — churn, not trend.
6 newly broken
xarray-6599 / 6938 · sklearn-11578 / 14053 / 25973 · sympy-12419
Per-repo up
psf 57→71% · sklearn 45→55% · mwaskom 0→50%
Per-repo down / drift signals
pydata 64→45% · sympy 40→30%
no-edit 7→11 · hit80 20→24 · error-status 5→10 · grep_search +28% (491→628) with run_command flat — more searching, less committing.
Item 2: precise errors correct fast
apply_patch attempts 32→21, errors 19→21, successes 13→0 — but every rejection audits clean: the "0 successes" is an infra story, not over-blocking.
All 21 rejections audited clean
9× infra write-fail (§4) · 4× unified-diff (correct) · 3× absolute-path (correct) · 4× unterminated block (correct) · 1× no-pair (correct). edit_file absorbed the load; the model corrects in 1 round on precise hints — zero misdirection observed.
Item 4: factual catalog, no did-you-mean
Unknown-tool hallucinations get a short factual list instead of a guess — the model recovers on the next round. (The full effect lands in run 3: bare-grep hallucinations remain variance, not zero.)
write_file embeds base64 as shell argv
The raise-unification (item 2) exposed a pre-existing backend failure that run 1 counted as success observations.
E2BIG, hidden until now
HarborWorkspace.write_file embeds base64 content as shell argv (printf %s <blob>, execution/harbor.py:129-134); files past the exec arg limit fail with OSError: [Errno 7] Argument list too long: 'docker'. Run 2: 10 trials hit it (3 pass / 7 fail — recovery possible but burns rounds). Run 1 had the identical failure in 12 trials, hidden: old code returned {"error"} as a success observation, so it never appeared in tool.error counts. Change C (return→raise) is what made it visible — an observability win, not a regression.
Flat on the surface — the ground shifted underneath
Principle validated, ceiling argument born
"Generic validation with precise hints: no over-blocking, one-round corrections, zero misdirection. The flat headline is the binding constraints moving downstream — infra writes, gate friction, fix correctness."
「通用校验 + 精确提示」原则成立:零误拦、一轮纠错、零误导。分数持平是因为约束点后移——基础设施写入、gate 摩擦、修复正确性;n=100 的 ±5pp 噪音也吞掉个位数收益。
Ordering the backlog by measured friction
- Infra write fix first — chunked/stdin write in HarborWorkspace.write_file; biggest measured friction (~10 trials/round).
- Item 3 alias table — grep_search -i/output_mode tolerance; already the observed next blocker on the corrected path.
- Item 1 pipe normalization — then the code-layer pipe fix.
- Validation next run — write-fail trials 10→0; apply_patch success back above 0 with rejections still auditing clean; output_mode errors 7→0; pass as primary.