Benchmark home/code-v16 ling-flash run 2
EXPERIMENT · EVAL100 · CODE-V16 · PATCH VALIDATION

code-v16 run 2: patch validation + factual tool catalog — 43.0%

Generic validation, precise hints — and an infra bug with a 10-trials-per-round price.

Code delta 17d2fd2: patch.py + tool_dispatch.py Harness: code-v16 unchanged Model: Ling-3.0-flash Baseline: run 1 (43/97 · 0.443) 100/100 evaluated · zero infra losses jobs/eval100-code-v16-ling-flash-2/
PASS RATE
43/100 · 0.43
flat vs 44.3 — 6 fixed / 6 broken, net 0; run-to-run noise dominates at n=100
HIDDEN WRITE BUG
~10 trials/round
write_file embeds base64 as shell argv — E2BIG in 12 run-1 trials (hidden), 10 run-2 trials
TOOL ERRORS
15 / 100 trials
grep aliases ×7 (next blocker) · write-fails ×10 already counted separately · hallucinated bare-grep ×5 in run 3 — watch
VALIDATION STYLE
0 over-blocks
all 21 apply_patch rejections audit clean · 1-round corrections observed · zero misdirection
Verdict: the "generic validation, precise hints" principle holds — the flat headline is an infrastructure bug, not the gene. The raise-unification made a pre-existing write failure visible; fixing it is the next P0.
01 · SETUP

Two code-layer items, harness untouched

Committed as 17d2fd2 before the job; the skill is the template-removed version (per run 1's verdict). Item 2 = patch validation, item 4 = factual tool catalog.

17d2fd2:
  patch.py        unified-diff sniff · path legality precheck
                  read/write-failure raise unification (was: return {"error"})
                  create-requires-existing-parent
  tool_dispatch.py unknown-tool -> factual catalog, no did-you-mean
Working-tree delta under test. Harness stays code-v16 (genes unchanged); skill = template-removed version.
Harness
code-v16, unchanged
Model
Ling-3.0-flash
Evaluated
100/100 · zero infra losses
Counting
tool.start authoritative · unique run_commands deduped
02 · HEADLINE

Flat score, symmetric churn

Paired on 98 shared instances with run 1: 48×(0,0), 37×(1,1) — perfect symmetry around zero.

6 newly fixed

seaborn-3069 · requests-1142 · sklearn-12973 / 13124 / 13439 / 14496

Includes three of run 1's seven regressions — churn, not trend.

6 newly broken

xarray-6599 / 6938 · sklearn-11578 / 14053 / 25973 · sympy-12419

Per-repo up

psf 57→71% · sklearn 45→55% · mwaskom 0→50%

Per-repo down / drift signals

pydata 64→45% · sympy 40→30%

no-edit 7→11 · hit80 20→24 · error-status 5→10 · grep_search +28% (491→628) with run_command flat — more searching, less committing.

03 · MECHANISM

Item 2: precise errors correct fast

apply_patch attempts 32→21, errors 19→21, successes 13→0 — but every rejection audits clean: the "0 successes" is an infra story, not over-blocking.

All 21 rejections audited clean

9× infra write-fail (§4) · 4× unified-diff (correct) · 3× absolute-path (correct) · 4× unterminated block (correct) · 1× no-pair (correct). edit_file absorbed the load; the model corrects in 1 round on precise hints — zero misdirection observed.

Item 4: factual catalog, no did-you-mean

Unknown-tool hallucinations get a short factual list instead of a guess — the model recovers on the next round. (The full effect lands in run 3: bare-grep hallucinations remain variance, not zero.)

04 · THE BUG

write_file embeds base64 as shell argv

The raise-unification (item 2) exposed a pre-existing backend failure that run 1 counted as success observations.

E2BIG, hidden until now

HarborWorkspace.write_file embeds base64 content as shell argv (printf %s <blob>, execution/harbor.py:129-134); files past the exec arg limit fail with OSError: [Errno 7] Argument list too long: 'docker'. Run 2: 10 trials hit it (3 pass / 7 fail — recovery possible but burns rounds). Run 1 had the identical failure in 12 trials, hidden: old code returned {"error"} as a success observation, so it never appeared in tool.error counts. Change C (return→raise) is what made it visible — an observability win, not a regression.

Documented casualty: xarray-6938 — a correct patch blocked by write-fail at it42, then 30+ rounds of gate friction to an error end. Fix direction (needs approval): stdin/chunked write in HarborWorkspace.write_file (split base64, append with >>), backend-agnostic, zero task knowledge. Expected: removes a ~10%-of-trials friction source and un-blocks apply_patch success rate (9 of its 21 run-2 errors are this bug, not the model).
05 · COMPARISON

Flat on the surface — the ground shifted underneath

01020 304050 6070 code-v7 · 41/100 code-v7-2 · 44/100 · noise code-v8 · 41/100 code-v9 · 43/100 code-v9 (dsk) · 40/100 code-v11 (dsk) · 48/100 code-v12 (dsk) · 50/100 code-v16 ling run2 · 43/100 · THIS RUN code-v16 deepseek-v4.1 (run −6) · 66/100 code-v16 deepseek-v4.1 run2 · 69/100 41–44 plateau 41 44 41 43 40 48 50 43 66 69 v7v7-2v8 v9v9 (dsk)v11 (dsk) v12 (dsk)v16 (ling)v16 (dsk41)r2 (dsk41)
eval100 full runs · this run highlighted. Same headline as run 1 — the difference is under the hood (a visible bug worth ~10 trials/round).
06 · CONCLUSION

Principle validated, ceiling argument born

"Generic validation with precise hints: no over-blocking, one-round corrections, zero misdirection. The flat headline is the binding constraints moving downstream — infra writes, gate friction, fix correctness."

「通用校验 + 精确提示」原则成立:零误拦、一轮纠错、零误导。分数持平是因为约束点后移——基础设施写入、gate 摩擦、修复正确性;n=100 的 ±5pp 噪音也吞掉个位数收益。

07 · NEXT

Ordering the backlog by measured friction

  1. Infra write fix first — chunked/stdin write in HarborWorkspace.write_file; biggest measured friction (~10 trials/round).
  2. Item 3 alias table — grep_search -i/output_mode tolerance; already the observed next blocker on the corrected path.
  3. Item 1 pipe normalization — then the code-layer pipe fix.
  4. Validation next run — write-fail trials 10→0; apply_patch success back above 0 with rejections still auditing clean; output_mode errors 7→0; pass as primary.
08 · REPRODUCTION

Artifacts

Job
jobs/eval100-code-v16-ling-flash-2/
Baseline
jobs/eval100-code-v16-ling-flash/ (run 1)
Code
commit 17d2fd2
Report
evaluation/analysis/eval100-code-v16-ling-flash-2.md
No evaluation/records snapshot for this run; data as published in the analysis report. Counting basis: tool.start authoritative; unique run_commands deduped by _id.
← Previousrun 1 · early budget nudge — flat, gate sharpened Back to leaderboard ↑SunAgent Harness Benchmark home Next →run 3 · chunked write fix — 45.5%, best v16 Ling run