code-v16 deepseek-v4.1 run 2 — 69/100
Same model, same genes: the +3pp drift is a coin flip. The run's real product is discipline.
A same-model control with a small working-tree delta
Lineage and genes identical to the −6 run (which a prior chat analysis had mislabeled "ling-flash" — the 0.66→0.69 delta is same-model). The only deltas: skill text (+Search routing / +de-narrowed verify, skills hash 270b605895dcf217→f14bde0e829de480) and a single-exec repo-context probe — plus the apply_patch write-refusal message already reverted to the old text before this run.
+3pp is noise (McNemar)
Paired on the same 100 instances:
| run2 pass | run2 fail | |
|---|---|---|
| −6 pass | 63 | 3 regressed |
| −6 fail | 6 improved | 28 |
6 improved
django-12209 / 13212 · xarray-6599 · pylint-4970 · sphinx-11510 / 9229
Includes two former zero-edit-fail trials. Skill changes may be helping — but n=9 is compatible with anything: hypothesis generator, not evidence.
3 regressed
pylint-6386 · sympy-17630 / 18199
No shared pattern.
One flat, one directional, one broken
The working-tree deltas under test, −6 → run2:
| Metric | −6 → run2 | Verdict |
|---|---|---|
| shell-grep share | 19.8% → 19.4% | No effect — skill routing text doesn't move interactive search (same lesson as the pipe ban: obeyed in declarations only) |
| git log calls | 2.56 → 2.29 /trial | No effect |
| which/version/pip probes | 165 → 175 | No effect |
| declared −k narrowing | 12/42 → 3/42 submits | Directional, weak evidence (llm.start undercounts submits) |
| tool.error total | 12 → 3 | Healthy (1× ChangeLog refusal, 2× edit-match) |
| pipe-mask rejections | 1 → 6 | Uptick, small N — watch one more run |
| never-run / budget / file-less | 24/11/3 → 17/8/3 | Directional, all inside noise |
| zero-edit · iters · tools | 2→1 · 44.1/40.0→42.1/37.5 · 48.4→47.2 | Flat |
Live mechanism, wrong content — reverted
The probe reported the wrong interpreter
Fired in 100/100 trials, but its pytest section reported the base env: /opt/miniconda3/bin/python: No module named pytest — while agents actually test with /opt/miniconda3/envs/testbed/bin/python. Net effect: a misleading toolchain fact. Consistent with probes not dropping: the model kept probing the real environment itself.
Decision + lesson
Reverted in full (4ee829c — probe code + its tests). Lesson: a one-exec probe is only viable if it resolves the effective toolchain (testbed env), not PATH python. Not re-attempted until that is designed.
30/31 are fix-quality — structural blindness
The held-out F2P wall
30× verification.passed + official 0.0: held-out fail-to-pass tests added by the test patch at grading; agents ran the F2P file green without the new test. Same structural blindness as the −6 run — no tool-call intervention addresses it. 1× process failure (sympy-17630, 1 rejection). Budget-exhausted shrank 3→0 (noise-range, not claimed). The −k-narrowed subset persists inside the 30; the skill de-narrowing line is the only live countermeasure, still unproven.
The new series best — with an honest error bar
Adopted from this run, permanently
- Every future comparison reports the McNemar table — χ²>3.84 before celebrating.
- Do not credit mechanism deltas inside noise — this retracts the run-2-era reading of never-run/budget/-k drops as "improvements".
- Search-bypass stays deprioritized — costs tokens, not score — unless a non-text mechanism (soft runtime hint) is proposed with a powered test.