Benchmark home/code-v16 deepseek-v4.1 run 2
EXPERIMENT · EVAL100 · CODE-V16 · SAME-MODEL CONTROL

code-v16 deepseek-v4.1 run 2 — 69/100

Same model, same genes: the +3pp drift is a coin flip. The run's real product is discipline.

Harness: code-v16 · genes 99aea98bc27bef4b (identical to the −6 run) Model: deepseek/deepseek-v4.1-flash — same as the −6 run Working-tree delta only: skill routing + de-narrow, single-exec probe 100/100 · wall 2026-09-28 15:23→19:45 UTC (~4.4h) jobs/eval100-code-v16-deepseek-v4.1/
PASS RATE
69/100 · 0.69
+3pp vs the −6 run — McNemar χ²=0.44 (p≈0.5): coin-flip level
DISCORDANT PAIRS
6 ↑ / 3 ↓
significance needs ~net +8 at a 9:1 split (χ²≈4.9) at n=100
PROBE
reverted (4ee829c)
fired 100/100 but reported the base env's python — a misleading toolchain fact
FAILURE STRUCTURE
30/31 fix-quality
verification.passed + official 0.0: held-out F2P the agent never knew about
Verdict: the 0.66→0.69 delta is same-model noise. The run's value is methodological: standing rules that every future comparison reports its McNemar table, and nothing inside noise gets celebrated.
01 · SETUP

A same-model control with a small working-tree delta

Lineage and genes identical to the −6 run (which a prior chat analysis had mislabeled "ling-flash" — the 0.66→0.69 delta is same-model). The only deltas: skill text (+Search routing / +de-narrowed verify, skills hash 270b605895dcf217→f14bde0e829de480) and a single-exec repo-context probe — plus the apply_patch write-refusal message already reverted to the old text before this run.

Harness
code-v16 · genes 99aea98b
Model
deepseek-v4.1-flash
Evaluated
100/100
Counting
tool calls deduped by _id · submits via verification events
02 · HEADLINE

+3pp is noise (McNemar)

Paired on the same 100 instances:

run2 passrun2 fail
−6 pass633 regressed
−6 fail6 improved28

6 improved

django-12209 / 13212 · xarray-6599 · pylint-4970 · sphinx-11510 / 9229

Includes two former zero-edit-fail trials. Skill changes may be helping — but n=9 is compatible with anything: hypothesis generator, not evidence.

3 regressed

pylint-6386 · sympy-17630 / 18199

No shared pattern.

Decision line for n=100: treat any mean drift < ~0.08 as noise; significance needs roughly net +8 with few regressions (e.g. a 9:1 split, χ²≈4.9). Discordant = 9 here, χ² = 0.44 (bar 3.84, p≈0.5); two-sample z ≈ 0.45 — same conclusion.
03 · MECHANISM READOUT

One flat, one directional, one broken

The working-tree deltas under test, −6 → run2:

Metric−6 → run2Verdict
shell-grep share19.8% → 19.4%No effect — skill routing text doesn't move interactive search (same lesson as the pipe ban: obeyed in declarations only)
git log calls2.56 → 2.29 /trialNo effect
which/version/pip probes165 → 175No effect
declared −k narrowing12/42 → 3/42 submitsDirectional, weak evidence (llm.start undercounts submits)
tool.error total12 → 3Healthy (1× ChangeLog refusal, 2× edit-match)
pipe-mask rejections1 → 6Uptick, small N — watch one more run
never-run / budget / file-less24/11/3 → 17/8/3Directional, all inside noise
zero-edit · iters · tools2→1 · 44.1/40.0→42.1/37.5 · 48.4→47.2Flat
04 · PROBE POST-MORTEM

Live mechanism, wrong content — reverted

The probe reported the wrong interpreter

Fired in 100/100 trials, but its pytest section reported the base env: /opt/miniconda3/bin/python: No module named pytest — while agents actually test with /opt/miniconda3/envs/testbed/bin/python. Net effect: a misleading toolchain fact. Consistent with probes not dropping: the model kept probing the real environment itself.

Decision + lesson

Reverted in full (4ee829c — probe code + its tests). Lesson: a one-exec probe is only viable if it resolves the effective toolchain (testbed env), not PATH python. Not re-attempted until that is designed.

Change log of this cycle: d934cb0 (probe + skill routing + de-narrow + write-refusal message) → 5c63d03 (owner: reverted the refusal message) → 4ee829c (reverted probe + tests). Retained: skill search-routing + de-narrowed-verify lines — zero cost, unproven; kept, not credited.
05 · FAILURE STRUCTURE

30/31 are fix-quality — structural blindness

The held-out F2P wall

30× verification.passed + official 0.0: held-out fail-to-pass tests added by the test patch at grading; agents ran the F2P file green without the new test. Same structural blindness as the −6 run — no tool-call intervention addresses it. 1× process failure (sympy-17630, 1 rejection). Budget-exhausted shrank 3→0 (noise-range, not claimed). The −k-narrowed subset persists inside the 30; the skill de-narrowing line is the only live countermeasure, still unproven.

06 · COMPARISON

The new series best — with an honest error bar

01020 304050 6070 code-v7 · 41/100 code-v7-2 · 44/100 · noise code-v8 · 41/100 code-v9 · 43/100 code-v9 (dsk) · 40/100 code-v11 (dsk) · 48/100 code-v12 (dsk) · 50/100 code-v16 ling run3 · 45/99 code-v16 deepseek-v4.1 (run −6) · 66/100 code-v16 deepseek-v4.1 run2 · 69/100 · THIS RUN 41–44 plateau 41 44 41 43 40 48 50 45.5 66 69 v7v7-2v8 v9v9 (dsk)v11 (dsk) v12 (dsk)v16 (ling)v16 (dsk41)r2 (dsk41)
eval100 full runs · this run highlighted. 66→69 same-model: χ²=0.44 — the best honest claim is "the band is 66–69".
07 · STANDING RULES

Adopted from this run, permanently

  1. Every future comparison reports the McNemar table — χ²>3.84 before celebrating.
  2. Do not credit mechanism deltas inside noise — this retracts the run-2-era reading of never-run/budget/-k drops as "improvements".
  3. Search-bypass stays deprioritized — costs tokens, not score — unless a non-text mechanism (soft runtime hint) is proposed with a powered test.
08 · REPRODUCTION

Artifacts & record

Job
jobs/eval100-code-v16-deepseek-v4.1/
Baseline
jobs/eval100-code-v16-ling-flash-6/ (66.0)
Record
eval100-code-v16-deepseek-v4.1__code-v16__dataset-eval100.json
Report
evaluation/analysis/eval100-code-v16-deepseek-v4-1-run2.md
Counting conventions: the record snapshot logs 68 pass / 31 fail / 1 errored (rate 0.6869 over 99 judged); the report counts rewards keyed by task prefix (69×1.0 / 31×0.0) — same convention as the v12 report's note. Data and conclusions as published in the analysis report.
← Previousdeepseek-v4.1 swap · 66/100 — +20.5pp on the same harness Back to leaderboard ↑SunAgent Harness Benchmark home Next →terminal-bench-2 · tb-v1 — new benchmark, 44/89