Benchmark home/terminal-bench-2 · tb-v1
EXPERIMENT · TERMINAL-BENCH-2 · TB-V1 · FIRST RUN

terminal-bench-2, tb-v1: 44/89

Not a zero. Forty-five zeros, fully attributed — a config-and-protocol story, not a model story.

Harness: tb-v1 (parent code-v16) · budget reminder OFF · verify minus command_to_verify Model: deepseek/deepseek-v4.1-flash via OpenRouter 89 trials · 11 AgentTimeoutError · wall 2026-09-29 06:41→17:21 UTC (~10.7h) jobs/terminal-bench-2-tb-v1/
PASS RATE
44/89 · 0.494
44×1.0 / 45×0.0 — correcting the "0 score" misconception
CATEGORY C
15 budget-exhausted
81 steps burned, no submit — the reminder gene exists but is disabled
CATEGORY A
11 stream-idle deaths
"no chunk for 120 seconds" — dying in 1–2 steps, provider interruption
CATEGORY D
9 gate-blindness
verification.passed + official 0 — self-check ≠ official test; the most valuable class
Verdict: the 45 zeros decompose into 41 infra/protocol losses (A+B+C, config-fixable) and 9 gate-blindness losses (D, small-code-fixable). The model is rarely the binding constraint on this suite.
01 · SETUP

tb-v1: code-v16 generalized, and what got cut

A new harness for a new benchmark — the general-agent variant of code-v16: the system prompt rewritten, budget reminder configured but disabled, verification trimmed to [solution_description, evidence] (command_to_verify cut — empty means pass, rerun skipped), and a new execute_code tool.

Tools
run_command · read_file · edit_file · apply_patch · write_file · grep_search · find_files · execute_code · submit_result
Controls
max_iterations 80 · finish_violation_limit 3 · budget_reminder.enabled=false
Recovery
provider_error 0 retries · stream_idle 2 · empty/truncated 2
Run config
timeout_multiplier 3.0 · n_concurrent_trials 4

The timeout formula (not a flat 15 minutes)

base = config.agent.override_timeout_sec or task.config.agent.timeout_sec   # 900/1200/2400s, None = unlimited
resolved = min(base, max_sec or inf) * (agent_timeout_multiplier or timeout_multiplier)

This run: ×3.0 — caffe-cifar-10 3600s (1200×3), schemelike 7200s (2400×3), llm-batching 5399s (1800×3); verifier 600×3=1800s. The prompt's coding/non-coding split ("FOR CODING fill in command_to_verify / FOR NON-CODING leave empty") is left entirely to the model.

02 · ATTRIBUTION

Where the 45 zeros went

passed C · budget A · stream-idle B · timeout D · gate-blind 44 15 11 10 9
89 trials = 44 passed + 45 zero · A=stream idle, B=outer timeout, C=budget exhausted, D=internal pass + official fail. Method: full-script statistics over agent-runtime.jsonl + result.json + verifier outputs.

A · LLM stream idle timeout (11)

agent.end "LLM stream produced no chunk for 120 seconds", usually no verification events. Dying in 1–2 steps is not task difficulty — it is a provider/model stream interruption (retries: provider 0, idle 2; durations ~361–557s). Representative: path-tracing (official 5/5 FAIL, FileNotFound /jail/reconstructed.ppm — model hung after a wrong path).

B · outer AgentTimeoutError (10 zero + 1 passed)

Compile/train-style tasks (Caffe source + 500 iters, GPT2 training, count-dataset 39 execute_code calls) exceed base×3; n_concurrent 4 aggravates contention. extract-elf timed out but scored 1 — the work was done and files landed; timeout ≠ no work.

C · budget exhausted, no submit (15 — largest)

llm.end=81, verification.failed missing=[submit_result] — 80 steps burned, reminder gene disabled. Official failures: query-optimize (/app/sol.sql never written, 4/6 checks failed), sam-cell-seg (test_output.csv is a directory), train-fasttext (model.bin missing), mailman (service never started). Tool imbalance: mailman used execute_code 0 times — all piping run_command, violating the prompt's requirement.

D · verification.passed but official 0 (9 — most valuable)

The gate treats empty command_to_verify as pass and skips rerun, so the self-check ≠ official test: mteb-retrieve read the wrong dataset (HumanEval vs MTEB), pytorch-model-cli 5/6 (output format), sanitize-git-repo 2/3 (dangling SHA), mcmc (internal missing=[] but official stan file fails to compile). qemu×2: not an agent error — dead Debian bullseye EOL repos (curl not found).

Contrast: the PASS group build-cython-ext / fix-ocaml-gc / largest-eigenval / winning-avg-corewars also never submitted — but scored via files landing on disk. Writing artifacts beats declaring results.
03 · ROOT CAUSE

Coding/non-coding is inferred by nobody

The harness does not know the task classification: require has no command_to_verify, empty = pass (loop.py:535), and the prompt tells the model to choose ("FOR CODING fill in the command / FOR NON-CODING leave empty"). That prompt-only split is the direct cause of Category D false passes. The fix infers the class from code instead — has_source_edit / executed_commands (verification.py:118,140): force rerun for coding, force existence checks for artifact-style outputs (sol.sql / result.txt / model.bin).

04 · CHECKLIST (BY ROI)

0.49 → 0.65 → 0.73+

Target: 44/89 (0.494) → 65–70/89 (0.73–0.78).

  1. P0-1 · Turn on budget_reminder (tb-v1.yaml): enabled true, fractions [0.25, 0.6, 0.85] — injects "Budget x/80 + Source edit? + write to disk + submit" at steps 20/48/68. Addresses the 15 C trials; expected +5~6. Optionally max_iterations 80→100.
  2. P0-2 · Recovery settings: provider_error retries 0→3, stream_idle 2→5, empty/truncated 2→3; env AGENT_RUNTIME_STREAM_IDLE_TIMEOUT 120→300. Addresses the 11 A trials; expected +6~8.
  3. P0-3 · Timeouts & concurrency: --agent-timeout-multiplier 2.0 --timeout-multiplier 4.0 (or override 7200s); n_concurrent 4→2. Addresses the 10 B trials; expected +3~4.
  4. P0-4 · qemu bullseye repos (2 free points): sed deb.debian.org→archive.debian.org at build, preinstall curl/uvx — qemu-alpine-ssh / qemu-startup test.sh then pass.
  5. P1-5 · Automatic coding detection (loop.py:_validate_submission + verification.py): is_coding = has_source_edit or test-shaped executed commands → coding must declare a rerun command; force file-level rerun (timeout 600s). Expected +4~5.
  6. P1-6 · Flush to disk + artifact-existence gate: catch timeout and trace.flush first; sol.sql / test_output.csv / model.bin missing → verification.failed. extract-elf proves landing files earns points.
  7. P1-7 · Stronger model for hard tasks: deepseek-v4.1-flash struggles on long chains (path-tracing); switch last — highest cost; expected +2~3.
Execution order: 1 + 2 + 4 first — three lines of yaml plus a mirror source — and immediately see 0.49 → 0.65; then 5 + 3 to push 0.73+.
05 · REPRODUCTION

Artifacts & record

Job
jobs/terminal-bench-2-tb-v1/
Job id
db043075-984d-4865-be0d-8bdb12bf6d3d
Record
terminal-bench-2-tb-v1__tb-v1__terminal-bench-2.json
Report
evaluation/analysis/terminal-bench-2-tb-v1.md

Event model

agent.start → llm.start/llm.chunk/llm.end → tool.start/tool.end → …
         → verification.{failed, rerun, passed} → agent.end
# Category markers:
#   A: agent.end error='LLM stream produced no chunk for 120 seconds'
#   B: result.json exception_info.exception_type=AgentTimeoutError
#   C: verification.failed missing=[submit_result] reasons=budget exhausted
#   D: verification.passed + verifier reward 0
Counting conventions: the record snapshot logs 43 pass / 35 fail / 11 errored (rate 0.5513 over 78 judged); the report counts 44/89 = 0.494, including extract-elf — timed out, files landed, verifier still passed it. Data and conclusions as published in the analysis report.
← Previousdeepseek-v4.1 run 2 · 69/100 — same-model noise, standing rules Back to leaderboard ↑SunAgent Harness Benchmark home Next →tb-v2 · P0 config fixes — planned (0.49 → 0.65)