terminal-bench-2, tb-v1: 44/89
Not a zero. Forty-five zeros, fully attributed — a config-and-protocol story, not a model story.
tb-v1: code-v16 generalized, and what got cut
A new harness for a new benchmark — the general-agent variant of code-v16: the system prompt rewritten, budget reminder configured but disabled, verification trimmed to [solution_description, evidence] (command_to_verify cut — empty means pass, rerun skipped), and a new execute_code tool.
The timeout formula (not a flat 15 minutes)
base = config.agent.override_timeout_sec or task.config.agent.timeout_sec # 900/1200/2400s, None = unlimited resolved = min(base, max_sec or inf) * (agent_timeout_multiplier or timeout_multiplier)
This run: ×3.0 — caffe-cifar-10 3600s (1200×3), schemelike 7200s (2400×3), llm-batching 5399s (1800×3); verifier 600×3=1800s. The prompt's coding/non-coding split ("FOR CODING fill in command_to_verify / FOR NON-CODING leave empty") is left entirely to the model.
Where the 45 zeros went
A · LLM stream idle timeout (11)
agent.end "LLM stream produced no chunk for 120 seconds", usually no verification events. Dying in 1–2 steps is not task difficulty — it is a provider/model stream interruption (retries: provider 0, idle 2; durations ~361–557s). Representative: path-tracing (official 5/5 FAIL, FileNotFound /jail/reconstructed.ppm — model hung after a wrong path).
B · outer AgentTimeoutError (10 zero + 1 passed)
Compile/train-style tasks (Caffe source + 500 iters, GPT2 training, count-dataset 39 execute_code calls) exceed base×3; n_concurrent 4 aggravates contention. extract-elf timed out but scored 1 — the work was done and files landed; timeout ≠ no work.
C · budget exhausted, no submit (15 — largest)
llm.end=81, verification.failed missing=[submit_result] — 80 steps burned, reminder gene disabled. Official failures: query-optimize (/app/sol.sql never written, 4/6 checks failed), sam-cell-seg (test_output.csv is a directory), train-fasttext (model.bin missing), mailman (service never started). Tool imbalance: mailman used execute_code 0 times — all piping run_command, violating the prompt's requirement.
D · verification.passed but official 0 (9 — most valuable)
The gate treats empty command_to_verify as pass and skips rerun, so the self-check ≠ official test: mteb-retrieve read the wrong dataset (HumanEval vs MTEB), pytorch-model-cli 5/6 (output format), sanitize-git-repo 2/3 (dangling SHA), mcmc (internal missing=[] but official stan file fails to compile). qemu×2: not an agent error — dead Debian bullseye EOL repos (curl not found).
Coding/non-coding is inferred by nobody
The harness does not know the task classification: require has no command_to_verify, empty = pass (loop.py:535), and the prompt tells the model to choose ("FOR CODING fill in the command / FOR NON-CODING leave empty"). That prompt-only split is the direct cause of Category D false passes. The fix infers the class from code instead — has_source_edit / executed_commands (verification.py:118,140): force rerun for coding, force existence checks for artifact-style outputs (sol.sql / result.txt / model.bin).
0.49 → 0.65 → 0.73+
Target: 44/89 (0.494) → 65–70/89 (0.73–0.78).
- P0-1 · Turn on budget_reminder (tb-v1.yaml): enabled true, fractions [0.25, 0.6, 0.85] — injects "Budget x/80 + Source edit? + write to disk + submit" at steps 20/48/68. Addresses the 15 C trials; expected +5~6. Optionally max_iterations 80→100.
- P0-2 · Recovery settings: provider_error retries 0→3, stream_idle 2→5, empty/truncated 2→3; env AGENT_RUNTIME_STREAM_IDLE_TIMEOUT 120→300. Addresses the 11 A trials; expected +6~8.
- P0-3 · Timeouts & concurrency: --agent-timeout-multiplier 2.0 --timeout-multiplier 4.0 (or override 7200s); n_concurrent 4→2. Addresses the 10 B trials; expected +3~4.
- P0-4 · qemu bullseye repos (2 free points): sed deb.debian.org→archive.debian.org at build, preinstall curl/uvx — qemu-alpine-ssh / qemu-startup test.sh then pass.
- P1-5 · Automatic coding detection (loop.py:_validate_submission + verification.py): is_coding = has_source_edit or test-shaped executed commands → coding must declare a rerun command; force file-level rerun (timeout 600s). Expected +4~5.
- P1-6 · Flush to disk + artifact-existence gate: catch timeout and trace.flush first; sol.sql / test_output.csv / model.bin missing → verification.failed. extract-elf proves landing files earns points.
- P1-7 · Stronger model for hard tasks: deepseek-v4.1-flash struggles on long chains (path-tracing); switch last — highest cost; expected +2~3.
Artifacts & record
Event model
agent.start → llm.start/llm.chunk/llm.end → tool.start/tool.end → …
→ verification.{failed, rerun, passed} → agent.end
# Category markers:
# A: agent.end error='LLM stream produced no chunk for 120 seconds'
# B: result.json exception_info.exception_type=AgentTimeoutError
# C: verification.failed missing=[submit_result] reasons=budget exhausted
# D: verification.passed + verifier reward 0