Benchmark sources

Inspect the work behind the scores.

Frozen inputs, grading code, individual results and the limits of each comparison.

Download benchmark source and score audit files (ZIP, 18.9 MiB)

The archive includes a file-by-file checksum manifest. Archive SHA-256:

7f1ee926430269caa3126c46de7335259f25e722fe97e423c764aecc30fafeab

DeepResearch Bench: ten-task local reproduction

Four research systems, ten frozen tasks, four judges. RACE compares report quality with human-written reference articles; FACT checks cited claims against retrieved source text. This is a local reproduction with documented deviations, not the upstream 100-task leaderboard.

The source package includes task and reference data, captured report text, rendered judge prompts, retrieval and grading tools, per-judge results and aggregation artifacts. Read research/METHODOLOGY.md, research/metadata.yaml and the results folders for the exact protocol, upstream revision and denominators.

App-Bench: six tasks, four systems

The archive contains frozen prompts and rubrics, the build harness, grading scripts, source-hash records, blind-label mapping, original grades and harmonized grades. Run python3 verify_scores.py after extraction to recompute the displayed totals from the original item-level grades.

There was one build and one grader per app. Any item marked unverified for one system is excluded for all systems on that task, leaving 136 comparable items out of 151. The fairness report records grading and deployment exceptions, including API-assisted setup for some pharmacy checks. Read appbench/FAIRNESS.md alongside appbench/RESULTS.md.

Source snapshot and full evidence

This download is the compact source and score-audit package. It does not include every raw model log, captured source page, browser screenshot or generated application archive. The full public evidence copies are attached to the benchmark evidence release: DeepResearch evidence (445 MiB, tar.zst) and App-Bench generated source and final grading evidence (90 MiB, tar.gz). Check downloads against SHA256SUMS. Publication copies omit installed environments and runtime state and redact token-like strings in captured evidence; per-file hash ledgers document the changes. The source runners retain original environment assumptions, so this is not a one-command rerun of the hosted products.

Upstream task and reference material retains its original authorship and terms. Provenance is recorded in the included manifests and citation files.

Back to the comparisons