REQUEST SAMPLES

See graded work before you commit to volume

Pick the collections you want to see. We'll send samples in the format you'd run them in, with the grading attached.

what you can ask for
01
SWE-bench style
Task sets and full traces for reasoning, bug fixing, and code repair.
02
Terminal and agentic
Coding environments that capture tool use, shell interaction, and real execution feedback.
03
TAU-bench
Multi-domain evaluation for long-context reasoning and adaptive problem solving.
04
Cross-model evals
Standardized scoring to compare reasoning across models, releases, and fine-tuning runs.
05
Front-end and UI
UI generation tasks that test whether a model turns a prompt or a spec into a working interface.
06
Custom
Task collections and evaluation suites co-designed with your research team.
request samples