Train and evaluate everything frontier AI does through code

Apps, games, interfaces, infrastructure, and work beyond software in science, ML, ops, and security. If models produce it through code, it's in our domain.

Applications
Full problem and solution traces from real repositories, beyond single-file patches
repo-leveltestsdebugging
+ retry with queue lock held
− sleep-based retry loop
tests 412 passed · 0 flaky
01 · patch and test run
Interfaces and UX
UI generation tasks, with design quality judged blind by senior designers
generationblind rubricdesign craft
rubric score 4.8 / 5
3 senior designers · blind
02 · graded UI generation
Games
Design quality of AI-generated games, judged by working engineers
gameplayphysicsreplay
frame 0412 · 120 Hz✓ replay match
03 · deterministic gameplay frame
3D
Assets and scenes produced through code and real toolchains
assetsscenestoolchains
mesh watertight · 24k tris
rubric client-defined
04 · asset with client rubric
Infrastructure
Infrastructure-as-code tasks, deployed and verified
live cloudterraformverified
apply plan clean · 0 drift ✓
probe live endpoints healthy ✓
verify rollback path exercised ✓
05 · deploy pipeline, live AWS
Beyond software engineering
Terminal work in science, ML, ops, and security
scienceMLopssecurity
$ nextflow run rnaseq --profile docker
✓ checksums match reference run
06 · verified terminal episode

Revelo Data Lab

Choose one option per dimension. The cube shows where the project sits.

How to work with us

Type of data

Use case

No projects yet

Get this data
EXPERTS

The engineers behind the data

Senior engineers and designers from a network of 400,000+, built over ten years of hiring, testing, and placing them into real jobs. We match people to tasks by specialty.

work queue · matched by specialtylive
Staff infrastructure engineerstaff infrastructure engineerIaCAssigned · env_0412
Senior product designersenior product designerUI rubricGrading · blind review
Senior ML engineerGPU kernel engineerCUDA kernelsProfiling · kernel benchmark
Game engineergame engineergameplayReviewing · replay checks
network · ten years of vetting
400,000+
engineers assessed on real work
years of vetting10
countries18+
companies hired into2,500+
RESEARCH

Latest from Revelo Research

live leaderboard
claude-opus-557.1
gpt-5.6-sol51.2
claude-fable-550.0
Revelo Code Index
A live coding-benchmarks leaderboard built on private, contamination-free collections, with per-model score, steps, and cost analysis.
View the Code Index →
whitepaper + dataset
rubric weights
composition
30
typography
25
color
25
interaction
15
imagery
5
UI Quality Bench
Ten frontier models, 35 hard prompts, three senior designers ranking every result blind on design craft alone. 1,050 blinded ratings, Kendall's W 0.91. Dataset released CC-BY on Hugging Face.
Open the dataset →
benchmark contributions
$ tb run --suite core --task wal-recovery
episode 14 · replay to consistent LSN
✓ verifier pass · checksums clean
Terminal-Bench
Benchmark task collections in the Code Index include Terminal-Bench 2.1 and 3.0 sets, with per-model results and the grading method for every run.
See the benchmark →
CATALOG

Off-the-shelf and custom datasets

Benchmark-based dataset generation
data
SWE-bench-stylemulti-turnIaCUI generation
Expert-written tasks and full problem and solution traces, pulled from real repositories. The task families cover SWE-bench-style, multi-turn, infrastructure-as-code, and UI generation.
task · context · response · outcome · metadata
01 · full trace, every row
Model evaluation
evals
execution checksrubric scoringagent testing
Human evaluation on real software tasks. Execution-based checks, rubric scoring, multi-turn reasoning, and agent performance testing.
210 tasks · 6 dimensions · 5 languages
02 · audit case, frontier lab
Pairwise DPO and RLHF datasets
data
preference pairsreward inputsreasoned choices
Preference pairs with reasoning on every choice, reward model inputs, and code-focused preference signals across languages.
A ○ · B ● preferred · reason attached
03 · ranked pair, every sample
Data quality audit
evals
contaminationbias reviewformatting
Independent audits of training data you already hold. We check for contamination, review bias and diversity, and validate formatting.
✓ format · ✓ diversity · ! overlap flagged
04 · audit report, per delivery
DATA

Expert data for frontier AI

Deep code expertise, turned into training signal for reasoning, agency, and creation.

+Expert-demonstrated trajectories with reasoning, patches, test runs, and review notes
+Synthetic data generated at scale, then reviewed and graded by working engineers
+Pairwise preference data ready for DPO and RLHF, with reasoning on every choice
+Multi-turn task datasets from real repositories
+Coverage across the top 30 programming languages
+Hybrid delivery: your platform or ours, your QA process or ours
sample_04117one row of the dataset
REASONING"The flaky test masks a race in the retry queue, not the parser. Locking during drain fixes the cause."
PATCH+ 41 − 9 · src/queue/retry.ts
TEST RUN412 passed · 0 flaky across 50 reruns ✓
REVIEW"Fix addresses cause, adds regression guard, no scope creep." — senior engineer
every row ships this completepreferred over baseline · DPO ready
iac_environment · release v3.2live AWS
CONNECT
repos · tools
BUILD
state · tasks
RUN
agents
VERIFY
live state
IMPROVE
experts
1 provision_vpc()✓
2 apply_terraform_plan()✓
3 probe_live_endpoints()✓
4 rotate_task_role()✗
checks per task
~30
avg episode
44 steps · 32 min
Opus 4.6 · 8 rollouts
21%
ENVIRONMENTS

RL environments built on real work

Agents act, get verified against live state, and improve, on infrastructure that actually runs.

+RL environments and gyms for agent training
+Environment setup and workflow design, on your infra or ours
+Infrastructure-as-code task environments, deployed and verified
+Long-horizon terminal tasks
EVALS

Evals calibrated to the frontier

Public coding benchmarks are saturated. We measure difficulty by running frontier models against the set and reporting what they score.

+Execution-based evaluations on real code and repos
+Rubric scoring by working engineers and designers
+Multi-turn reasoning evaluation
+Agent performance testing: planning, tool use, code edits
+Contamination detection against public benchmarks
+Bias, diversity, and formatting review
calibration_run · terminal-bench 3 coreauto-verified
taskattempt 1 · 2 · 3verdict
wal-recovery-041✗✗✗kept · frontier-hard
eks-blue-green-027✗✓✗kept · frontier-hard
retry-queue-race-112✗✗✗kept · frontier-hard
rename-config-flag-009✓✓✓cut · too easy
fix-off-by-one-114✓✓✗cut · too easy
difficulty measured, not estimated · contamination screenedkept set: opus-5 <10%
MECHANISM

Engineers judge what tests can't

01 · THE TASK
Original, expert-written tasks
Written by vetted senior engineers from real repositories and real infrastructure, with the full expert trace: reasoning, patches, test runs, and review notes.
edge case +
02 · THE GRADES
Execution and rubric grades
Every task is graded by execution against real code and by rubric review from working engineers and designers.
92%rubric weighted
03 · THE EVIDENCE
Quality evidence attached
The quality review travels with the data itself, so what you receive is inspectable, not asserted.
✓audit trail attached
04 · THE SAMPLE
Start with a sample
Request samples and see the graded work before you commit to volume.
sample →
QUESTIONS RESEARCHERS ASK
How is this different from a horizontal expert network?
A horizontal network treats software as one category among fifty and staffs it from a resume. We've spent ten years deciding which engineers are good in one domain, and the pipeline reflects that. Every task gets graded by execution and reviewed by someone who does that work for a living. Plenty of vendors say they vet. Ask how long they've been doing it, and on what evidence.
Is this only software engineering?
No. The domain is anything a model produces through code. Applications, interfaces, games, 3D, and infrastructure, plus terminal work in science, ML, ops, and security.
Are these repackaged public benchmarks?
No. Our engineers write every task from real repositories and real infrastructure, in the formats you already run: SWE-bench-style, multi-turn, infrastructure-as-code, UI generation. The Code Index runs on private collections that stay unpublished, and we screen every delivery for overlap with public benchmarks.
How is task difficulty calibrated?
We run frontier models against a collection before it ships, then report the pass rate with the model, the attempt count, and the grading method attached. If a task turns out broken or ambiguous, solvability review catches it first.
Who are your experts, and how are they vetted?
Senior engineers and designers from a network of 400,000+, built over ten years of hiring and placing them into real jobs. We match people to tasks by specialty. The engineers who grade design quality ship interfaces for a living; the ones who write infrastructure tasks run infrastructure.
Do you provide environments and evals, or just data?
All of it. Datasets, RL environments that run on real systems, and evaluation. The evaluation work covers execution-based checks, rubric scoring, and independent audits of training data you already hold.
What formats do datasets ship in?
SWE-bench-style task sets, Harbor-format terminal tasks, full traces, and DPO-ready preference pairs. If you need a format that isn't on that list, we'll match it.
Can you work inside our stack and QA process?
Yes. Your platform or ours, your QA process or ours, or a mix of both.
Can you build custom benchmarks and evals for our use case?
Yes, and most of our work is custom. We build task collections calibrated to a difficulty you set, execution-graded environments, and evals scored against your rubric.

The data your model needs for code and everything code produces.