How is this different from a horizontal expert network?
A horizontal network treats software as one category among fifty and staffs it from a resume. We've spent ten years deciding which engineers are good in one domain, and the pipeline reflects that. Every task gets graded by execution and reviewed by someone who does that work for a living. Plenty of vendors say they vet. Ask how long they've been doing it, and on what evidence.
Is this only software engineering?
No. The domain is anything a model produces through code. Applications, interfaces, games, 3D, and infrastructure, plus terminal work in science, ML, ops, and security.
Are these repackaged public benchmarks?
No. Our engineers write every task from real repositories and real infrastructure, in the formats you already run: SWE-bench-style, multi-turn, infrastructure-as-code, UI generation. The Code Index runs on private collections that stay unpublished, and we screen every delivery for overlap with public benchmarks.
How is task difficulty calibrated?
We run frontier models against a collection before it ships, then report the pass rate with the model, the attempt count, and the grading method attached. If a task turns out broken or ambiguous, solvability review catches it first.
Who are your experts, and how are they vetted?
Senior engineers and designers from a network of 400,000+, built over ten years of hiring and placing them into real jobs. We match people to tasks by specialty. The engineers who grade design quality ship interfaces for a living; the ones who write infrastructure tasks run infrastructure.
Do you provide environments and evals, or just data?
All of it. Datasets, RL environments that run on real systems, and evaluation. The evaluation work covers execution-based checks, rubric scoring, and independent audits of training data you already hold.
What formats do datasets ship in?
SWE-bench-style task sets, Harbor-format terminal tasks, full traces, and DPO-ready preference pairs. If you need a format that isn't on that list, we'll match it.
Can you work inside our stack and QA process?
Yes. Your platform or ours, your QA process or ours, or a mix of both.
Can you build custom benchmarks and evals for our use case?
Yes, and most of our work is custom. We build task collections calibrated to a difficulty you set, execution-graded environments, and evals scored against your rubric.