Frontier models are great at verifiable tasks like math and coding.
We bring expertise and taste to the real-world domains beyond them.
Models are moving from answering questions to doing work. Exam-style tests are close to saturated, but passing a test is not the same as running an operation.
Frontier capability is now limited mostly by data. There's little training signal for the messy, rule-bound decisions people make in real-world work. That is the gap we work on.
Answering questions
One prompt, one answer, checked against a key. Today's models do this well.
Doing real work
Many steps, hard constraints and a long horizon, where one bad decision early on spreads through the rest of the day.
Where we work.
We focus on long-horizon, rule-bound work where decisions are coupled over time and outcomes can be checked mechanically.
Real-world simulations
Mission-critical operations, simulated at full complexity.
Operations · Physical world
Verifiers and data
Graders that can't be gamed, and the tasks, trajectories and eval sets to train models on real work.
Rewards · Evals · Training data
RL environments
Sandboxed, multi-step environments with tiered tasks, from small debugging cases to cascading crisis scenarios.
Long horizon · tool use