About Build AI
Build AI is the data hyperscaler for Physical AI. We co-design hardware, collection, infrastructure, and research to scale the physical labor dataset orders of magnitude faster than anyone in the world.
Job Summary
We’re hiring an Evals Lead to benchmark video and world-model capabilities with and without our data. You need to deeply understand the philosophy of evals: what an eval is allowed to claim, contamination, leakage, construct validity, and whether a number actually corresponds to a capability. More practically, you should have pushed consequential evals before — evals that changed what a lab trained, shipped, or collected, not a weekend leaderboard.
This is not Head of Dataset & Quality. That seat owns the dataset objective. You own model capability evals.
Key Responsibilities
Design and run benchmarks for video and world models, with Build data and without it, under the same protocol
Make the comparison honest: held-out tasks, contamination and leakage checks, no cooking the numbers
Build the analytics layer: model performance, failure patterns, and whether scaling the dataset moves capability — and on which axes
Work with research and Head of Dataset & Quality so collection and evals inform each other
Push evals that are consequential enough that people change plans when the number moves
Design processes that increase evaluation quality, repeatability, and scale as we add tasks and countries
You may be a good fit if you have (Must-have qualifications)
You have shipped or driven evals that mattered: they changed training, hiring, product, or data decisions
You understand eval philosophy well enough to argue about validity, not only to plot a curve
Ideally video, robotics, world models, or multimodal, but the bar is consequential evals more than a specific domain
You will not confuse a pretty dashboard with an eval that is allowed to decide things
Strong candidates may also have experience with (Nice-to-have qualifications)
Video, robotics, world models, or multimodal evals
You have designed evals used in a paper, a product launch, or a data decision
Background in construct validity, contamination, or leakage
Exposure to evaluation operations: throughput, failure taxonomy, repeatability
Benefits
Competitive pay
Medical, dental, and vision packages with generous premium coverage
$500 per month credit for waiving medical benefits
Housing subsidy of $2k per month for those living within walking distance of the office
Relocation support for those moving to San Francisco (Financial District) or Shenzhen (Nanshan)
Various wellness benefits covering fitness, mental health, and more
Daily lunch and dinner in our office
Unlimited compute budget subject to ROI justification
Travel
How we're different
Build believes in the Bitter Lesson. By taking a general approach of learning from humans, our addressable market is all physical labor.
We are a fully in-person team in San Francisco (Financial District) and Shenzhen (Nanshan), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.
Build AI is an equal opportunity employer. We review every application. If you do not meet every bullet, still apply. Questions: research@build.ai
Learn more about this Employer on their Career Site
