About the Role
We are a small, fast-moving AI training data company building the next generation of verifier-grounded training data for frontier AI models. Our core conviction is that model capability breakthroughs require data breakthroughs — and we're tackling that problem head-on using formal proof systems, simulators, executable tests, and oracle databases to produce data that is correct by construction.
We work closely with frontier model teams across LLMs, robotics, and AI for Science, injecting high-quality, information-dense content — covering physics, biological facts, and self-consistent logic — directly into AI training pipelines. We're hiring both Research Scientists and Research Engineers to join our founding team in San Francisco.
What You'll Do
Design and build automated pipelines for generating verifiable, high-quality AI training data using formal methods, simulators, and oracle databases.
Research and develop novel approaches to producing training data that encodes fundamental natural-world knowledge (physics, biology, logic) in natural language.
Collaborate directly with frontier AI lab partners to understand data needs and translate them into scalable data generation solutions.
Run experiments to evaluate the downstream impact of synthetic and verifier-grounded data on model capability and performance.
Contribute to the technical direction of the company as an early team member with significant scope and ownership.
What We're Looking For
Required:
Based in San Francisco, CA, or willing to relocate — this is a fully on-site role.
Strong background in machine learning, AI research, or related engineering disciplines (0–12 years of experience considered across both tracks).
Experience with one or more of: large language models, formal verification, robotics simulation, scientific computing, or data pipeline engineering.
Comfort working in a very early-stage environment with high ambiguity and a small team.
Rigorous, first-principles thinking about data quality, model training, and evaluation.
Nice to Have:
Familiarity with formal proof systems (e.g., Lean, Coq, Isabelle) or symbolic reasoning tools.
Experience in robotics, physics simulation, or AI for Science domains.
Prior work at or close collaboration with frontier AI labs.
Background in synthetic data generation or data-centric AI research.
Compensation & Benefits
Salary: $100,000 – $300,000 USD annually, depending on experience and role track (Scientist vs. Engineer).
Early-stage equity commensurate with founding-team proximity.
Visa sponsorship: Not available — candidates must be authorized to work in the United States.
Location
On-site, full-time in San Francisco, CA, United States.
Relocation support available for candidates outside the Bay Area.
Learn more about this Employer on their Career Site
