Responsibilities
- Work hands-on across the full LLM post-training stack
- Build high-quality training data and design and run rigorous, product-relevant evaluations
- Execute, analyze, and iterate on large-scale post-training runs
- Own end-to-end model capability hill-climbing, from identifying gaps through data, training strategy, evaluation, deployment, and product feedback
- Advanced long-horizon agent capabilities, including tool use, full-stack coding, search, planning, and personalization
- Develop realistic harnesses, environments, and evaluations for agentive tasks spanning code understanding, implementation, testing, debugging, and tool use
- Research improved training, evaluation, synthetic-data generation, and data curation strategies
- Translate ambiguous user and product needs into tractable research questions, technical plans, and measurable outcomes
- Lead complex cross-functional projects end-to-end while remaining deeply involved in implementation, experimentation, and analysis
Minimum Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- Bachelor’s or Master’s degree in a relevant technical field, or equivalent practical experience
- 6 years of experience in machine learning engineering, AI research, software engineering, or a related field
- 4 years of providing technical leadership for complex, multi-person projects
- Experience in developing or improving frontier-quality large language models or related foundation models
- Deep, hands-on experience with state-of-the-art LLM post-training, data generation, evaluation, experimentation, and model behavior analysis
- Track record of solving ambiguous, real-world problems and delivering measurable impact within defined timelines
- Ability to work independently, lead across functions, and adapt quickly as evidence and priorities evolve
Preferred Qualifications
- Publications at leading peer-reviewed venues such as NeurIPS, ICML, ICLR, ACL, or EMNLP, or equivalent have demonstrated industry impact in AI
- Experience taking model capabilities from research prototypes to production products
- Experience with large-scale distributed systems and high-throughput data pipelines
- Experience developing agent harnesses, realistic environments such as web browsers or coding sandboxes, and associated evaluations
- Extensive experience with long-horizon agents, agentive coding, tool use, personalization, or search
$219,000/year to $301,000/year + bonus + equity + benefits
Learn more about this Employer on their Career Site
