Responsibilities
- Curate and integrate publicly available and internal benchmarks to direct the capabilities of frontier model development
- Develop and implement evaluation environments, including environments for novel model capabilities and modalities
- Collaborate with external data vendors to source and prepare high-quality evaluation datasets
- Execute on the technical vision of research scientists designing new benchmarks and evaluations
- Build robust, reusable evaluation pipelines that scale across multiple model lines and product areas
- Contribute to evaluation tooling that measures the quality and reliability of evaluation suites
- Mentor and support other engineers on the team by providing technical guidance and feedback, and helping raise the quality and velocity of evaluation development
Minimum Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- 5+ years of industrial experience in machine learning engineering, machine learning research, or a related technical role
- Proficiency in Python and experience with ML frameworks such as PyTorch
- Experience identifying, designing and completing medium to large technical features independently, without guidance
- Demonstrated software engineering practices including version control, testing, and code review practices
- Ability to work independently and adapt to rapidly changing priorities
Preferred Qualifications
- Publications at peer-reviewed venues (NeurIPS, ICML, ICLR, ACL, EMNLP, or similar) related to language model evaluation, benchmarking, or deep learning
- Hands-on experience with language model post-training and deep learning systems, or building reinforcement learning environments
- Experience implementing or developing evaluation benchmarks for large language models and multimodal models (e.g., vision-language, audio, video)
- Experience working with large-scale distributed systems and data pipelines
- Familiarity with language model evaluation frameworks and metrics
- Track record of open-source contributions to ML evaluation tools or benchmarks
$219,000/year to $301,000/year + bonus + equity + benefits
Learn more about this Employer on their Career Site
