Description
* Design and iterate on end-to-end post-training strategies (including Reinforcement Learning) to unlock model capacities toward achieving specific model behaviors. * Pioneer novel algorithms for preference optimization, model steering, and safety. * Drive our data strategy by researching methods for high-quality human and synthetic data generation, automated data filtering, and curriculum learning to improve instruction following and reasoning. * Design robust evaluation methodologies to measure model helpfulness, factuality, and utility, moving beyond static benchmarks to accurately capture real-world performance. * Partner closely with pre-training teams to inform architecture choices, and with product teams to translate user requirements into model capabilities.
Minimum Qualifications
Demonstrated expertise in deep learning with a focus on LLMs, post-training, or reinforcement learning, backed by a strong record of academic or real-world accomplishments in these or closely related domains. Proficient programming skills in Python and a major deep learning framework such as JAX or PyTorch. Masters/PhD, or equivalent practical experience, in Computer Science, Machine Learning, or a related technical field.
Preferred Qualifications
Experience training state-of-the-art large models at scale, with familiarity in distributed training challenges and trade-offs. Experience improving model performance on complex reasoning tasks (math, coding, logic). Experience with various transformers architectures and its transformations. Strong communication skills and a passion for working cross-functionally across Research and Product teams.
Learn more about this Employer on their Career Site
