Responsibilities
- Define and own the technical architecture of critical machine learning systems, including model training pipelines, inference infrastructure, and feature engineering platforms, ensuring reliability and scalability across billions of users
- Identify and solve the most complex ML systems challenges across multiple product areas, including issues that span model quality, training efficiency, serving latency, and data integrity
- Develop and establish extensible ML frameworks, modeling standards, and engineering practices that drive consistency and velocity across multiple engineering organizations
- Lead cross-functional technical strategy for machine learning initiatives, aligning research, infrastructure, and product teams around multi-year roadmaps that balance short-term delivery with long-term architectural health
- Apply AI-native workflows and tooling as a force multiplier to accelerate model development cycles, automate evaluation pipelines, and expand the scope of what engineering teams can deliver
- Define new metrics and data-driven decision-making principles for long-term ML projects, connecting model performance signals to organization-level business outcomes
- Proactively identify systemic reliability, privacy, and integrity risks in ML systems and build robust technical safeguards, partnering with compliance and policy teams to ensure responsible AI deployment
- Mentor engineers across the organization on ML systems design, debugging complex model behavior, and building production-grade AI systems, establishing yourself as a sought-after technical coach and technical leader
- Drive performance improvements across large-scale ML systems by identifying bottlenecks that span training, data loading, model serving, and hardware utilization, and leading cross-org efforts to resolve them
- Influence the broader ML engineering community through technical publications, design frameworks, and cross-industry engagement that advances the field
Minimum Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- 12+ years of experience designing, building, and deploying large-scale machine learning systems in production environments
- Experience architecting end-to-end ML platforms spanning data pipelines, distributed training, model evaluation, and low-latency inference serving
- Experience identifying and resolving complex, cross-system ML failures including issues in model quality, training stability, feature consistency, and serving correctness
- Experience defining technical strategy and gaining organizational alignment across multiple engineering teams and cross-functional stakeholders
- Experience communicating complex ML system designs and trade-offs in writing to both technical and non-technical audiences, including executive leadership
Preferred Qualifications
- Experience applying ML to multiple product domains such as ranking and recommendation, generative AI, computer vision, or natural language understanding
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
- Track record of industry-recognized contributions to machine learning systems, such as publications, open-source frameworks, or widely adopted architectural patterns
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Experience building AI-native developer tooling or automation that measurably accelerates ML experimentation and production deployment cycles
- Experience with large-scale foundation model training, fine-tuning, or inference optimization across distributed hardware clusters
$347,000/year to $403,000/year + bonus + equity + benefits
Learn more about this Employer on their Career Site
