Description
As a Sr Engineering Specialist on the Evaluation team, you'll keep our evaluation projects running clean from design through delivery. You'll be responsible for understanding how well our products perform, advocating for the changes that matter, and representing how our customers will experience them. We're a nimble, fast-paced team that runs on a tight partnership between Feature/ Tooling Engineering, Program Management, Operations, and Data Science — and this role sits at the center of it. In this role, you will: Own device configuration for evaluation projects, and coordinate with the Operations team to roll out configurations to all testers. Monitor in-progress projects and flag quality issues early, before they affect downstream results. Partner with Evaluation Data Science to review evaluation design for quality and alignment ahead of launch. Drive triage with the evaluation tooling team when technical issues surface mid-project, and follow through to resolution. Turn recurring problems into durable fixes — checklists, guardrails, and process improvements that hold up across future projects. Success in this role looks like projects that launch on a sound design, run without avoidable surprises, and produce results the team can act on with confidence.
Minimum Qualifications
BA/BS in Management Information Systems, Computer Science, or a related technical field — or equivalent practical experience. 7+ years of relevant experience in roles such as technical program manager, consultant, QA/device tester, or evaluation operations. Hands-on manual testing experience, including designing and executing test passes on physical devices. Highly organized, with a track record of balancing multiple concurrent efforts and meeting firm deadlines. Excellent written and verbal communication skills. This role involves a high level of interaction with engineering teams, management, and other organizations across Apple, and we expect you to build and navigate those relationships autonomously.
Preferred Qualifications
10+ years of related experience, including 3+ years leading quality evaluation projects end-to-end. Experience evaluating LLM-based assistants, voice assistants, or other generative AI features. Proven track record of building and scaling efficient quality processes across cross-functional teams. Experience managing device configurations at scale — MDM, profiles, or build rollouts — in coordination with an operations team. Working data literacy: SQL, dashboards, or light scripting to spot anomalies in evaluation data yourself rather than waiting on a report. Experience partnering with data science on evaluation or experiment design, or supporting multi-locale/international evaluation programs. Experience testing or evaluating AI-powered features, and comfort using AI tools in your own day-to-day work. Strong debugging and troubleshooting skills, with the analytical judgment to assess project quality and prioritize which issues to fix first.
Learn more about this Employer on their Career Site
