SonicJobs Logo
Left arrow iconBack to search

Senior Reliability Engineer

NVIDIA
Posted 3 months ago, valid for 13 days
Location

Santa Clara, CA 95052, US

Salary

$116,000 - $184,000 per year

Contract type

Full Time

By applying, a NVIDIA account will be created for you. NVIDIA's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • NVIDIA is seeking a Senior HTOL Reliability Engineer for its Santa Clara lab, requiring 5+ years of experience in HTOL test system operation and data analysis for semiconductor devices.
  • The role involves building next-generation HTOL boards, optimizing test programs, and ensuring the reliability of silicon for the AI era.
  • Candidates should possess a Master's or Bachelor's degree in Electrical Engineering or a related field, along with expertise in HTOL stress testing and environmental stress tests.
  • The base salary for this position ranges from 116,000 USD to 184,000 USD, depending on location and experience, with additional eligibility for equity and benefits.
  • NVIDIA is committed to diversity and inclusion, encouraging applications from all qualified candidates until at least June 14, 2026.

NVIDIA is the world leader in accelerated computing, developing breakthroughs that tackle challenges no one else can solve. Our work in AI and digital twins is transforming the world's largest industries and profoundly impacting society. Come join the team and help build the next era of computing!

We're seeking an outstanding Senior HTOL Reliability Engineer to join our Santa Clara lab. This role requires deep device-circuitry knowledge and hands-on hardware development. You will build next-generation HTOL boards and run HTOL processes on advanced ovens. This ensures world-class reliability of the silicon powering the AI era.

What you'll be doing:

  • Implement and optimize HTOL test programs aligned with JEDEC standards.

  • Operate and maintain HTOL ovens, ensuring efficient test conditions and high data accuracy.

  • Build and debug burn-in boards, resolving signal-integrity issues and optimizing thermal performance.

  • Apply sophisticated thermal management techniques to deliver detailed temperature control and mitigate thermal stress in HTOL environments.

  • Work alongside lab technicians, build engineers, and reliability engineers to solve technical challenges and continuously improve test processes.

  • Contribute to multi-functional teams to debug and resolve hardware and software product issues.

  • Maintain and improve our reliability database, finding opportunities for improvement.

  • Collaborate with vendors to develop and implement improvements to burn-in boards, HTOL systems, and thermal interface materials.

What we need to see:

  • Master's or Bachelor's degree in Electrical Engineering or a related field (or equivalent experience).

  • 5+ years of experience in HTOL test system operation and data analysis for semiconductor devices.

  • Proven expertise in HTOL stress testing, JEDEC standards, and environmental stress tests including Temperature Cycling (TC), Reflow, Thermal Shock, and HAST.

  • Hands-on experience with MCC HTOL chamber operation, repairs, and preventative maintenance.

  • Proficiency with oscilloscopes, current probes, and other test equipment for data acquisition and analysis.

  • Skill in vector debugging, test-script development/modification, and data-analysis tools. ATE experience is a plus.

  • Programming experience with Python or MATLAB for data analysis and automation.

  • Excellent communication, teamwork, and problem-solving skills, with strong attention to detail.

Ways to stand out from the crowd:

  • Experience with dual-die or multi-die configurations and the associated thermal challenges.

  • Background crafting burn-in boards for high-power GPU or SoC devices.

  • Familiarity with reliability analytics platforms (JMP) and statistical lifetime modeling (e.g., Weibull, Arrhenius).

  • Track record driving vendor qualification and component selection for reliability test hardware.

  • Exposure to AI/ML-based approaches for reliability data analysis or predictive failure modeling.

With highly competitive salaries and a comprehensive benefits package, NVIDIA is widely considered to be one of the technology world’s most desirable employers. We have some of the most forward-thinking and hardworking people on the planet working for us. If you're creative and autonomous, with a genuine passion for technology, we want to hear from you.

Your base salary will be determined based on your location, experience, and the pay of employees in similar positions. The base salary range is 116,000 USD - 184,000 USD.

You will also be eligible for equity and benefits.

Applications for this job will be accepted at least until June 14, 2026.

This posting is for an existing vacancy. 

NVIDIA uses AI tools in its recruiting processes.

NVIDIA is committed to fostering an inclusive work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.




Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a NVIDIA account will be created for you. NVIDIA's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.