SonicJobs Logo
Left arrow iconBack to search

Software Engineer II

Microsoft
Posted 20 hours ago, valid for 11 days
Location

Mountain View, CA, US

Salary

$102,100 - $219,200 per year

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • The AI Infrastructure team at Microsoft is seeking a Software Engineer to design and maintain large-scale GPU infrastructure for AI applications.
  • Candidates must have a Bachelor's Degree in Computer Science or a related field, along with at least 2 years of technical engineering experience.
  • The role involves working with languages such as Go, Rust, Python, C++, and C# on Kubernetes clusters to support AI model workloads.
  • The typical base salary for this position ranges from USD $102,100 to $202,200 per year, with higher ranges for specific locations like San Francisco and New York City.
  • Microsoft values diversity and inclusion, providing equal opportunity for all qualified applicants regardless of various personal characteristics.
Overview

The AI Infrastructure team is responsible for building and operating large-scale, highly reliable, and efficient GPU infrastructure that powers Microsoft’s AI ecosystem. We host the training and inference platforms behind many of Microsoft’s flagship AI offerings, including Microsoft 365 Copilot, GitHub Copilot, Microsoft Copilot, and Azure AI Foundry’s inference and fine-tuning services for both OpenAI and open-source models. Our infrastructure enables AI innovation at hyperscale and supports some of the most demanding workloads across the company.

As a Software Engineer on the AI infrastructure team, you will work on cutting edge infrastructure and tools to support large scale model deployments, pre-training, post-training and fine-tuning on latest generation of NVIDIA and AMD GPUs in Azure and Microsoft partner clouds on some of the world’s largest AI Supercomputers.  

Microsoft’s mission is to empower every person and every organization on the planet to achieve more. As employees we come together with a growth mindset, innovate to empower others and collaborate to realize our shared goals. Each day we build on our values of respect, integrity, and accountability to create a culture of inclusion where everyone can thrive at work and beyond.



Responsibilities

As an engineer on the AI infrastructure team, your responsibilities include:

•    Design, develop, and maintain AI infrastructure services in Go, Rust, Python, C++, and C#, deployed on large-scale Kubernetes clusters to support inference, pre-training, and post-training workloads for state-of-the-art AI models.
•    Collaborate with engineers, researchers, and external partners to troubleshoot issues, improve reliability, and optimize the performance of large-scale AI training and inference systems.
•    Build and enhance distributed systems that deliver high reliability, low latency, operational efficiency, and strong security across Azure and partner cloud environments.
•    Develop automation and tooling to improve GPU capacity utilization, streamline fleet operations, and enable efficient scaling of AI infrastructure.
•    Provide operational support, technical leadership, and vision while contributing to the deployment, monitoring, and continuous improvement of engineering systems and practices.



Qualifications

Required Qualifications:

  • Bachelor's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python OR equivalent experience.

Preferred Qualifications: 

•    2+ years designing, developing, and shipping high quality software.  
•    2+ years of experience with distributed systems and cloud-based infrastructure. 
•    1+ year of experience with DevOps practices (CI/CD, automated testing, deployment, etc.).   
•    2+ years of software development experience in C#, C++, Python, or similar languages.  
•    2+ years of experience with containerization tools (e.g., Docker, Kubernetes).
•    Knowledge and hands on experience with production ML systems, large-scale training infrastructure, NCCL, CUDA libraries and tools

#AIINFRA



Software Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.



Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.




Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.