Microsoft is a company where passionate innovators come to collaborate, envision what can be and take their careers further. This is a world of more possibilities, more innovation, more openness, and the sky is the limit thinking in a cloud-enabled world.Â
Â
Microsoft’s Azure Data engineering team is leading the transformation of analytics in the world of data with products like databases, data integration, big data analytics, messaging & real-time analytics, and business intelligence. The products our portfolio include Microsoft Fabric, Azure SQL DB, Azure Cosmos DB, Azure PostgreSQL, Azure Data Factory, Azure Synapse Analytics, Azure Service Bus, Azure Event Grid, and Power BI. Our mission is to build the data platform of AI, powering a new class of data-first applications and driving a data culture.Â
​​Within Microsoft Fabric, the Azure Monitor team builds services that enable customers to monitor, detect, troubleshoot, and mitigate issues with their services through an increasingly agentic experience. Azure Monitor includes Log Analytics, Application Insights, Container Insights, Hosted Prometheus, Azure Managed Grafana, and more. Additionally, Azure Monitor is the platform upon which Microsoft Sentinel is built. We have a multi-billion dollar business that is growing rapidly, and we run some of the world’s highest scale observability services both for Microsoft internally and for our external customers, processing over 1.5 Exabytes of logs daily and tracking over 100 billion active metrics.​Â
​​The Observability Site Reliability Team is looking for a Software Engineer II to build and improve software features for Microsoft Monitoring solutions used by external customers and internal Microsoft engineering teams. This role focuses on monitoring, alerting, telemetry, diagnostics, automation, and agentic AI-assisted workflows that improve service reliability and supportability.Â
We do not just value differences or different perspectives. We seek them out and invite them in so we can tap into the collective power of everyone in the company. As a result, our customers are better served.Â
Responsibilities
- ​​Build and maintain software components for monitoring, alerting, dashboards, diagnostics, telemetry collection, and service health validation. Â
- Implement automation for alert enrichment, incident context generation, health checks, and operational reporting.Â
- Use logs, metrics, traces, telemetry, tests, and debugging tools to investigate issues and improve diagnosability.Â
- Participate in code reviews and apply team standards for maintainability, reliability, testability, security, and privacy.Â
- Contribute to AI-assisted workflows for incident classification, evidence collection, root-cause investigation, and remediation recommendations. Â
- Participate in on-call rotations, incident retrospectives, and follow-up work that improves monitoring, tests, documentation, and troubleshooting guides. Â
- Escalate blockers, design issues, reliability risks, and production concerns with clear impact and timeline ​
Other:
- Embody our Culture and Values
Qualifications
Required QualificationsÂ
- Bachelor's Degree in Computer Science or related technical field AND 2+ years technical engineering experience with coding in languages including, but not limited to, C, C++, C#, Java, JavaScript, or Python
- OR equivalent experience.Â
Additional Job Requirements:Â
Ability to meet Microsoft, customer and/or government security screening requirements are required for this role. These requirements include, but are not limited to the following specialized security screenings: Microsoft Cloud Background Check:Â
This position will be required to pass the Microsoft Cloud background check upon hire/transfer and every two years thereafter.Â
Preferred QualificationsÂ
- ​​Software engineering experience building features, services, tools, or automation using languages such as C#, Python, Go, Java, PowerShell, or TypeScript. Â
- Foundational understanding of distributed systems, cloud services, APIs, monitoring, logging, telemetry, and production debugging. Â
- Ability to write maintainable, testable, diagnosable, and secure code with guidance from experienced engineers. Â
- Collaboration and communication skills in engineering and live-site situations. ​Â
- Â Experience or interest in Azure Monitor, Log Analytics, Application Insights, Kusto/KQL, Geneva, IcM, or similar observability and incident-management systems. Â
- Exposure to AI-assisted development, runbook automation, alerting, service health dashboards, or telemetry-driven diagnostics. Â
- Interest in reliability concepts such as SLIs, SLOs, availability, latency, error rates, operational toil reduction, and safe deployment practices. Â
Â
Candidate Signals Â
- Software engineering fundamentals and ability to learn production systems quickly. Â
- Interest in cloud-scale monitoring, telemetry, observability, reliability, and safe automation. Â
- Clear communication and collaboration during investigations and customer-impacting situations. ​Â
#azdatÂ
#azuredataÂ
​​#AzureMonitor #Observability #SRE #Log Analytics #Azure ​Â
Software Engineering IC3 - The typical base pay range for this role across the U.S. is USD $102,100 - $202,200 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $133,800 - $219,200 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
Learn more about this Employer on their Career Site
