Responsibilities
- Identify and root cause systemic issues across the server fleet and drive resolutions to maximize uptime and utilization by leveraging hardware failure data and diagnostic telemetry
- Write, review, and maintain code for diagnostic and automation tooling that supports quality and efficient delivery of production servers at hyperscale
- Own and develop diagnostic tooling requirements that enable frontline operations teams to efficiently manage and repair the server fleet
- Drive the escalation process for Data Center Operations to identify, root cause, and resolve complex tooling and hardware issues affecting fleet health
- Execute operational validation and verification activities for new product integration into the production environment
- Collaborate with cross-functional tooling teams to provide an operations-centric perspective on open issues and contribute to their development roadmaps
- Perform deep data analysis to prioritize automation opportunities for server repair workflows in a large-scale, heterogeneous hardware environment
- Build cross-functional relationships and influence policies and procedures to improve global data center operations consistency and efficiency
- Mentor other engineers on evaluating and resolving fleet issues and defining improvements to tools and operational processes
- Travel up to 25% to support global data center operations and new site deployments
Minimum Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- 6+ years of experience in production systems engineering, infrastructure engineering, or systems software development for large-scale hardware environments
- 6+ years of experience with hardware lifecycle management, fleet automation, or data center operations systems spanning compute, storage, or networking infrastructure
- Experience developing systems software or automation tooling in Python, Bash, PHP, C, or C++ for Linux-based production environments at scale
- Experience with configuration and maintenance of production systems including web servers, load balancers, relational databases, storage systems, and messaging systems
- Experience communicating technical designs and infrastructure decisions through written documentation and cross-functional stakeholder alignment across engineering and operations teams
Preferred Qualifications
- Experience designing or operating configuration management and infrastructure-as-code systems for large heterogeneous hardware fleets
- Experience supporting global, multi-site data center infrastructure deployments including hardware qualification and regional rollout coordination
- Experience with data analysis and visualization tools used to prioritize fleet health initiatives and drive operational decision-making
- Familiarity with distributed systems monitoring, alerting, and automated remediation pipelines at hyperscale
$144,000/year to $204,000/year + bonus + equity + benefits
Learn more about this Employer on their Career Site
