Mount Thor is hiring an engineer to design and improve the data center systems behind our compute fleet. This is an on-site engineering role with direct responsibility for production infrastructure.
The Team
Fleet Operations Engineering designs and operates the physical infrastructure behind Mount Thor’s compute fleet. The team owns data center architecture, hardware integration, commissioning, observability, diagnostics, maintenance, and reliability. It works across racks, power, cooling, networks, operating systems, and Apple hardware to bring capacity online and keep it performing in production. The team builds the instrumentation, tools, standards, and automated workflows needed to find failures, shorten recovery time, and improve each new deployment.
Fleet Operations Engineering partners with Fleet at the machine readiness and health boundary: this team establishes and validates those conditions in the data center, while Fleet represents and acts on them through the software control plane.
In this role you will
Own the technical direction for data center architecture and operations engineering. Define standards for rack design, power, cooling, cabling, network integration, serviceability, safety, and security.
Design and automate the commissioning of new capacity. Build repeatable systems for installation checks, inventory validation, connectivity testing, hardware qualification, and production readiness.
Build observability for the physical fleet. Collect and validate hardware, power, thermal, network, and environmental telemetry. Make that data useful for diagnosis, capacity planning, and automated health decisions.
Lead complex failure investigations across hardware and software boundaries. Work directly with production systems to isolate issues involving hardware, firmware, macOS, networking, storage, power, or cooling.
Turn failures into engineering improvements. Build better diagnostics, test tools, hardware designs, maintenance procedures, and automated workflows.
Design the systems behind maintenance and hardware lifecycle management. Improve preventive maintenance, repair, spare planning, vendor escalation, hardware refresh, and secure decommissioning.
Make agentic development part of daily engineering. Use agents to analyze telemetry, develop tools, investigate failures, and improve documentation. Build structured data and workflows that agents can use safely.
Lead production incidents and planned infrastructure changes. Participate in on-call coverage and coordinate with Fleet, networking, security, data center partners, and hardware vendors.
You might thrive in this role if you have
Designed or operated production data center, HPC, cloud, or large-scale compute infrastructure.
Strong knowledge of data center systems, including racks, power, cooling, structured cabling, networks, and environmental monitoring.
Deep systems knowledge across macOS, Linux, or Unix. You can troubleshoot boot flows, firmware, storage, networking, system performance, and hardware-software interactions.
Experience designing commissioning, qualification, diagnostic, or maintenance systems for physical infrastructure.
Strong programming skills in Python, Go, Rust, Bash, or a similar language.
Experience building telemetry pipelines, dashboards, alerts, or diagnostic tools using metrics, logs, and time-series data.
A record of leading root-cause analysis and turning individual failures into broader reliability improvements.
Strong operational judgment. You plan changes carefully, validate outcomes, document decisions, and design for safe recovery.
Experience using coding agents or other AI tools for engineering, investigation, and data analysis.
Comfort working in active data center environments, joining an on-call rotation, and traveling to other sites when needed.
Bonus Skills
Operated Apple Silicon or macOS infrastructure at data center scale.
Worked with macOS recovery, firmware, secure boot, hardware diagnostics, or automated device restoration.
Designed custom racks or data center environments for dense, nontraditional compute hardware.
Integrated building, power, environmental, or hardware telemetry through industrial protocols and vendor APIs.
Applied hardware reliability methods such as failure analysis, predictive maintenance, component-life tracking, or data-driven spare planning.
Built commissioning or diagnostic systems used across multiple data center sites.
Designed AI-assisted operational systems with clear permissions, audit trails, validation, and human escalation.
Who We Are
Mount Thor makes Apple hardware (macOS and Apple Silicon) available and performant at datacenter scale for AI workloads. We take consumer hardware and build the infrastructure platform around it to enable consumption as elastic compute exposed through developer-friendly interfaces. Our customers use us to develop computer-use model capabilities, deploy long-running agents, and accelerate agentic engineering.
Learn more about this Employer on their Career Site
