SonicJobs Logo
Left arrow iconBack to search

Member of Technical Staff, Fleet

Mount Thor
Posted 6 days ago, valid for 20 days
Location

San Francisco, San Francisco, CA

Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • Mount Thor is seeking a software engineer to develop systems that manage a large compute fleet efficiently.
  • The role requires significant experience in building and operating large production systems, particularly with strong software engineering skills in languages such as Go, Python, or Rust.
  • Candidates should have a background in distributed systems and control planes, with familiarity in Linux, macOS, or Unix environments.
  • The position offers a competitive salary of $160,000 per year and requires at least 5 years of relevant experience.
  • The engineer will lead the technical strategy for the fleet control plane, ensuring reliability and efficiency as capacity scales.

Mount Thor is hiring a software engineer to build the systems that manage our compute fleet at scale.

The Team

The Fleet team owns the software layer that turns raw compute capacity into a coherent fleet. The team defines how machines enter, operate within, recover, and leave the fleet. It connects physical capacity to workload demand. It establishes the systems and policies that keep the fleet reliable, secure, and efficient as it grows. Fleet determines how quickly new capacity becomes usable and how reliably workloads run. Our goal is to maximize healthy, schedulable capacity and minimize capacity stranded by provisioning failures, hardware faults, or incomplete recovery.

Fleet is also a proving ground for agentic engineering at Mount Thor. We build systems that software agents can inspect and operate safely. Agents should handle routine investigation, change, and recovery. Humans set policy, manage risk, and resolve novel failures.

The team works across hardware, operating systems, networking, scheduling, security, and data center operations.

In this role you will

  • Own the technical strategy and roadmap for the fleet control plane and full machine lifecycle. Lead complex work across teams and systems.

  • Build the distributed control plane and node-level software that manage inventory, configuration, health, and lifecycle state. Make every action safe, observable, auditable, and recoverable.

  • Automate capacity ingestion across Apple hardware generations. This includes provisioning, validation, configuration, updates, reimaging, diagnostics, repair, and return to service.

  • Connect fleet health and capacity to workload scheduling. Improve availability, placement, utilization, recovery time, and the speed at which new capacity reaches production.

  • Make agentic development and operation core to the team. Use coding agents throughout investigation, implementation, testing, and operations. Build interfaces that let software agents inspect state, take safe action, verify results, and escalate exceptions.

  • Establish strong operational practices. Define health signals and service objectives, lead incidents, improve on-call health, and turn failures into lasting system improvements.

You might thrive in this role if you have

  • Built and operated large production systems that other teams depend on.

  • Strong software engineering skills in Go, Python, Rust, or a similar language.

  • Experience with distributed systems, control planes, state machines, controllers, or durable workflows.

  • Strong knowledge of Linux, macOS, or Unix systems. You are comfortable with boot flows, processes, networking, storage, containers, and system performance.

  • Experience with bare-metal compute, machine provisioning, Kubernetes, workload schedulers, or large server fleets.

  • Experience connecting node-level software to distributed control planes or automated operators.

  • Deep experience using coding agents to build production software. You know how to provide the context, tools, tests, and constraints required for reliable results.

  • A track record of leading complex, multi-team infrastructure work from strategy through production.

  • Strong operational judgment. You design for partial failure, safe retries, auditability, and recovery.

Bonus Skills

  • Built fleet-management systems for thousands of machines across multiple sites.

  • Operated Apple Silicon or macOS infrastructure at scale.

  • Built host agents, health daemons, provisioning pipelines, or automated repair systems.

  • Managed scheduling and capacity across several hardware generations.

  • Built infrastructure designed to be operated by software agents, including permissions, validation, rollback, and human escalation.

  • Delivered measurable improvements in availability, utilization, provisioning speed, or recovery time.

Who We Are

Mount Thor makes Apple hardware (macOS and Apple Silicon) available and performant at datacenter scale for AI workloads. We take consumer hardware and build the infrastructure platform around it to enable consumption as elastic compute exposed through developer-friendly interfaces. Our customers use us to develop computer-use model capabilities, deploy long-running agents, and accelerate agentic engineering.




Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.