About us
We are building AI systems that can reason, use tools, and complete meaningful work in the real world. Our team works across model post-training, reinforcement-learning infrastructure, large-scale training, and product engineering. We believe the fastest path to more capable and reliable agents is an integrated loop: challenging environments, rigorous evaluations, efficient training, reliable inference, and products that make those capabilities useful.
About the role
You will build and optimize the systems that train large language models across multi-node accelerator clusters. The work spans training frameworks, parallelism strategies, data pipelines, kernels and compilers, communication, checkpointing, fault tolerance, numerical correctness, and fleet utilization.
The role is ideal for an engineer who enjoys following a performance or stability problem across every layer—from the model graph and optimizer to the runtime, collective communication library, network, and GPU. You will partner closely with model researchers so that new architectures and experiments can scale without sacrificing correctness or iteration speed.
What you'll do
Design, build, and maintain distributed training systems for pre-training and full-parameter model training at increasing model, data, and cluster scale.
Implement and tune combinations of data, tensor, pipeline, context, sequence, and expert parallelism based on model architecture and hardware topology.
Improve accelerator utilization and end-to-end step time through profiling, graph and kernel optimization, communication overlap, memory management, compilation, and input-pipeline improvements.
Build high-throughput, reproducible data pipelines with clear dataset versioning, deterministic sampling, sharding, mixing, validation, and observability.
Protect numerical correctness and training quality across precision formats, optimizer changes, parallel layouts, fused operations, and system upgrades; create loss-parity and convergence tests.
Improve reliability for long-running jobs through robust checkpointing, fast restart, fault detection, elastic recovery, and automated diagnosis of hardware and software failures.
Measure and improve scaling efficiency, model FLOPs utilization, cluster utilization, failure recovery time, and cost per trained token.
Develop tooling for experiment launch, configuration, performance analysis, capacity planning, and safe rollout of training-stack changes.
Evaluate new accelerators, interconnects, libraries, and framework capabilities, then integrate the ones that create durable performance or reliability gains.
You may be a good fit if you have
Experience building or operating large-scale ML training systems, distributed computing platforms, or high-performance systems.
Strong knowledge of a modern deep-learning framework such as PyTorch or JAX and practical understanding of automatic differentiation, optimizers, mixed precision, and distributed execution.
Hands-on experience with GPUs or other accelerators, including the ability to interpret profiles, utilization metrics, memory behavior, and communication traces.
Strong programming and debugging skills in Python and a systems language such as C++ or Rust.
A rigorous approach to performance and correctness: benchmark design, measurement discipline, regression testing, and root-cause analysis.
Ability to collaborate closely with researchers and infrastructure teams in a fast-moving environment where architectures and scale targets change quickly.
Strong pluses
Experience training very large dense or mixture-of-experts language models, or scaling workloads across hundreds or thousands of accelerators.
Familiarity with Megatron-LM, FSDP, torch.distributed, torchtitan, DeepSpeed, XLA, or comparable large-scale training stacks.
Experience with CUDA, Triton, torch.compile, custom kernels, compiler stacks, or performance engineering near the hardware-software boundary.
Knowledge of NCCL/RCCL, RDMA, InfiniBand, NVLink, collective algorithms, network topology, storage systems, or cluster schedulers.
Contributions to open-source ML systems, training frameworks, compilers, kernels, or performance tooling.
How we work
Mission first. We choose work for its impact on the mission and take responsibility for the outcome, not just our assigned tasks.
High agency. We identify what is missing, form a plan, and move without waiting for perfect clarity.
Speed with rigor. We ship, measure, and iterate quickly while protecting correctness, safety, and reliability.
Flexible scope. We cross team and technical boundaries when that is the fastest way to solve the real problem.
Low ego, high standards. We give direct feedback, change our minds when the evidence changes, and help the whole team win.
Continuous learning. The stack changes quickly; we are willing to learn unfamiliar systems, methods, and domains as the work demands.
Location, visa sponsorship & benefits
Location-based hybrid policy. This is a location-based hybrid role. We currently expect all staff to work from one of our offices at least three days per week. Exact office options will be confirmed during the recruiting process.
Visa sponsorship. We do sponsor visas. However, we cannot successfully sponsor a visa for every role and every candidate. If we make you an offer, we will make every reasonable effort to secure the necessary visa, and we retain immigration counsel to support the process.
Meals and office benefits. We provide complimentary lunch and dinner in our offices, along with snacks and beverages.
A note on qualifications. We care more about exceptional evidence than a perfect keyword match. If the work excites you and you can show unusual strength, learning speed, or ownership, we encourage you to apply even if your background does not match every preferred qualification.
Equal opportunity
We are an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, genetic information, or any other characteristic protected by applicable law. We provide reasonable accommodations for candidates who need them during the hiring process.
Learn more about this Employer on their Career Site
