About us
We are building AI systems that can reason, use tools, and complete meaningful work in the real world. Our team works across model post-training, reinforcement-learning infrastructure, large-scale training, and product engineering. We believe the fastest path to more capable and reliable agents is an integrated loop: challenging environments, rigorous evaluations, efficient training, reliable inference, and products that make those capabilities useful.
About the role
You will build the core platform for agentic reinforcement learning: rollout inference, environment orchestration, distributed training, scheduling, data movement, observability, and recovery. Your goal is to make ambitious RL experiments easy to launch, fast to iterate, efficient to scale, and reliable enough to run for days without constant intervention.
This role sits at the intersection of distributed systems, ML infrastructure, inference, and performance engineering. You will work directly with researchers to find the bottlenecks that matter, then redesign the system so that a local optimization becomes durable leverage for every future run.
What you'll do
Architect and implement end-to-end RL pipelines that coordinate asynchronous rollout generation, environment execution, reward computation, training, evaluation, and checkpoint promotion.
Build environment infrastructure for sandboxed and stateful agent workloads, including lifecycle management, isolation, retries, timeouts, replay, checkpoint and restore, and clear distinction between task outcomes and infrastructure failures.
Scale high-throughput rollout inference through batching, scheduling, caching, load balancing, disaggregated execution, and efficient model-weight updates.
Design resource-management and scheduling systems that place heterogeneous RL workloads efficiently across GPU, CPU, memory, network, and storage constraints.
Improve training-inference synchronization, checkpointing, fault recovery, elastic scaling, and long-running job resilience.
Profile the full stack and remove bottlenecks in GPU utilization, kernels, communication, serialization, data transfer, storage, and environment throughput.
Establish correctness and reproducibility through versioned artifacts, data lineage, idempotent operations, invariant checks, and tests for silent failure modes.
Build observability and debugging tools that let researchers answer why a run is slow, unstable, or behaviorally wrong without depending on an infrastructure specialist.
Create simple APIs and abstractions that make correct, efficient system use the default while preserving the flexibility needed for fast-moving research.
You may be a good fit if you have
A strong record building and operating distributed systems, ML infrastructure, high-performance computing platforms, or performance-critical backend systems.
Excellent programming ability in Python plus at least one systems language such as C++, Rust, or Go.
Experience reasoning about concurrency, partial failure, backpressure, scheduling, consistency, retries, and observability in production systems.
Ability to profile a complex workload, identify the limiting resource, and deliver optimizations that hold under realistic scale and failure conditions.
Comfort working across boundaries: research code, model runtimes, infrastructure services, cluster schedulers, and accelerator behavior.
Strong ownership and communication, including the ability to turn loosely defined research pain into maintainable platform capabilities.
Strong pluses
Experience with distributed RL, large-scale pre-training or post-training, online inference, or asynchronous actor-learner architectures.
Familiarity with PyTorch or JAX and systems such as FSDP, Megatron, DeepSpeed, Ray, Kubernetes, vLLM, SGLang, TensorRT-LLM, or similar tools.
Knowledge of GPU architecture, CUDA or Triton, NCCL/RCCL, RDMA, InfiniBand, NVLink, or topology-aware scheduling.
Experience building sandboxed execution, workflow engines, actor systems, durable runtimes, or multi-agent orchestration.
Meaningful contributions to open-source ML systems or infrastructure projects.
How we work
Mission first. We choose work for its impact on the mission and take responsibility for the outcome, not just our assigned tasks.
High agency. We identify what is missing, form a plan, and move without waiting for perfect clarity.
Speed with rigor. We ship, measure, and iterate quickly while protecting correctness, safety, and reliability.
Flexible scope. We cross team and technical boundaries when that is the fastest way to solve the real problem.
Low ego, high standards. We give direct feedback, change our minds when the evidence changes, and help the whole team win.
Continuous learning. The stack changes quickly; we are willing to learn unfamiliar systems, methods, and domains as the work demands.
Location, visa sponsorship & benefits
Location-based hybrid policy. This is a location-based hybrid role. We currently expect all staff to work from one of our offices at least three days per week. Exact office options will be confirmed during the recruiting process.
Visa sponsorship. We do sponsor visas. However, we cannot successfully sponsor a visa for every role and every candidate. If we make you an offer, we will make every reasonable effort to secure the necessary visa, and we retain immigration counsel to support the process.
Meals and office benefits. We provide complimentary lunch and dinner in our offices, along with snacks and beverages.
A note on qualifications. We care more about exceptional evidence than a perfect keyword match. If the work excites you and you can show unusual strength, learning speed, or ownership, we encourage you to apply even if your background does not match every preferred qualification.
Equal opportunity
We are an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, genetic information, or any other characteristic protected by applicable law. We provide reasonable accommodations for candidates who need them during the hiring process.
Learn more about this Employer on their Career Site
