Responsibilities
- Design, implement, and optimize high-performance compute and communication kernels for MTIA accelerators, taking ownership from architectural analysis through production deployment
- Profile and root-cause performance across the full stack — instruction scheduling, memory hierarchy and DMA behavior, on-chip interconnect, multi-device collectives — and drive the fixes to the right layer, whether that is the kernel, the compiler, the runtime, or the hardware
- Build and extend kernel authoring frameworks, templates, and libraries so that other engineers can reach high performance without deep architectural expertise; raise the ceiling and lower the floor at the same time
- Deliver and maintain broad kernel coverage for PyTorch operators across recommendation, ranking, and generative AI workloads, in both eager and compiled execution paths
- Partner with silicon architecture and design teams on hardware/software co-design: quantify the value of proposed features with real kernels, characterize rooflines pre-silicon, and advocate for the changes the software stack actually needs
- Work with compiler, runtime, framework, and product-facing teams to land end-to-end wins on production models rather than isolated microbenchmark improvements
- Investigate numerics and precision trade-offs, and design software mitigations that recover performance or accuracy lost to hardware limitations
- Set technical direction for a kernel domain, write the design documents that align cross-functional partners, and mentor engineers on performance methodology and accelerator programming
Minimum Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- Bachelor's degree in Computer Science, Computer Engineering, a related technical field, or equivalent practical experience
- 6+ years of professional experience in high-performance computing, accelerator kernel development, compiler backends, or systems performance engineering
- Proficiency in C++ and Python, including low-level systems programming, templates and generic programming, and comfort reading and writing performance-critical code
- Demonstrated experience writing and optimizing kernels for a parallel architecture — GPU (CUDA, ROCm/HIP, SYCL/OpenCL), TPU or other AI ASICs, or SIMD/vector CPU targets
- Working knowledge of computer architecture as it applies to performance: memory hierarchies and bandwidth, latency hiding, occupancy and scheduling, vectorization, and synchronization
- A rigorous, measurement-driven approach to performance: the ability to build a roofline or analytical model, profile against it, and explain the residual gap
Preferred Qualifications
- Experience mentoring engineers and setting technical direction across teams
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
- Experience with distributed execution and collective communication (NCCL/RCCL-class primitives, tensor and expert parallelism, overlapping communication with compute)
- Experience with low-precision numerics and quantization — FP8/E4M3/E5M2, MX and other block-scaled formats, INT8/INT4 — including error analysis and calibration
- Experience with compiler and codegen technologies relevant to kernels: MLIR, LLVM, TVM, XLA, Halide, or polyhedral scheduling
- 8+ years of experience in accelerator software, HPC, or ML systems performance (or equivalent with an advanced degree)
- Deep familiarity with transformer and attention kernel design: FlashAttention-class algorithms, KV-cache management, paged and chunked attention, linear and state-space attention variants, MoE routing and expert dispatch
- Track record of open-source contribution in the kernel, compiler, or ML systems ecosystem
- Experience with pre-silicon software development — architectural simulators, FPGA emulation, performance modeling — and with hardware/software co-design cycles
- Familiarity with ML framework internals: PyTorch dispatch and eager execution, torch.compile / Inductor, custom operator integration, and inference serving stacks such as vLLM or SGLang
- Experience building or contributing to high-performance kernel libraries or frameworks — CUTLASS, cuBLAS, cuDNN, CUTE, Triton, Helion, ThunderKittens, oneDNN, Composable Kernel, or comparable internal equivalents
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
$183,997/year to $257,000/year + bonus + equity + benefits
Learn more about this Employer on their Career Site
