Responsibilities
- Own the technical vision and roadmap for key areas of MTIA's developer tooling ecosystem, focusing on debugging, workload error analysis, and product / fleet reliability
- Design and develop debugging tools — including graph-mode debugging, kernel-level diagnostics, and multi-rank fault isolation
- Contribute to overall MTIA SW tooling infrastructure — enabling profiling, performance debugging, memory sanitization, and reliability analysis for MTIA training and inference workloads
- Collaborate closely with MTIA compiler, runtime, kernel, and hardware teams to integrate tooling hooks throughout the MTIA software stack
- Drive AI-native tooling approaches by leveraging automation and LLM-guided diagnostics to improve developer productivity and reduce time-to-root-cause
- Partner with internal product teams across advertising, recommendations, and generative AI to understand developer pain points and prioritize tooling investments
- Advise on tooling best practices, debugging methodologies, and systems-level analysis for accelerator software; communicate architectural decisions clearly through design documents and cross-team reviews
Minimum Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- 6+ years of experience in systems software engineering, performance engineering, developer tooling, or a closely related field
- Experience building debugging, profiling, or diagnostic tools for complex software/hardware systems
- Proficiency in C++ and Python, including low-level systems programming and scripting for tool automation
- Experience working across multiple layers of a system stack (compiler, runtime, OS/driver, hardware)
- Experience leading the technical design and delivery of tooling or infrastructure projects from inception through production deployment
- Experience using data-driven methods and experimentation to evaluate and validate tooling effectiveness and systems performance improvements
Preferred Qualifications
- Familiarity with ML framework internals (PyTorch graph execution, torch.compile, operator dispatch) and AI compiler stacks (MLIR, LLVM, TVM, Triton)
- Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements)
- Experience with accelerator ecosystems (GPU/CUDA, TPU, custom ASICs) including performance profiling, memory analysis, and runtime debugging using their toolchains (cuda-gdb, nsight-compute, nsight-systems, cuda-memcheck)
- Demonstrated cross-stack debugging ability, including Linux kernel and driver-level debugging, with capacity to trace issues across application, OS, and hardware boundaries
- Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies
- Experience with distributed systems debugging or profiling (multi-device, multi-node/multi-rank)
- Experience with Linux debugging and profiling infrastructure (gdb, perf, eBPF, ftrace, coredump analysis, hardware performance counters) and familiarity with binary formats and debugging metadata (ELF/DWARF)
- Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews)
- 7+ years of experience in systems software, developer tooling, or accelerator software development (or equivalent with advanced degree)
- Track record of building developer tools adopted by large engineering populations, ideally with contributions to open-source projects (gdb, LLVM sanitizers, Valgrind, Triton, etc.)
$183,997/year to $257,000/year + bonus + equity + benefits
Learn more about this Employer on their Career Site
