About the Role
Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable.
We're looking for a deeply hands-on, senior engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.
What You'll Do
Performance Optimization
- Optimize system and GPU performance for high-throughput AI workloads across multi-node training and inference
- Analyze and improve latency, throughput, memory usage, and compute efficiency
- Profile system performance to detect and resolve GPU- and kernel-level bottlenecks
- Implement low-level optimizations using CUDA, Triton, and other performance tooling
- Improve support for mixed precision, quantization, and model graph optimization
- Build and maintain performance benchmarking and monitoring infrastructure
- Scale inference and training systems across multi-GPU, multi-node environments
Inference Systems & Reliability
- Own and evolve our inference engine, enabling reliability and performance at scale
- Develop and optimize runtime inference services for large-scale AI applications
- Implement robust, fault-tolerant systems for data ingestion and processing
Requirements
Must-have:
- 5+ years of experience writing high-quality, high-performance code
- Familiarity with NVIDIA GPU architecture and CUDA
- Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling
- A research-leaning or systems background in LLM, Speech-to-Text, Text-to-Speech, or Speech-to-Speech inference, with work you can point to
- A record of shipping research or systems that other people build on, whether in a lab or in industry
Nice-to-have:
- Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes
- Experience developing large-scale, high-load production systems
- Experience maintaining or contributing to open-source ML projects
- Experience managing machine learning workloads on Kubernetes clusters
- Experience with InfiniBand or RoCE networking
- Experience with bare-metal provisioning and lifecycle management
- Experience operating large-scale AI training or inference clusters
- Experience with hardware health monitoring and predictive failure detection
- Experience with distributed storage systems
Learn more about this Employer on their Career Site
