SonicJobs Logo
Left arrow iconBack to search

Member of Technical Staff, ML Inference Engineering

Sanas
Posted 8 days ago, valid for 13 days
Location

Palo Alto, CA, US

Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • Sanas is seeking a senior engineer with over 5 years of experience to lead the development of real-time speech and language models deployed in private data centers.
  • The role involves optimizing system and GPU performance for high-throughput AI workloads and improving latency, throughput, and compute efficiency.
  • Candidates should have expertise in NVIDIA GPU architecture, CUDA, and the LLM serving stack, along with a strong background in LLM or related inference systems.
  • The position requires a hands-on approach to core infrastructure decisions and the ability to enhance the capabilities of fellow engineers.
  • Salary details are not specified, but the role demands a high level of expertise in developing reliable, scalable AI applications.

About the Role

Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable.

We're looking for a deeply hands-on, senior engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.

What You'll Do

Performance Optimization

  • Optimize system and GPU performance for high-throughput AI workloads across multi-node training and inference
  • Analyze and improve latency, throughput, memory usage, and compute efficiency
  • Profile system performance to detect and resolve GPU- and kernel-level bottlenecks
  • Implement low-level optimizations using CUDA, Triton, and other performance tooling
  • Improve support for mixed precision, quantization, and model graph optimization
  • Build and maintain performance benchmarking and monitoring infrastructure
  • Scale inference and training systems across multi-GPU, multi-node environments

Inference Systems & Reliability

  • Own and evolve our inference engine, enabling reliability and performance at scale
  • Develop and optimize runtime inference services for large-scale AI applications
  • Implement robust, fault-tolerant systems for data ingestion and processing

Requirements

Must-have:

  • 5+ years of experience writing high-quality, high-performance code
  • Familiarity with NVIDIA GPU architecture and CUDA
  • Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling
  • A research-leaning or systems background in LLM, Speech-to-Text, Text-to-Speech, or Speech-to-Speech inference, with work you can point to
  • A record of shipping research or systems that other people build on, whether in a lab or in industry

Nice-to-have:

  • Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes
  • Experience developing large-scale, high-load production systems
  • Experience maintaining or contributing to open-source ML projects
  • Experience managing machine learning workloads on Kubernetes clusters
  • Experience with InfiniBand or RoCE networking
  • Experience with bare-metal provisioning and lifecycle management
  • Experience operating large-scale AI training or inference clusters
  • Experience with hardware health monitoring and predictive failure detection
  • Experience with distributed storage systems



Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.