SonicJobs Logo
Left arrow iconBack to search

Member of Technical Staff, ML Inference Engineering

Sanas
Posted 8 days ago, valid for 13 days
Location

Palo Alto, CA, US

Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

About the Role

Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable.

We're looking for a deeply hands-on, senior engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.

What You'll Do

Performance Optimization

  • Optimize system and GPU performance for high-throughput AI workloads across multi-node training and inference
  • Analyze and improve latency, throughput, memory usage, and compute efficiency
  • Profile system performance to detect and resolve GPU- and kernel-level bottlenecks
  • Implement low-level optimizations using CUDA, Triton, and other performance tooling
  • Improve support for mixed precision, quantization, and model graph optimization
  • Build and maintain performance benchmarking and monitoring infrastructure
  • Scale inference and training systems across multi-GPU, multi-node environments

Inference Systems & Reliability

  • Own and evolve our inference engine, enabling reliability and performance at scale
  • Develop and optimize runtime inference services for large-scale AI applications
  • Implement robust, fault-tolerant systems for data ingestion and processing

Requirements

Must-have:

  • 5+ years of experience writing high-quality, high-performance code
  • Familiarity with NVIDIA GPU architecture and CUDA
  • Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling
  • A research-leaning or systems background in LLM, Speech-to-Text, Text-to-Speech, or Speech-to-Speech inference, with work you can point to
  • A record of shipping research or systems that other people build on, whether in a lab or in industry

Nice-to-have:

  • Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes
  • Experience developing large-scale, high-load production systems
  • Experience maintaining or contributing to open-source ML projects
  • Experience managing machine learning workloads on Kubernetes clusters
  • Experience with InfiniBand or RoCE networking
  • Experience with bare-metal provisioning and lifecycle management
  • Experience operating large-scale AI training or inference clusters
  • Experience with hardware health monitoring and predictive failure detection
  • Experience with distributed storage systems



Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.