SonicJobs Logo
Left arrow iconBack to search

Model Bring-up Engineer / ML Compiler Engineer

General Compute
Posted 7 days ago, valid for 23 days
Location

San Francisco, San Francisco, CA

Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • The company is seeking a Senior IC Engineer with 5+ years of experience in systems or ML systems, particularly in ML compilers, model porting, or high-performance kernels.
  • The role involves taking new model architectures from reference weights to first correct tokens running on ASICs in a matter of days.
  • Candidates should be comfortable with compiler stacks like MLIR/LLVM and have a strong understanding of modern LLM inference internals.
  • The position emphasizes correctness before optimization, requiring the development of verification harnesses and performance improvements after the initial bring-up.
  • Salary details are not explicitly mentioned, but the role is positioned within a high-stakes startup environment that recently secured significant funding.

About Us

We are the first AI inference neocloud, using ASIC compute to generate tokens 5–7Ɨ faster than existing GPU-based competitors. We just closed an oversubscribed seed round and quadrupled our compute allocation to $97M. Multiple lender conversations are live on a $200M asset-backed equipment facility.

Ā 

About the role

You'll take a new model and get it running — correctly — on our ASIC in record time. When a frontier model drops, the only question that matters is how fast we can land it on our silicon and start serving it. You own that loop: from reference weights, through the compiler, to first correct tokens. The low-level runtime is co-owned with our hardware partner today; your job is everything it takes to get a brand-new architecture compiled, verified, and fast on top of it.

The bet of this role is that bring-up should be an agentic loop, not a hand-port. You'll build the harness of agents that compiles, runs, diffs against reference, and localizes failures — so the marginal model comes up faster than the last one did. Correctness first, optimization second: get it right, prove it's right, then make it cheap. This is a senior IC role on a small team. You'll own the bring-up pipeline, not tickets.

What You'll Do:

  • Own model bringup end-to-end. Take a new architecture — a frontier LLM, an MoE, a multimodal model — from reference weights to first correct tokens running on our ASIC, in days, not quarters.

  • Build the agentic bringup loop. The differentiator isn't hand-porting one model — it's the harness of agents that compiles, runs, diffs against reference, localizes the failing op, and iterates without you in the inner loop. Each model you land should make the loop better at landing the next one.

  • Live in the compiler. Graph capture, IR lowering, op coverage, kernel selection — when a model won't compile or produces wrong numbers, the fix is yours, whether it's a missing lowering, a fused-kernel bug, or a numerics mismatch.

  • Own correctness before speed. Build the verification harness — layer-by-layer activation diffs, logit parity, end-to-end evals — that proves a freshly brought-up model matches reference before anyone trusts a token of it.

  • Then optimize. Once it's correct, make it fast: operator fusion, quantization, memory layout, batching and KV-cache behavior on our hardware. Bringup gets it running; this is where it earns its cost-per-token.

  • Work shoulder-to-shoulder with our hardware partner's compiler and runtime team. You're the person who turns 'the chip can technically run this' into 'this model is live and correct in production.

What we need from you:

  • 5+ years in systems or ML systems, with real depth in at least one of: ML compilers, model porting/bringup, or high-performance kernels.

  • You've taken a model architecture you didn't design and made it run — and run correctly — on a target it wasn't written for. Numerics debugging doesn't scare you.

  • Strong on the internals of modern LLM inference: transformers, attention, KV cache, MoE routing, quantization, batching. You can read a new model's reference implementation and know what will be hard to lower.

  • Comfortable inside a compiler stack — MLIR/LLVM, XLA, or a vendor graph compiler — at the level of IR, lowering, and op coverage, not just calling into one.

  • Fluent with agentic tooling. You'd rather build the agent that runs the tedious bringup loop than run it by hand — and you have the taste to know where the loop still needs a human.

  • Self-directed. We don't assign tickets — you'll see the next model coming and have it half brought-up before anyone asks.

Nice-to-haves:

  • Have worked on a non-NVIDIA accelerator — TPU, Trainium/Inferentia, Tenstorrent, Groq, Cerebras, or similar — at the compiler or model-bringup layer.

  • Kernel-level experience in CUDA, Triton, or a vendor kernel language. You know why a fused attention kernel beats three unfused ops.

  • Have built eval and numerics-verification harnesses (logit parity, activation diffing) for models in production.

  • Contributed to a graph compiler or serving runtime — XLA, TVM, MLIR, vLLM, TGI, TensorRT-LLM, or SGLang.

  • Have built agent loops or LLM-driven tooling that did real engineering work, not demos.




Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.