SonicJobs Logo
Left arrow iconBack to search

LLM Inference Engineer

Near AI
Posted 24 days ago, valid for 12 days
Location

San Francisco, CA, US

Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • The NEAR AI team is seeking an expert in high-performance LLM serving systems and inference optimization for their decentralized machine learning infrastructure.
  • This role involves architecting and maintaining production high-traffic LLM serving systems while optimizing throughput, latency, and costs for open-source LLMs.
  • Candidates should have strong hands-on experience in LLM inference and be proficient in debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.
  • A deep knowledge of GPU architectures and proficiency in tools like PyTorch, Triton, and CUDA are essential for this position.
  • The role is based in San Francisco or can be performed remotely, with a competitive salary range that requires several years of relevant experience.

Locations: San Francisco or Remote

About The Role

The NEAR AI team is building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our mission is to build highly scalable and efficient infrastructure for open-source AI at a global scale.

We are specifically seeking an expert in high-performance LLM serving systems and inference optimization. In this role, you will push the boundaries of how large language models are served.

What You'll Be Doing

  • Architect and maintain production high-traffic LLM serving systems.
  • Optimize throughput, latency, and cost for leading open-source LLMs.

What We're Looking For

  • Strong hands-on experience in LLM inference, with expertise debugging and optimizing major inference engines such as SGLang, vLLM, or TensorRT.
  • Deep knowledge of state-of-the-art GPU architectures, and effectively exploit them using PyTorch, Triton, CuTe, CUDA, etc.
  • Proven track record in designing and maintaining end-to-end high-traffic LLM serving systems.
  • Strong problem-solving skills and ability to communicate technical ideas clearly.

We'd Love If You Have

  • Experience with Trusted Execution Environments (TEE).
  • Active contributor to open-source LLM inference engines.

Please let us know if you require any special requirements for your interview and we'll do our best to accommodate.




Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.