SonicJobs Logo
Left arrow iconBack to search

MLOps / Serving Engineer

DATAECONOMY
Posted 12 days ago, valid for 11 days
Salary

Competitive

Contract type

Full Time

Health Insurance
Life Insurance

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • The position is for an MLOps/Serving Engineer with over 5 years of experience, located in Hyderabad or Pune.
  • The role involves designing and operating production serving infrastructure for fine-tuned LLMs on AWS, including deploying models and optimizing inference engines.
  • Candidates should have hands-on experience with technologies like vLLM, TensorRT-LLM, or Triton Inference Server, along with familiarity with various AWS instance families and auto-scaling.
  • The job offers comprehensive medical coverage of INR 5.0 Lakhs for employees and their families, along with other benefits such as retirement plans and flexible work options.
  • The ideal candidate should be able to start with a notice period of 0-30 days and will enjoy a generous leave policy of 21 days annually.
Job Title: MLOps/ Serving Engineer
Experience : 5+ years
Location : Hyderabad OR Pune
Notice Period: 0-30 days
Work mode - Hybrid
 
We are seeking an experienced MLOps/ Serving Engineer who can design and operate the production serving infrastructure forfine-tuned LLMs on AWS — optimised inference engines, shadow-mode and stagedrollout pipelines, monitoring dashboards, and the path from experimental modelto full production traffic.
Key Responsibilities:
  • Deploy fine-tuned LLMs using vLLM, TensorRT-LLM, or Triton with continuous batching on AWS GPU instances
  • Build shadow-mode deployment: run fine-tuned model alongside production, log comparison data without impacting live traffic
  • Execute staged rollout: canary (5%) → gradual ramp (25% → 50% → 100%) with automated rollback on quality degradation
  • Optimize inference for input-heavy workloads (~17K token inputs, ~130 token outputs): prefill throughput, KV-cache, INT8 quantization
  • Build monitoring dashboards: latency, throughput, accuracy metrics, cost per request
  • Design auto-scaling; implement high-availability (2× instances); automated rollback triggers on end-to-end quality metrics

Requirements

  • 5+ years MLOps or ML infrastructure engineering
  • Hands-on with vLLM, TensorRT-LLM, or Triton Inference Server
  • Deep familiarity with g5, p4de, p5 instance families, EC2 auto-scaling
  • Have worked on Deployment patterns like Shadow-mode, canary, A/B traffic routing, automated rollback
  • Experience onto Continuous batching, INT8 quantization, KV-cache management
  • Expertise on Docker, Kubernetes (EKS) for ML workloads
  • Worked on CloudWatch, Prometheus, Grafana



Benefits

  • Comprehensive Medical Coverage:  
    Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.
  • Robust Protection Plans:  
    Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.
  • Retirement Benefits: 
    PF and Gratuity provided as per standard government regulations.
  • Flexible Work Options: 
    Enjoy hybrid work arrangements & flexible working hours
  • Generous Leave Policy:  
    21 days of annual leave, in addition to 10 company-declared holidays.
  • Employee Well-being Spaces: 
    Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.




Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.