SonicJobs Logo
Left arrow iconBack to search

Data Engineer — ML Training Data Pipeline

DATAECONOMY
Posted 12 days ago, valid for 11 days
Salary

Competitive

Contract type

Full Time

Health Insurance
Life Insurance

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • The job title is Data Engineer - ML Training Data Pipeline, with a requirement of 5+ years of experience in data engineering focused on ML data pipelines.
  • The position is located in Hyderabad or Pune and involves building and maintaining data pipelines for transforming raw production traces into high-quality training datasets for LLM fine-tuning.
  • Candidates should have strong skills in Python, experience with AWS services, and familiarity with ML data libraries and formats.
  • The salary for this position is competitive, and the company offers comprehensive medical coverage, robust protection plans, and retirement benefits.
  • Flexible work options, a generous leave policy, and dedicated employee well-being spaces are also part of the benefits package.
Job Title:DataEngineer - ML Training Data Pipeline
Notice period: 0-30 Days
Experience : 5+ Years
Location: Hyderabad OR Pune 

We are looking for DataEngineer - ML Training Data Pipeline who can Build and maintain the data pipeline that transforms rawproduction traces into high-quality training datasets for LLM fine-tuning-ingestion, deduplication, format conversion, quality filtering, and train/testsplitting at scale on AWS.


What We Expect:

  • Build end-to-end data pipelines: raw trace ingestion → dedup → format conversion → quality gating → training-ready datasets
  • Process large-scale JSONL data on AWS S3 (tens of thousands of traces per batch)
  • Convert between chat-completion formats (e.g., OpenAI → Llama 3.1 tool-calling format)
  • Implement smart deduplication and sampling to balance training distribution
  • Design identity-aware train/test splits that measure true generalization
  • Build data validation gates to detect schema drift and format anomalies
  • Create a continuous pipeline that auto-processes new production traces for retraining


Requirements

  • Experience: 6+ years data engineering focused on ML data pipelines
  • Python: Strong — pandas, pyarrow, JSONL processing at scale
  • ML Data Libraries: HuggingFace Datasets, Arrow-based storage
  • Data Formats: Multi-turn conversation/chat data structures and tokenizer-specific formatting
  • Deduplication: Content hashing, identity-based grouping strategies
  • AWS: S3, EC2, batch processing workflows

Preferred (Not Required): LLM training data prep(chat templates, tool-calling schemas); Axolotl or similar dataset formats;data versioning (DVC, LakeFS); browser-automation trace data or Playwright.



Benefits

  • Comprehensive Medical Coverage:
    Health insurance of INR 5.0 Lakhs for you and your family (up to 6 members), ensuring complete peace of mind.
  • Robust Protection Plans:
    Group Personal Accident Insurance and Group Term Life Insurance to safeguard you and your loved ones.
  • Retirement Benefits:
    PF and Gratuity provided as per standard government regulations.
  • Flexible Work Options:
    Enjoy hybrid work arrangements & flexible working hours.
  • Generous Leave Policy:
    21 days of annual leave, in addition to 10 company-declared holidays.
  • Employee Well-being Spaces:
    Access to a dedicated break-out area with round-the-clock refreshments for relaxation and rejuvenation.





Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.