SonicJobs Logo
Left arrow iconBack to search

Machine Learning Engineer (Egocentric 3D Human Pose)

Glint Tech Solutions LLC
Posted 12 hours ago, valid for 11 days
Location

Santa Clara, CA, US

Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • We are seeking a Machine Learning Engineer with at least 3 years of experience to join our research team focused on 3D human body and hand motion recovery from egocentric video.
  • The role involves developing models and production pipelines for accurate 3D pose representations, addressing challenges such as self-occlusion and motion blur.
  • Candidates should have a Bachelor's, Master's, or PhD in a relevant field and strong skills in Python, PyTorch or TensorFlow, along with experience in 3D human pose estimation.
  • The position offers a competitive salary and equity package, along with comprehensive benefits including medical insurance and a 401(k) retirement plan.
  • This is an opportunity to work with leading researchers in robotics and AI, contributing to cutting-edge perception systems for next-generation robotics.

About the Role

We are looking for a Machine Learning Engineer to join our core research and development team focused on recovering accurate 3D human body and hand motion from egocentric (first-person) video.

Human demonstration data is the foundation of robot learning, and its quality depends on accurately reconstructing human motion. In this role, you will develop models and production pipelines that transform head-mounted and body-mounted camera streams—including wide-FOV, stereo, motion-blurred, and heavily self-occluded video—into metrically accurate, temporally consistent 3D pose representations for robot policy training and human-to-robot motion retargeting.

You will work across the entire perception stack, including camera calibration, data annotation, model training, evaluation, and large-scale deployment. This role is ideal for engineers with strong expertise in both computer vision and deep learning who enjoy solving challenging real-world perception problems.

Responsibilities

  • Develop state-of-the-art 3D body and hand pose estimation models for egocentric video using monocular and stereo camera systems.
  • Build models for 2D/3D keypoint estimation, SMPL/SMPL-X, MANO, and full-body motion reconstruction.
  • Address challenging egocentric vision problems, including severe self-occlusion, motion blur, rolling shutter artifacts, truncated limbs, extreme viewpoints, and hand-object interaction.
  • Design and maintain camera geometry and calibration pipelines, including fisheye and wide-FOV camera models, stereo calibration, triangulation, and coordinate frame alignment.
  • Improve temporal consistency and physical plausibility using filtering, kinematic constraints, multi-view fusion, and multi-modal sensor integration.
  • Build scalable annotation, evaluation, and quality assurance pipelines for large-scale human motion datasets.
  • Optimize large-scale model training and high-throughput inference pipelines for production environments.
  • Collaborate closely with robotics engineers to convert reconstructed human motion into high-quality robot training data.
  • Contribute to system architecture, engineering best practices, and the long-term evolution of the perception platform.

Minimum Qualifications

  • Bachelor's, Master's, or PhD in Computer Science, Machine Learning, Computer Vision, Robotics, or a related field.
  • 3+ years of experience building and deploying machine learning systems.
  • Hands-on experience with 3D human pose estimation, hand pose estimation, or human motion tracking from video.
  • Strong understanding of multi-view geometry, camera calibration, triangulation, coordinate transformations, and projection models.
  • Strong Python programming skills and proficiency with PyTorch or TensorFlow.
  • Solid knowledge of modern deep learning techniques, model training, evaluation, and production ML workflows.
  • Strong analytical and problem-solving skills with the ability to thrive in a fast-paced collaborative environment.

Preferred Qualifications

  • Experience with egocentric perception systems, AR/VR headsets, smart glasses, or wearable capture rigs.
  • Expertise with SMPL, SMPL-X, MANO, inverse kinematics, markerless motion capture, or hand-object pose estimation.
  • Familiarity with egocentric vision datasets and benchmarks.
  • Experience with modern video and 3D learning architectures, including Video Transformers, diffusion-based motion models, or 3D CNNs.
  • Experience building multi-camera capture systems, synchronization, and calibration infrastructure.
  • Experience with human-to-robot motion retargeting, teleoperation, imitation learning, or dexterous manipulation.
  • Experience developing annotation tools, active learning pipelines, or large-scale data quality systems.
  • Publications at leading conferences such as CVPR, ICCV, ECCV, NeurIPS, SIGGRAPH, or 3DV, open-source contributions, or demonstrated impact in applied AI systems.

What We Offer

  • Competitive salary and equity package
  • Comprehensive medical, dental, and vision insurance
  • 401(k) retirement plan
  • Generous paid time off and company holidays
  • Paid sick leave
  • Opportunity to work alongside leading researchers and engineers in robotics, computer vision, and AI
  • High-impact role building cutting-edge perception systems for next-generation robotics



Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.