Description
You will work closely with Multimodal GenAI researchers and engineers to develop world-class audio-video perception and representation algorithms. The role focuses on building, integrating and maintaining complex real-time systems within a large production software stack. Deep understanding of machine learning inference and performance oriented development is highly desirable. As a member of a fast-paced team, you have the unique and rewarding opportunity to shape upcoming products that will delight and inspire millions of people every day. The ideal candidate should possess these qualities: * Be highly-motivated and take initiative to achieve goals, while delivering on schedule. * Has a sense of curiosity and willingness to learn new things in order to improve the quality of their solutions. * Works well in a collaborative setting.
Minimum Qualifications
Bachelor’s degree in Computer Vision, Computer Graphics, Machine Learning, Computer Science, Computer Engineering or related fields. 5+ years of relevant industry experience. Strong proficiency in C/C++ or Swift, writing clean and well-structured code. Solid mathematical foundation in linear algebra and numerical analysis. Experience breaking down large problems and delivering robust solutions. Experience developing user-facing APIs for complex systems. Curiosity about new technologies and flexibility to work across different layers of the software stack.
Preferred Qualifications
MS/PhD with relevant industry experience. Experience with audio, computer vision, 3D graphics and animation algorithms. Experience deploying algorithms to GPUs or other hardware accelerators. Understanding of SW/HW parallelism, threads, processes and asynchronous processing. iOS or macOS development experience using frameworks like Foundation Models, MLX, ARKit, REKit. Knowledge of diffusion transformer models especially in context of on-device inference.
Learn more about this Employer on their Career Site
