Responsibilities
- Develop joint audio-visual understanding systems that integrate visual and auditory signals for advanced perception
- Build and evaluate audiovisual language models for social interactions and understanding, including predicting social intent, semantic function, and reasoning from human-centric inputs
- Contribute to benchmarks and evaluation frameworks for visual social understanding and interactions
- Train and optimize state-of-the-art machine learning and neural network methodologies
- Conduct and collaborate on research projects within a globally-based team
Minimum Qualifications
- Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience
- A PhD in AI, computer science, data science, or related technical fields
- Experience holding an industry, postdoctoral, faculty, or government researcher position
- Research background in machine learning, artificial intelligence, computational statistics, or applied mathematics, or related areas
- Research publications reflecting experience in theoretical or empirical research
- Experience in developing and debugging in Python or similar programming languages
- Experience in analyzing and collecting data from various sources
- Must obtain work authorization in country of employment at the time of hire, and maintain ongoing work authorization during employment
Preferred Qualifications
- Demonstrated research and software engineering experience via an internship, work experience, coding competitions, or widely used contributions in open source repositories (e.g. GitHub)
- Experience with audio-visual learning or multimodal fusion techniques
- Familiarity with human action recognition, social signal processing, or human-centric video understanding
- Experience with long-form video understanding, video-language models, or streaming perception systems
- Experience with vision-language models (VLMs) such as LLaVA, GPT-4V, Gemini, or similar architectures
- Experience with temporal modeling, video transformers, or recurrent architectures for sequential data
$154,000/year to $217,000/year + bonus + equity + benefits
Learn more about this Employer on their Career Site
