SonicJobs Logo
Left arrow iconBack to search

Research Scientist, Gemini Audio i18n, DeepMind

Google
Posted 4 days ago, valid for 18 days
Location

New York, NY, US

Salary

$174,000 - $252,000 per year

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • Google is seeking a Research Scientist with a Bachelor's degree in Computer Science or a related field and at least 3 years of experience in Large Language Models or multimodal foundation models.
  • The role involves conducting research and development in Speech Recognition, Text-to-Speech, or related areas, and requires proficiency in Python or C++ along with deep learning frameworks like PyTorch, JAX, or TensorFlow.
  • Responsibilities include creating evaluation sets for audio model performance, proposing novel modeling techniques, and collaborating with teams to address performance gaps in multilingual audio models.
  • The position offers a salary range of $174,000 to $252,000 USD, along with a 15% bonus target, equity, and benefits.
  • Candidates with a Master's degree or Ph.D. in relevant fields and a publication record in speech or machine learning conferences are preferred.

Minimum qualifications:

  • Bachelor's degree in Computer Science, Speech Recognition, Computational Linguistics, a related technical field, or equivalent practical experience.
  • Experience conducting research or development in Speech Recognition, Text-to-Speech (TTS), or Large Language Models (LLMs).
  • Experience coding in Python or C++ and using deep learning frameworks such as PyTorch, JAX, or TensorFlow.
  • Experience working with audio data, speech processing, or multilingual datasets.

Preferred qualifications:

  • Master's degree or Ph.D. in Automatic Speech Recognition (ASR), Text-to-Speech (TTS), Computational Linguistics, or a related field.
  • 3 years of experience in Large Language Models (LLMs) or multimodal foundation models.
  • Experience developing audio-to-audio (A2A) architectures, end-to-end speech models, or spoken dialog systems.
  • Experience scaling speech models across international languages, accents, or low-resource locales.
  • Publication record in speech or machine learning conferences (e.g., ICASSP, INTERSPEECH, NeurIPS, or ACL).

About the job:

As an organization, Google maintains a portfolio of research projects driven by fundamental research, new product innovation, product contribution and infrastructure goals, while providing individuals and teams the freedom to emphasize specific types of work. As a Research Scientist, you'll setup large-scale tests and deploy promising ideas quickly and broadly, managing deadlines and deliverables while applying the latest theories to develop new and improved products, processes, or technologies. From creating experiments and prototyping implementations to designing new architectures, our research scientists work on real-world problems that span the breadth of computer science, such as machine (and deep) learning, data mining, natural language processing, hardware and software performance analysis, improving compilers for mobile platforms, as well as core search and much more.

As a Research Scientist, you'll also actively contribute to the wider research community by sharing and publishing your findings, with ideas inspired by internal projects as well as from collaborations with research programs at partner universities and technical institutes all over the world.

Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.

We are pushing the boundaries across multiple domains. Our global teams offer diverse learning opportunities and varied career pathways for those driven to achieve exceptional results through collective effort.
Individual pay is determined by factors including job-related skills, experience, and relevant education or training.

US: $174000 - $252000 (USD) + 15% bonus target + equity + benefits

Learn more about benefits at Google.

Responsibilities:

  • Create comprehensive evaluation sets and benchmarks to measure audio-to-audio (A2A) model performance across international languages, accents, and regional dialects.
  • Propose, prototype, and evaluate novel modeling techniques to improve A2A audio understanding, dialog, and audio generation capabilities with a focus on scalability.
  • Identify performance gaps in current multilingual audio models and collaborate with cross-functional research and engineering teams to deploy solutions.



Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.