SonicJobs Logo
Left arrow iconBack to search

Research Scientist, Text-to-Speech

Oddin
Posted 6 months ago, valid for 21 days
Location

Malabar, Brevard County 32950, FL

Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • Valka.ai, a spin-off from the Realms Group, is seeking a candidate to revolutionize digital content creation through AI-driven interactive platforms.
  • The role involves researching and training state-of-the-art text-to-speech (TTS) models for applications in entertainment and education.
  • Candidates should have experience in training TTS or voice cloning models, along with proficiency in Python and knowledge of key machine learning frameworks.
  • A minimum of 3 years of experience in relevant fields is required, and the position offers a salary of $80,000 to $120,000 per year.
  • The ideal candidate will collaborate with a team of researchers and stay updated with the latest advancements in speech synthesis technology.

 

About Valka.ai

 

Valka, a visionary spin-off from the Realms Group (the parent company of Oddin.gg), is on a mission to revolutionize the way people create and experience digital content.

 

Our team believes that content shouldn’t just be consumed; it should be co-created in real time, blurring the lines between imagination and reality. By harnessing the power of cutting-edge AI, we aim to build an interactive human-digital platform where virtual characters respond dynamically to each user’s voice, text, gestures, and more.

 

This is your chance to join a diverse group of innovators who are driven to redefine what’s possible in generative content. Together, we’re changing the paradigm from passive viewing to active participation, unlocking new creative frontiers across gaming, entertainment, education, and beyond.

\n


What you will be doing
  • Research and train fast and quality SOTA TTS models for realistic and emotional voice generation for entertainment and education applications.
  • You will be experimenting with different architectures / data to improve the quality and speed of the TTS model(s) and put the best results to production.
  • Staying up to date with current research and coming up with new ideas / what to improve is very important for us!
  • You will be in immediate collaboration with a team of 3 researchers specializing in TTS, and the product is supported by engineering and hardware stuff to ensure deployment


Skills you need
  • Experience with training some text-to-speech / voice cloning models
  • Solid knowledge of transformers, diffusion models, GANs
  • Understanding of human speech and audio processing (sampling, spectrograms, vocoders)
  • Proficiency in Python and key libraries (e.g., PyTorch, Hugging Face Transformers).
  • Ability to keep up to date with research, understand papers, implement approaches; strong ML fundamentals and critical thinking
 
Nice-to-have:
 
  • Familiarity with modern speech synthesis models (GPT-based, flow matching… such as Vevo, StyleTTS, IndexTTS, Maskgct etc.)
  • Contributions to open-source AI tools or research publications in Speech processing field
  • Familiarity with AWS / similar clusters


\n



Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.