SonicJobs Logo
Left arrow iconBack to search

Member of Technical Staff, Research — Early Career(PHD)

Goaly
Posted 10 days ago, valid for 14 days
Location

Menlo Park, San Mateo, CA

Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • The company is seeking a researcher with a focus on agentic reinforcement learning, requiring a PhD or equivalent research experience.
  • Candidates should have a strong background in machine learning, excellent Python skills, and experience with modern deep-learning frameworks such as PyTorch or JAX.
  • This role involves designing experiments, implementing new methods, and collaborating with engineers to scale ideas while ensuring they meet production constraints.
  • The position offers a competitive salary, with specifics not mentioned, and is suitable for those with a strong research record and experimental rigor.
  • The work environment promotes high agency, continuous learning, and a mission-first approach, with a hybrid policy requiring in-office presence at least three days a week.

About us

We are building AI systems that can reason, use tools, and complete meaningful work in the real world. Our team works across model post-training, reinforcement-learning infrastructure, large-scale training, and product engineering. We believe the fastest path to more capable and reliable agents is an integrated loop: challenging environments, rigorous evaluations, efficient training, reliable inference, and products that make those capabilities useful.

About the role

You will own research at the intersection of agentic reinforcement learning, post-training, evaluation, environments, and scaling. You will identify high-leverage questions, design and run decisive experiments, and translate results into model improvements and reusable systems.

This role is designed for researchers completing or recently completing a PhD, as well as candidates with an equivalent record of original research. It is not a purely academic position: strong candidates write excellent code, work closely with systems engineers, and care whether an idea survives realistic evaluation and production constraints.

What you'll do

  • Formulate high-leverage research questions about agentic capability and reliability, RL algorithms, reward and verifier design, exploration, curricula, environment design, task distributions, and scaling behavior.

  • Design rigorous experiments, ablations, controls, and evaluations that separate real model improvement from noise, data leakage, reward hacking, or benchmark overfitting.

  • Implement new methods in modern deep-learning frameworks and integrate them with production training, rollout, environment, and evaluation systems.

  • Build or improve datasets, agent environments, verifiers, and evaluations for domains such as coding, tool use, reasoning, long-horizon tasks, or computer interaction.

  • Analyze trajectories and model behavior, develop useful failure taxonomies, and turn observations into testable hypotheses and prioritized experiments.

  • Partner closely with Post-Training, RL Systems, Training, and Backend & Product engineers to scale promising ideas and expose them to realistic product constraints.

  • Communicate findings in clear internal documents and technical reviews; contribute to papers, technical reports, blog posts, or open-source releases when aligned with company goals.

  • Help shape the research roadmap by identifying compounding capabilities, reusable evaluation assets, and experiments that retire the most important uncertainties.

You may be a good fit if you have

  • Completing or recently completed a PhD in computer science, machine learning, statistics, mathematics, or a related field—or an equivalent record of original, technically rigorous research.

  • A strong research record in machine learning, reinforcement learning, large language models, agents, or ML systems, demonstrated through publications, preprints, open-source work, or substantial independent projects.

  • Excellent Python skills and hands-on experience with a modern deep-learning framework such as PyTorch or JAX.

  • Experimental rigor: you can define a falsifiable question, build the right measurement, control confounders, interpret noisy results, and communicate uncertainty honestly.

  • The engineering ability to navigate an unfamiliar codebase, build reliable research infrastructure, and turn a promising idea into a working system.

  • Clear written and verbal communication and the ability to collaborate across research, systems, and product disciplines in a fast-moving environment.

Strong pluses

  • Research experience in agentic reinforcement learning, post-training, preference learning, reward or verifier modeling, evaluation, or environment design.

  • Experience with large-model training, distributed inference, high-throughput rollout systems, or performance-sensitive ML infrastructure.

  • Notable publications, open-source contributions, datasets, benchmarks, or research artifacts that other people use.

  • Domain expertise in coding agents, mathematical reasoning, scientific discovery, tool use, long-horizon planning, or computer interaction.

  • Experience transferring a research result into a production model, product, or dependable shared system.

How we work

  • Mission first. We choose work for its impact on the mission and take responsibility for the outcome, not just our assigned tasks.

  • High agency. We identify what is missing, form a plan, and move without waiting for perfect clarity.

  • Speed with rigor. We ship, measure, and iterate quickly while protecting correctness, safety, and reliability.

  • Flexible scope. We cross team and technical boundaries when that is the fastest way to solve the real problem.

  • Low ego, high standards. We give direct feedback, change our minds when the evidence changes, and help the whole team win.

  • Continuous learning. The stack changes quickly; we are willing to learn unfamiliar systems, methods, and domains as the work demands.

Location, visa sponsorship & benefits

  • Location-based hybrid policy. This is a location-based hybrid role. We currently expect all staff to work from one of our offices at least three days per week. Exact office options will be confirmed during the recruiting process.

  • Visa sponsorship. We do sponsor visas. However, we cannot successfully sponsor a visa for every role and every candidate. If we make you an offer, we will make every reasonable effort to secure the necessary visa, and we retain immigration counsel to support the process.

  • Meals and office benefits. We provide complimentary lunch and dinner in our offices, along with snacks and beverages.

A note on qualifications. We care more about exceptional evidence than a perfect keyword match. If the work excites you and you can show unusual strength, learning speed, or ownership, we encourage you to apply even if your background does not match every preferred qualification.

Equal opportunity

We are an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity or expression, national origin, age, disability, veteran status, genetic information, or any other characteristic protected by applicable law. We provide reasonable accommodations for candidates who need them during the hiring process.




Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.