SonicJobs Logo
Left arrow iconBack to search

Data Scientist — Agent Evaluations & Quality

Clera
Posted a day ago, valid for 20 days
Location

Palo Alto, CA, US

Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • The role is for a Data Scientist focused on Agent Evaluations & Quality at an AI startup in Palo Alto, CA, with a team of approximately 25 people.
  • Candidates must have at least 5 years of experience in data science, machine learning, or analytics, specifically in evaluation systems and quality measurement for production systems.
  • The position involves architecting automated evaluation pipelines, defining success criteria, and building datasets and metrics to improve AI agent capabilities.
  • Proficiency in Python and SQL is required, along with strong statistical and experimental design skills, and the ability to communicate evaluation results effectively to diverse stakeholders.
  • Compensation will be competitive with market rates for senior applied data science roles at early-stage AI startups, with specific details shared during the interview process.

About the Role

We're a ~25-person AI startup building autonomous agents that handle real, complex work — email, calendar, browser automation, business software, and more. We're looking for a Data Scientist focused on Agent Evaluations & Quality to join us on-site in Palo Alto, CA.

Your mission: measure, understand, and continuously improve the quality of our AI agent capabilities. You'll turn ambiguous product behavior into rigorous, measurable definitions of success — then build the evaluation infrastructure, datasets, metrics, and feedback loops that drive engineering and product decisions. This is applied data science at the intersection of LLM systems, evaluation design, and production quality engineering.

Visa sponsorship is not available for this role.

What You'll Do

  • Architect and maintain automated evaluation pipelines that measure agent quality across capabilities and product surfaces.

  • Translate agent capabilities into explicit success criteria — defining pass, partial-pass, and failure conditions for complex multi-step tasks.

  • Build representative gold datasets and regression suites covering common workflows, ambiguous requests, long-tail behavior, edge cases, and adversarial scenarios.

  • Define and track metrics spanning task success, partial completion, tool-selection accuracy, tool-use correctness, instruction adherence, factual consistency, user corrections, latency, cost, and reliability.

  • Design deterministic graders, model-based graders, and human-review processes; calibrate LLM-as-a-judge systems and measure false positives, false negatives, variance, and grader agreement.

  • Analyze traces, tool calls, model outputs, user context, and production outcomes to identify root causes and build a useful failure taxonomy.

  • Compare models, prompts, tools, and capability implementations using rigorous offline experiments and production evidence.

  • Build dashboards, reports, and release-quality signals that make evaluation results clear and actionable for engineering, product, and leadership.

  • Partner with capability engineers to recommend improvements and verify that fixes raise quality without unacceptable regressions in cost, latency, or reliability.

What We're Looking For

Must-haves:

  • 5+ years of experience in data science, machine learning, or analytics roles — with demonstrated focus on evaluation systems, metrics frameworks, or quality measurement for production systems.

  • Proven experience designing and implementing evaluation frameworks, grading systems, and metrics for ML or AI systems in production.

  • Production-quality Python and SQL proficiency; ability to build automated data pipelines and analysis code at scale.

  • Deep expertise in evaluation methodology: success criteria definition, dataset construction, metric selection, and distinguishing meaningful benchmarks from misleading ones.

  • Strong statistical and experimental design skills: sampling, variance, uncertainty quantification, bias detection, confounding, and significance testing for non-deterministic systems.

  • Experience with ground-truth data development: labeling guideline design, annotation quality control, ambiguity resolution, and dataset maintenance as product behavior evolves.

  • Working knowledge of LLM behavior, model-based graders, tool use, retrieval systems, multi-step execution, partial completion, and real-world failure modes of language model systems.

  • Analytical debugging ability: connecting quantitative patterns to individual system traces and identifying failure origins across model, prompt, context, tools, data, and application logic.

  • Experience communicating evaluation results, methodology, uncertainty, and trade-offs to both technical and non-technical stakeholders.

Nice to have:

  • Experience with LLM-as-a-judge systems, calibration, and grader agreement measurement.

  • Prior work on evaluation or benchmarking platforms for AI systems.

  • Experience with agentic systems, multi-step task execution, or tool-use evaluation.

  • Background working on customer-facing consumer software or production ML systems with real-world user outcomes.

Location & Work Arrangement

This is a full-time, on-site role based in Palo Alto, CA. Remote work is not available for this position.

Compensation & Benefits

Compensation details will be shared during the interview process and will be competitive with market rates for senior applied data science roles at early-stage AI startups.




Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.