SonicJobs Logo
Left arrow iconBack to search

Software Engineer, AI Runtime & Platform Services

CrewAI
Posted 2 months ago, valid for 12 days
Location

San Francisco, CA 94102, US

Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • CrewAI is seeking a backend/platform engineer to enhance its enterprise runtime layer, focusing on building secure and observable production systems.
  • The role requires strong Python experience, particularly with FastAPI, Celery, and Redis, alongside a solid understanding of distributed systems and security practices.
  • Candidates should possess a minimum of 5 years of relevant experience in production services and have a strong testing background.
  • The position offers a competitive salary of $120,000 to $150,000 per year, depending on experience and qualifications.
  • Ideal applicants will be adept at collaborating across teams while maintaining the integrity of both open-source and enterprise components.
About CrewAI

CrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production.

The Role

You'll work on the enterprise runtime layer that turns CrewAI's open-source Crews and Flows into secure, observable, remotely executable production systems. This is the layer between the framework and the platform: APIs, workers, checkpoints, webhooks, auth, deployment behavior, telemetry, and enterprise extensions that make CrewAI run reliably in real customer environments.

You'll partner closely with the open-source, product, and infrastructure teams, but your center of gravity is production execution: making agent workflows resumable, inspectable, authenticated, observable, and safe to operate at scale.

What You'll Do
  • Build and maintain the Python enterprise runtime around CrewAI: FastAPI services, Celery workers, Redis-backed state, execution APIs, and deployment-facing tools.
  • Extend open-source CrewAI behavior for enterprise environments while preserving compatibility with upstream framework changes.
  • Own production execution flows: crew and flow kickoff, status, retries, cancellation, checkpoint restore and fork, chat/session state, and human-in-the-loop resume paths.
  • Build secure integration surfaces: JWT auth, signed webhooks, token refresh, file handling, secret fetching, and workload identity across AWS, GCP, and Azure.
  • Improve observability across distributed execution: OpenTelemetry traces, structured logs, Sentry, event tracking, and debuggability across API, worker, and platform boundaries.
  • Maintain strong test coverage for async/runtime behavior using pytest, mypy, ruff, mocks/fakes, and e2e deployment harnesses.
  • Partner with the Agent Management Platform team on API contracts, versioning, enterprise client behavior, deployment status, and failure reporting.
What We're Looking For
  • Strong Python backend/platform engineering experience, especially building production services rather than only libraries.
  • Experience with FastAPI or similar API frameworks, Celery or other job systems, Redis, Pydantic, and typed Python.
  • Good instincts for distributed systems: retries, idempotency, async execution, status tracking, race conditions, and failure recovery.
  • Comfort with auth and security-sensitive systems: JWTs, webhooks, signatures, secrets, IAM/workload identity, and least-privilege thinking.
  • Practical observability experience: tracing, structured logging, metrics, Sentry/OpenTelemetry, and debugging multi-service failures.
  • Ability to work at the boundary between an open-source framework and a hosted enterprise platform without creating brittle coupling.
  • Strong testing habits and comfort with CI, package/version management, and release discipline.
Bonus
  • Experience operating AI/agent runtimes, workflow engines, or distributed task systems.
  • Cloud platform experience with AWS ECS/ECR, Kubernetes, Helm, GCP/Azure identity, or secret managers.
  • Experience with enterprise SaaS constraints: auditability, tenant isolation, customer environments, deployment rollbacks, and supportability.
  • Familiarity with Rails/SaaS platforms is useful, but not required.



Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.