About the Role
This is a senior applied AI engineering role focused on building the AI systems layer behind a B2B process intelligence platform. You'll own the production systems that transform messy enterprise data into structured, reliable insights — sitting at the intersection of retrieval, orchestration, evaluation, and model integration. The work directly shapes the quality and usefulness of AI outputs delivered to real customers.
What You'll Do
Build and own production systems on top of hosted frontier LLM APIs (OpenAI, Anthropic, Gemini, Cohere).
Design and implement retrieval and context construction pipelines, including RAG, vector search, hybrid search, chunking, and reranking.
Create evaluation datasets, quality gates, regression tests, and LLM-as-judge workflows to measure and prevent model quality degradation.
Own structured extraction, model/provider selection, prompt optimization, and schema design to improve output quality.
Build and iterate on agent and tool orchestration workflows, including multi-step pipelines and planner/executor patterns.
Instrument systems for observability, cost tracking, latency monitoring, and failure mode handling.
Partner closely with backend, product, and forward-deployed engineering teams to solve real customer workflow problems.
What We're Looking For
3–6+ years of experience building and shipping production software, ML systems, or applied AI systems (not internal tooling or notebooks only).
Strong Python skills for building reliable services, data pipelines, eval workflows, and AI tooling.
Hands-on experience with LLM APIs, including prompting, structured outputs, embeddings, retrieval, and workflow orchestration.
Experience building evals, benchmarks, or CI/CD pipelines for model quality assessment.
Background in RAG, retrieval systems, or agent orchestration.
Breadth across multiple ML domains — not a narrow specialist in a single area.
Strong systems thinking: ability to reason end-to-end from raw data through model outputs to user-facing behavior.
Comfort with ambiguity — translating fuzzy product goals into experiments, implementations, and shipped improvements.
BS or higher in Computer Science, Engineering, or a related technical field.
Nice to have: experience with multimodal inputs (documents, screenshots, transcripts); familiarity with Go; experience with GCP, PostgreSQL, Redis, or Terraform.
Compensation & Benefits
Salary range: $180,000 – $275,000 USD annually. Visa sponsorship is not available for this role.
Location
On-site in New York, NY, United States. Candidates based in San Francisco may also be considered.
Learn more about this Employer on their Career Site
