World Emblem is a $100M+ manufacturer scaling to $250M across the US, Mexico, Dominican Republic, and Canada. We are building an owned AI agent stack inside the company — not buying a tool, not renting a black box.
Job Summary
As AI Engineer, you design, build, and ship multi-step agentic systems into live WEI workflows across sales, customer service, operations, and ecommerce — reaching real business data in NetSuite, HubSpot, Business Central, and our ecommerce platforms. You will work on the shared platform layer (orchestration, retrieval, evals, observability) and own agents end to end: from specification through evaluation to production operation. The scoreboard is agents running in production with a measured operational delta — not slideware.
Essential Duties & Responsibilities
Design and ship multi-step agentic systems (planner/executor, tool-using, multi-agent, human-in-the-loop) for live business workflows across sales, customer service, operations, and ecommerce. Â
Architect agent graphs in LangGraph (or comparable — CrewAI, AutoGen, Claude Agent SDK) with explicit state, durable execution, retries, and safe fallbacks. Â
Build the retrieval layer powering our agents — chunking, embeddings, hybrid search, reranking, and grounded citation. Â
Run eval and override discipline on every build: golden sets, offline regression suites, LLM-as-judge, online A/B and shadow evals, and red-teaming for jailbreaks, prompt injection, and data leakage. No agent ships without a measured pass bar. Â
Expose agents to production systems (NetSuite, HubSpot, Business Central, ecommerce) via well-typed tools and MCP servers. Treat tool surface area as a product. Â
Drive production MLOps: deployment, versioning, traffic shaping, cost/latency budgets, tracing, and on-call playbooks for agent incidents. Â
Partner with IT, Security, and Legal to keep agents inside WEI’s data-handling, access-control, and acceptable-use posture — auditability, explainability, and human-override paths built in, not bolted on. Â
Keep IP hygiene tight: prompts, code, orchestration, and eval scaffolding are WEI-owned assets — documented, version-controlled, and reusable, so each build makes the next one faster.
Technology Stack
Languages: Python, Node.js, TypeScript Â
Agent / LLM Frameworks: LangGraph, LangChain, Claude Agent SDK, MCP, OpenAI SDK Â
Models: Anthropic Claude, OpenAI, open-weight where appropriate — behind a model-agnostic abstraction Â
Retrieval & Data: PostgreSQL, pgvector, OpenSearch, Redis Â
Business Systems: NetSuite, HubSpot, Microsoft Business Central, ecommerce Â
Infra: AWS, Kubernetes (EKS), Terraform Â
Evals & Observability: LangSmith / Langfuse / Braintrust-style tooling, DataDog.
What We’re Looking For (Required Requirements)
5+ years of software engineering experience, with 2+ years building production LLM or agentic systems — not just notebooks, prototypes, or demos. Â
Hands-on experience with a modern agent framework (LangGraph strongly preferred) and a track record of shipping agents that run, fail gracefully, and recover. Â
Strong RAG fundamentals — chunking, embeddings, hybrid retrieval, reranking, grounding — and judgment about when RAG isn’t the right answer. Â
Real eval experience: golden sets, offline and online evaluations, used to make ship/no-ship calls. Â
Production MLOps fluency: has deployed LLM workloads under real latency, cost, and reliability constraints. Â
Strong Python; comfortable in TypeScript / Node.js. Â
Solid systems engineering instincts: APIs, async patterns, queues, databases, and distributed-system failure modes. Â
Calibrated communicator who thrives in ambiguous, fast-moving environments and can translate business needs into shipped solutions. Â
Bachelor’s in CS, MIS, Data Science, Engineering, Mathematics, or equivalent hands-on experience.
Preferred (Nice to Have Requirements)
Experience building MCP servers or other structured tool interfaces for LLMs. Â
Experience integrating with ERP / CRM / ecommerce platforms (NetSuite, HubSpot, Business Central). Manufacturing background optional. Â
Background in classical ML (ranking, scoring, calibration). Â
Experience designing explainable / auditable AI workflows with guardrails for LLM output. Â
Open-source contributions to agent frameworks, eval tooling, or retrieval libraries. Â
AWS depth (EKS, RDS, S3, Lambda) and infrastructure-as-code with Terraform. Â
Multilingual (Spanish).
This is Not the Role If You Want…
Research, slideware, or pilots that never ship. Â
To stay in notebooks. Agents here run in production against live workflows, and you own them when they break. Â
A narrow lane. You will touch retrieval, evals, tooling, infra, and the business process itself.
Success Metrics
Agent Quality: Measurable improvements in task success rate, grounding accuracy, and hallucination rate on our eval suites. Â
Production Reliability: Agents you own meet defined SLOs for latency (P90/P99), tool-call success, and cost per task. Â
Velocity: New agent capabilities go from prototype to production in weeks — without skipping evals or guardrails. Â
Risk Posture: Zero material incidents tied to prompt injection, data leakage, or unsafe tool use on agents you own. Â
Force Multiplier: Patterns, tools, and eval scaffolding you build get adopted across the AI engineering unit and our build partners.
Supervisory / Working EnvironmentIndividual contributor; no direct reports. Reports to the Head of AI Engineering and collaborates cross-functionally and with external build partners. Standard office / home-office environment; occasional after-hours support for deployments. Travel up to ~10%.
World Emblem is an Equal Opportunity Employer (EOE). Qualified applicants are considered for employment without regard to age, race, color, religion, sex, national origin, sexual orientation, disability, or veteran status. World Emblem is proud to be a drug-free workplace. All applicants will undergo a criminal background check and pre-placement drug screen, and we participate in E-Verify.
Learn more about this Employer on their Career Site
