Agent Reliability Engineer, GTM
Skills
About the role
This is an Agent Reliability Engineer role on LangChain's GTM Engineering team, responsible for the health, cost, performance and business impact of the internal GTM Agent. It suits someone with SRE instincts and strong production Python/SQL skills who wants to operate an LLM agent in production and build the monitoring, evals and reporting around it. The role is based in San Francisco, CA, on-site or hybrid.
What you’ll do
- Monitor production health across every graph, catching errors, slow runs, expensive runs and silent failures
- Triage incoming issues from Slack, tickets and rep reports
- Run the weekly eval suite and turn production bugs into permanent regression tests
- Track cost and latency by model, graph, use case and role, recommending changes to model choice and caching
- Track usage and adoption per rep and feature, owning the weekly health report
- Build business metrics showing leadership the agent's value, from reply rates to ROI
- Build monitoring and alerting on LangSmith and write up how the team runs its own agent
What they’re looking for
- Strong production Python and SQL, comfortable working in traces, logs and warehouse tables
- Real experience running LLM applications, including tracing, evals, and prompt/cache mechanics
- SRE or production operations instincts around percentiles, SLOs and separating noise from real patterns
- Healthy skepticism about metrics
- Clear writing and interest in publishing what you learn
- High agency and initiative
Nice to have
- LangGraph or LangSmith experience
- Experience building an eval suite from scratch
- BigQuery or dbt
- Prior DevRel-adjacent writing
- Empathy for sales and go-to-market users
What’s on offer
- $150,000-$190,000 base salary
- Medical, dental and vision coverage
- Flexible vacation
- 401(k) plan
- Meals on in-office days in the US
Questions about this role
Is this role remote?
No, the posting lists it as on-site/hybrid in San Francisco, CA.
How much does it pay?
The stated salary range is $150,000 to $190,000.
Related roles
Deployed Engineer, Professional Services
Remote Deployed Engineer role for an experienced Python engineer who has shipped production AI agents. You'll work directly with enterprise customers on architecture, evaluation, implementation, and deployment.
Engineer - Engine Reliability
Onsite engine-reliability engineering role in Chennai for an automotive product-development professional with two to nine years of relevant experience. The work covers powertrain calibration, testing, diagnostics, reliability, emissions, and vehicle integration.
Software Engineer (AI-Powered Advertising Agents - AgenticOS)
A software engineering role at PubMatic in Pune building AI agents for programmatic advertising, open to candidates with 2-10 years of experience.
Agentic Python Engineer
Evaboot is hiring a remote Agentic Python Engineer to build SDR agents and self-improving company systems. The role is suited to strong Python developers who use agentic coding tools.
Cloud Engineer 3 - AI Agents
A Bangalore-based cloud engineering role at Adobe building AI agents and tooling for cloud cost optimization.
Agentic AI Engineer
Remote full-time Agentic AI Engineer role at General Dynamics Information Technology. The posting is aimed at engineers building enterprise AI systems with retrieval, model operations and automated decision support.