LLM Evaluation jobs
14 open roles for llm evaluation, updated continuously as employers post them.
Hiring most here right now:
Senior Machine Learning Engineer
A senior machine learning engineer role at Atlassian building production ML and evaluation systems for the Rovo AI agent platform.
Applied AI Engineer, Digital Natives
A hybrid, London-based Applied AI Engineer role at OpenAI, deploying production AI systems for Digital Native customers.
AI Data Readiness Lead
A remote, senior data governance role at Deepgram owning metric definitions and verifying how people and AI agents use company data.
Agent Reliability Engineer, GTM
An on-site/hybrid San Francisco role at LangChain operating and improving the internal AI agent that powers its go-to-market team.
AI Test Engineer
Qentelli is hiring an AI Test Engineer in Hyderabad to establish quality engineering for production agentic and multi-agent AI systems. The role suits a senior QA practitioner with Python, distributed-systems testing and LLM evaluation experience.
Senior GenAI / Agentic AI Engineer
Inspironlabs is hiring a Senior GenAI / Agentic AI Engineer for onsite work in Bengaluru, Gurugram or Mumbai. The role focuses on production RAG systems, multi-agent architectures, Python services and enterprise AI integration.
Senior LLM Engineer
An onsite senior engineering role building and operating production LLM and multi-agent systems in Ahmedabad or Pune. It suits engineers with Python, LangGraph, cloud, MLOps, and agent-orchestration experience.
AI Engineer – Generative AI
Onsite AI Engineer role in Hyderabad, Ahmedabad, or Indore building generative-AI and agentic systems. The position suits an engineer with at least six years of experience across LLMs, RAG, Python, and cloud deployment.
AI Engineer - LLM
AI Engineer role in Bangalore for an experienced software engineer building production LLM, RAG, and enterprise AI systems. It is suited to candidates with Python, retrieval, NLP, and generative AI application experience.
Generative AI Engineer
Senior hybrid generative AI engineer role in Bengaluru for a professional with at least six years of engineering experience. The work covers LLM applications, evaluations, guardrails, RAG, observability, prompt optimization, and production delivery.
Prompt Engineer
Remote worldwide prompt engineer role for a mid-level practitioner with production LLM experience. You'll work on prompts, agent workflows, evaluations, and reliable GenAI product delivery.
Applied AI Engineer (LLM Systems)
Nference is hiring a mid-level Applied AI Engineer in Bengaluru to build and evaluate retrieval-augmented clinical reasoning systems. The role is suited to engineers working with Python, RAG, LLM evaluation, and PyTorch.
Senior Data Scientist, AI Scoring & Evaluation
A senior data science role at Workera owning the accuracy and fairness of AI-driven skills assessment scoring. Suits a data scientist with production experience who wants to build evaluation systems for LLM-based assessment products.
Applied AI Engineer
Perplexity AI is hiring a mid-level Applied AI Engineer for its India team to build retrieval, model-orchestration and evaluation systems remotely from a Bengaluru anchor. The role requires production engineering, retrieval or RAG experience and LLM evaluation skills.