Machine Learning Engineer - Agentic AI Evaluation Frameworks

Machine Learning and AI Cupertino, United States

Cupertino, United States SeniorOn-siteFull-time $184,700 – $277,600 a year

About the role

This is a senior machine learning evaluation engineering role at Apple, based in Cupertino, building evaluation frameworks for generative AI, conversational AI and agentic products. It suits an engineer who wants to build the datasets, tooling and quality metrics that let teams measure and improve LLM-powered experiences. The role sits on the Channel Sales AI Product Engineering team, covering Store AI, Shopping AI, Learning AI, Content GenAI, Conversational AI and Platform Self-Service.

What you’ll do

  • Design and build automated evaluation frameworks and pipelines for AI-powered products.
  • Define evaluation methodologies and quality metrics across accuracy, relevance, groundedness, consistency and task completion.
  • Build and maintain golden datasets, benchmark sets, regression suites and adversarial test sets.
  • Develop Auto Eval capabilities for rapidly evaluating models, prompts, retrieval systems and agents.
  • Implement model-based evaluation approaches such as LLM-as-a-Judge, with calibration and validation.
  • Develop human-in-the-loop evaluation for complex or subjective quality dimensions.
  • Define evaluation rubrics, annotation guidelines and grading criteria with product and domain teams.
  • Calibrate automated evaluators against human judgment and measure evaluator consistency.
  • Evaluate end-to-end AI systems, including retrieval, prompts, tool use and downstream product experiences.
  • Develop evaluation methods for multi-turn conversations, personalization and agentic task execution.
  • Perform error analysis and failure-mode investigation.
  • Build reusable evaluation infrastructure, APIs, dashboards and developer tooling.

What they’re looking for

  • Typically 7+ years in ML engineering, ML evaluation, software engineering, data science or quality engineering.
  • Strong Python skills and experience building production-quality ML systems or evaluation infrastructure.
  • Experience developing or evaluating LLMs, generative AI, conversational AI, NLP or recommendation systems.
  • Experience designing automated ML evaluation frameworks, metrics, benchmarks or datasets.
  • Understanding of modern LLM application architectures, including RAG, tool use and agentic workflows.
  • Experience with model-based evaluation techniques such as LLM-as-a-Judge.
  • Experience with human-in-the-loop evaluation, annotation or data-quality workflows.
  • Strong grounding in statistical analysis, experimentation and measurement methodologies.
  • Experience with model error and root-cause analysis.
  • Ability to work across ML, engineering, product, quality and data teams.
  • Bachelor's degree in Computer Science, ML, AI, Data Science, Statistics, Electrical Engineering or related field, or equivalent experience.

Nice to have

  • Experience building evaluation infrastructure for production-scale LLM or GenAI applications.
  • Experience evaluating RAG, conversational systems, AI agents, personalization or multimodal AI.
  • Experience building golden datasets, regression suites and automated quality gates.
  • Experience integrating ML evaluation into CI/CD and release processes.
  • Experience with prompt evaluation, model comparison and AI observability.
  • Experience evaluating multilingual AI experiences.
  • Familiarity with responsible AI evaluation, including robustness, safety and bias testing.
  • Experience building internal ML platforms or self-service evaluation tooling.
  • Experience with large-scale datasets and distributed ML infrastructure.
  • Master's degree in a related technical field.

What’s on offer

  • Base pay range of $184,700 to $277,600, depending on skills, qualifications, experience and location.
  • Eligibility for Apple's discretionary stock programs and Employee Stock Purchase Plan.
  • Comprehensive medical and dental coverage and retirement benefits.
  • Discounted products and free services.
  • Educational expense reimbursement, including tuition.
  • Possible discretionary bonuses, commission or relocation.

Questions about this role

How much does it pay?

The base pay range is $184,700 to $277,600, depending on skills, experience and location.

How many years of experience do I need?

Typically a minimum of 7 years of related experience.

Do I need a degree?

A bachelor's degree in a related technical field, or equivalent industry experience.

Is this role remote?

The posting lists Cupertino, California as the work location; it does not describe remote or hybrid arrangements.

Related roles

Machine Learning-AI and Data Science Engineer III

Deloitte Consulting · Bengaluru, India · Hyderabad, India · Pune, India · Chennai, India

A senior data science and ML engineering leadership role at Deloitte Consulting, based across Bengaluru, Hyderabad, Pune or Chennai. Suits someone with 6-10 years of experience who can lead a small DS/ML team through client delivery.

On-siteSeniorFull-time $33,700 – $53,000 a year
Machine LearningData ScienceStatistical ModelingPython +11
4 days ago

Senior Machine Learning Engineer, Ad Tech Engineering

Roku · Bengaluru, India

Senior hybrid machine learning role in Bengaluru for an experienced engineer building AI systems for Roku's advertising platform. You'll work on multimodal ad understanding, brand safety, moderation, creative generation, and large-scale model serving.

HybridSeniorFull-time
Machine LearningDeep LearningComputer VisionGenerative AI +13
5 days ago

Senior Machine Learning Engineer - Advertising Engineering

Roku · Bengaluru, India

Senior hybrid machine learning engineering role in Bengaluru focused on Roku's advertising data and AI platform. You'll build large-scale ML systems and agentic generative AI features for targeting, measurement, reporting, and optimization.

HybridSeniorFull-time
Machine LearningPythonScalaJava +16
5 days ago

Machine Learning Engineer, Dubbing

Sarvam AI · Bengaluru, India

Sarvam AI is hiring a mid-level Machine Learning Engineer in Bengaluru to build production systems for dubbing and live translation. The role requires speech and NLP experience, production ML skills and Python.

On-siteMid level
Machine LearningPythonSpeech RecognitionNatural Language Processing +4
Date not stated

Machine Learning Engineer

Weekday Partner Network · Bengaluru, India

Machine Learning Engineer role in Bengaluru for someone with 5–12 years of experience designing models, data pipelines, and production ML systems. The position also involves shaping the roadmap and delivering scalable AI applications.

On-siteMid levelFull-time ₹5,000,000 – ₹12,000,000 a year
Machine LearningData PipelinesModel TrainingML Deployment +2
12 days ago

Machine Learning Engineer

Nykaa · Bengaluru, India

Nykaa is hiring a mid-level machine learning engineer in Bengaluru to productionize ML systems used in e-commerce. The role requires 5 to 7 years of experience along with Python, MLOps, model serving, monitoring, and AWS or cloud skills.

HybridMid level
Machine LearningMLOpsPythonAWS +3
Date not stated