Machine Learning Engineer - Agentic AI Evaluation Frameworks
Skills
About the role
This is a senior machine learning evaluation engineering role at Apple, based in Cupertino, building evaluation frameworks for generative AI, conversational AI and agentic products. It suits an engineer who wants to build the datasets, tooling and quality metrics that let teams measure and improve LLM-powered experiences. The role sits on the Channel Sales AI Product Engineering team, covering Store AI, Shopping AI, Learning AI, Content GenAI, Conversational AI and Platform Self-Service.
What you’ll do
- Design and build automated evaluation frameworks and pipelines for AI-powered products.
- Define evaluation methodologies and quality metrics across accuracy, relevance, groundedness, consistency and task completion.
- Build and maintain golden datasets, benchmark sets, regression suites and adversarial test sets.
- Develop Auto Eval capabilities for rapidly evaluating models, prompts, retrieval systems and agents.
- Implement model-based evaluation approaches such as LLM-as-a-Judge, with calibration and validation.
- Develop human-in-the-loop evaluation for complex or subjective quality dimensions.
- Define evaluation rubrics, annotation guidelines and grading criteria with product and domain teams.
- Calibrate automated evaluators against human judgment and measure evaluator consistency.
- Evaluate end-to-end AI systems, including retrieval, prompts, tool use and downstream product experiences.
- Develop evaluation methods for multi-turn conversations, personalization and agentic task execution.
- Perform error analysis and failure-mode investigation.
- Build reusable evaluation infrastructure, APIs, dashboards and developer tooling.
What they’re looking for
- Typically 7+ years in ML engineering, ML evaluation, software engineering, data science or quality engineering.
- Strong Python skills and experience building production-quality ML systems or evaluation infrastructure.
- Experience developing or evaluating LLMs, generative AI, conversational AI, NLP or recommendation systems.
- Experience designing automated ML evaluation frameworks, metrics, benchmarks or datasets.
- Understanding of modern LLM application architectures, including RAG, tool use and agentic workflows.
- Experience with model-based evaluation techniques such as LLM-as-a-Judge.
- Experience with human-in-the-loop evaluation, annotation or data-quality workflows.
- Strong grounding in statistical analysis, experimentation and measurement methodologies.
- Experience with model error and root-cause analysis.
- Ability to work across ML, engineering, product, quality and data teams.
- Bachelor's degree in Computer Science, ML, AI, Data Science, Statistics, Electrical Engineering or related field, or equivalent experience.
Nice to have
- Experience building evaluation infrastructure for production-scale LLM or GenAI applications.
- Experience evaluating RAG, conversational systems, AI agents, personalization or multimodal AI.
- Experience building golden datasets, regression suites and automated quality gates.
- Experience integrating ML evaluation into CI/CD and release processes.
- Experience with prompt evaluation, model comparison and AI observability.
- Experience evaluating multilingual AI experiences.
- Familiarity with responsible AI evaluation, including robustness, safety and bias testing.
- Experience building internal ML platforms or self-service evaluation tooling.
- Experience with large-scale datasets and distributed ML infrastructure.
- Master's degree in a related technical field.
What’s on offer
- Base pay range of $184,700 to $277,600, depending on skills, qualifications, experience and location.
- Eligibility for Apple's discretionary stock programs and Employee Stock Purchase Plan.
- Comprehensive medical and dental coverage and retirement benefits.
- Discounted products and free services.
- Educational expense reimbursement, including tuition.
- Possible discretionary bonuses, commission or relocation.
Questions about this role
How much does it pay?
The base pay range is $184,700 to $277,600, depending on skills, experience and location.
How many years of experience do I need?
Typically a minimum of 7 years of related experience.
Do I need a degree?
A bachelor's degree in a related technical field, or equivalent industry experience.
Is this role remote?
The posting lists Cupertino, California as the work location; it does not describe remote or hybrid arrangements.
Related roles
Machine Learning-AI and Data Science Engineer III
A senior data science and ML engineering leadership role at Deloitte Consulting, based across Bengaluru, Hyderabad, Pune or Chennai. Suits someone with 6-10 years of experience who can lead a small DS/ML team through client delivery.
Senior Machine Learning Engineer, Ad Tech Engineering
Senior hybrid machine learning role in Bengaluru for an experienced engineer building AI systems for Roku's advertising platform. You'll work on multimodal ad understanding, brand safety, moderation, creative generation, and large-scale model serving.
Senior Machine Learning Engineer - Advertising Engineering
Senior hybrid machine learning engineering role in Bengaluru focused on Roku's advertising data and AI platform. You'll build large-scale ML systems and agentic generative AI features for targeting, measurement, reporting, and optimization.
Machine Learning Engineer, Dubbing
Sarvam AI is hiring a mid-level Machine Learning Engineer in Bengaluru to build production systems for dubbing and live translation. The role requires speech and NLP experience, production ML skills and Python.
Machine Learning Engineer
Machine Learning Engineer role in Bengaluru for someone with 5–12 years of experience designing models, data pipelines, and production ML systems. The position also involves shaping the roadmap and delivering scalable AI applications.
Machine Learning Engineer
Nykaa is hiring a mid-level machine learning engineer in Bengaluru to productionize ML systems used in e-commerce. The role requires 5 to 7 years of experience along with Python, MLOps, model serving, monitoring, and AWS or cloud skills.