Principal Engineer, Cluster Orchestration
Skills
About the role
This is a Principal Engineer role in Bellevue, Washington or Sunnyvale, California, leading orchestration systems for large AI and GPU clusters. You'll set architecture and technical direction across Kubernetes, Slurm, SUNK, Kueue, and related control planes supporting training, inference, and model onboarding. The role suits a senior distributed-systems engineer with at least 15 years of experience and deep Kubernetes and Slurm expertise.
What you’ll do
- Define the long-term architecture for cluster orchestration platforms.
- Set technical direction for scheduling, quotas, fairness, pre-emption, and GPU isolation.
- Lead Kubernetes-native control-plane and operator development.
- Remove scaling limits across schedulers, registries, networking, storage, and control planes.
- Establish reliability, observability, SLO, alerting, and incident practices.
- Write and review production code for controllers, schedulers, and admission systems.
- Improve scheduling latency, startup time, image distribution, and cold-start performance.
- Lead design reviews and mentor senior and staff engineers.
What they’re looking for
- At least 15 years building and operating large-scale distributed systems.
- Deep practical knowledge of Kubernetes and Slurm internals.
- Experience operating GPU-heavy AI training, inference, or HPC platforms.
- Strong Go and cloud-native systems development skills.
- Ability to influence technical direction across teams without direct authority.
- Comfort making high-impact decisions in complex systems.
- A bachelor's or master's degree in a relevant field, or equivalent experience.
Nice to have
- Experience with Kueue, Kubeflow, Argo Workflows, Ray, Istio, or Knative.
- Background in ML platform engineering, model onboarding, or lifecycle management.
- Knowledge of scheduling, pre-emption, quota enforcement, and elastic scaling.
- A record of operating reliable systems with SLOs and incident processes.
- Contributions to Kubernetes, ML infrastructure, or related open-source projects.
- Experience mentoring senior engineers and raising engineering standards.
What’s on offer
- Onsite work in Bellevue, WA or Sunnyvale, CA.
- Base salary of $206,000 to $303,000 per year.
- Additional discretionary bonus, equity awards, and benefits subject to eligibility.
- Benefits include insurance, retirement matching, paid parental leave, flexible PTO, and other US-based programs.
- Work on infrastructure that determines GPU utilization, workload reliability, and AI deployment speed.
Questions about this role
Where is this Principal Engineer role based?
The role is listed in Bellevue, Washington and Sunnyvale, California.
How much experience is required?
The posting asks for 15 or more years of experience with large-scale distributed systems.
What is the base salary?
The base salary range is $206,000 to $303,000 per year.
Which orchestration systems are involved?
The role covers Kubernetes, Slurm, SUNK, Kueue, and related control planes.
Is a degree required?
A bachelor's or master's degree in a relevant field is listed, with equivalent experience accepted.
Related roles
Senior Manager, SOX-Business Process
Senior SOX business-process role in Bellevue for an accounting, finance or controls professional with at least 8 years of experience. The position leads control design, testing, remediation and financial-reporting compliance.
Senior Manager, Operations Communications
On-site Senior Manager role in Bellevue leading executive and internal communications for CoreWeave's operations organization.
Senior Manager, Operations Accounting - Fixed Assets
Senior fixed-asset accounting manager role in Dallas overseeing close, reporting, controls, audits, and a global data-center asset portfolio. The role suits an accounting leader with at least six years of experience and a professional accounting qualification.
Senior Manager, Operations Accounting Data Center Infrastructure
Senior onsite accounting leadership role in Dallas focused on financial governance for large-scale data center infrastructure construction. The role requires at least seven years of accounting experience and covers CapEx, CIP, fixed assets, SOX controls, and audit readiness.
Senior Manager, Joint Venture & VIE
CoreWeave is hiring a Senior Manager for Joint Venture and VIE Accounting in New York, Sunnyvale, or Dallas. The role requires a CPA and seven or more years of accounting experience, including hands-on ASC 810, ASC 323, and JV reporting work.
Senior Firmware Engineer, FPGA
Senior firmware and FPGA engineering role in New York or Sunnyvale focused on integrating platform logic with BMC, BIOS and server hardware. It suits an experienced RTL and low-level software engineer who can handle board bring-up and cross-layer debugging.