Principal Engineer - Perf and Benchmarking
Skills
About the role
This is a principal engineering leadership role in Sunnyvale, California, or Bellevue, Washington, for someone with at least ten years of distributed-systems, HPC, or cloud-services experience. You'll lead CoreWeave's performance and benchmarking function, covering a global performance data warehouse, MLPerf submissions, Kubernetes-based benchmarking, and GPU workload analysis. The role combines technical strategy, team leadership, systems architecture, developer tooling, and external collaboration across the AI infrastructure ecosystem.
What you’ll do
- Set the long-term benchmarking roadmap across models, workloads, and hardware tiers.
- Lead and mentor performance engineers and data analysts.
- Establish reproducible, versioned, and auditable benchmarking methods.
- Own MLPerf Training and Inference submissions from planning through publication.
- Coordinate optimization work across CUDA, TensorRT, Triton, NCCL, and related systems.
- Design a Kubernetes-native benchmarking service across SUNK, Kueue, and Kubeflow.
- Measure latency, throughput, jitter, token rates, startup behavior, and cost.
- Automate comparisons across software releases, hardware generations, precisions, and batch sizes.
- Build CI/CD pipelines and Kubernetes controllers for large-scale benchmark scheduling.
- Integrate Prometheus, Grafana, OpenTelemetry, and results warehouses.
- Secure benchmark artifacts with SBOMs and Cosign signatures.
- Work with NVIDIA, ISVs, and open-source projects on optimizations and upstream fixes.
What they’re looking for
- At least ten years building distributed systems, HPC platforms, or cloud services.
- Deep experience with large-scale machine learning training or comparable high-performance workloads.
- Track record designing planet-scale telemetry, observability, warehouse, or OLAP systems.
- Strong understanding of CUDA, NCCL, RDMA, NVLink, PCIe, and GPU memory bandwidth.
- Experience with Triton, vLLM, TensorRT-LLM, TorchServe, or similar model-serving stacks.
- Knowledge of PyTorch FSDP, DeepSpeed, Megatron-LM, or distributed training systems.
- Production experience with Kubernetes and ML control planes.
- Excellent communication with executives, customers, auditors, and open-source communities.
- Ability to lead teams and define technical strategy.
Nice to have
- Time-series database, LSM-tree, or custom storage-engine experience.
- Experience running MLPerf or other audited benchmarks at scale.
- Contributions to MLPerf, Triton, vLLM, PyTorch, KServe, or similar projects.
- Experience benchmarking multi-region fleets and clusters with thousands of GPUs.
- Publications or talks on ML performance, latency engineering, or benchmarking methodology.
What’s on offer
- Base salary of $206,000 to $333,000 per year.
- Discretionary bonus and equity awards.
- Medical, dental, and vision insurance paid by CoreWeave for eligible US employees.
- Life, disability, flexible spending, and health savings benefits.
- Tuition reimbursement and employee stock purchase program participation.
- Mental wellness, family-forming, parental leave, and childcare support.
- 401(k) with employer match and flexible paid time off.
- Catered meals at offices and data centers.
Questions about this role
Where is this principal engineer role based?
The listed locations are Sunnyvale, California, and Bellevue, Washington, with no remote arrangement stated.
How much does the role pay?
The base salary range is $206,000 to $333,000 per year, with bonus and equity also included in the total rewards package.
How much experience is required?
The posting asks for at least ten years building distributed systems, HPC, or cloud services.
What does the benchmarking team work on?
It covers MLPerf submissions, large-scale performance data, Kubernetes-native benchmarks, GPU workloads, and performance analysis across models and hardware.
Related roles
Senior Manager, SOX-Business Process
Senior SOX business-process role in Bellevue for an accounting, finance or controls professional with at least 8 years of experience. The position leads control design, testing, remediation and financial-reporting compliance.
Senior Manager, Operations Communications
On-site Senior Manager role in Bellevue leading executive and internal communications for CoreWeave's operations organization.
Senior Manager, Operations Accounting - Fixed Assets
Senior fixed-asset accounting manager role in Dallas overseeing close, reporting, controls, audits, and a global data-center asset portfolio. The role suits an accounting leader with at least six years of experience and a professional accounting qualification.
Senior Manager, Operations Accounting Data Center Infrastructure
Senior onsite accounting leadership role in Dallas focused on financial governance for large-scale data center infrastructure construction. The role requires at least seven years of accounting experience and covers CapEx, CIP, fixed assets, SOX controls, and audit readiness.
Senior Manager, Joint Venture & VIE
CoreWeave is hiring a Senior Manager for Joint Venture and VIE Accounting in New York, Sunnyvale, or Dallas. The role requires a CPA and seven or more years of accounting experience, including hands-on ASC 810, ASC 323, and JV reporting work.
Senior Firmware Engineer, FPGA
Senior firmware and FPGA engineering role in New York or Sunnyvale focused on integrating platform logic with BMC, BIOS and server hardware. It suits an experienced RTL and low-level software engineer who can handle board bring-up and cross-layer debugging.