Home  /  Careers  /  SRE

Senior Site Reliability Engineer

Operations · Ashburn, VA · Full-time · $165–195k + uptime bonus

The role

You own reliability for 75+ petaflops of client GPU fleets across our three Ashburn halls — from BMC firmware to Slurm queues. When a training run degrades at 2am, you're the person the alert finds, and the person empowered to fix it.

Responsibilities

  • Fleet health ownershipThermal, power and node-level monitoring with 60-second alerting.
  • Burn-in & bring-upValidate every new GPU delivery before it touches client workloads.
  • Automation firstKill toil with tooling: provisioning, firmware rollouts, capacity checks.

Requirements

5+ years running Linux fleets in production; comfort with BMC/IPMI, PXE provisioning and one config-management system; GPU or HPC exposure strongly preferred; US work authorization and ability to badge on site in Ashburn.

Apply for this role

We reply to every application within two weeks.