The role
You own reliability for 75+ petaflops of client GPU fleets across our three Ashburn halls — from BMC firmware to Slurm queues. When a training run degrades at 2am, you're the person the alert finds, and the person empowered to fix it.
Responsibilities
- Fleet health ownershipThermal, power and node-level monitoring with 60-second alerting.
- Burn-in & bring-upValidate every new GPU delivery before it touches client workloads.
- Automation firstKill toil with tooling: provisioning, firmware rollouts, capacity checks.
Requirements
5+ years running Linux fleets in production; comfort with BMC/IPMI, PXE provisioning and one config-management system; GPU or HPC exposure strongly preferred; US work authorization and ability to badge on site in Ashburn.
Apply for this role
We reply to every application within two weeks.