Compute Engineer, Deployment
- Discuss this with your agent
- Open in Claude
- Open in ChatGPT
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Role details
Tech stack
Job description
Own compute turn-up from facility availability to ready-for-service: the stretch after the network hands off and before customers run workloads. Qualify racks at scale: establish firmware baselines, configure BMC and BIOS, run burn-in, and validate at node and cluster level across hundreds of racks per site on GPU and custom accelerator platforms. Drive qualification through the base-management Kubernetes platform and provisioning stack (discovery, imaging, firmware updates, shared services), burning down qualification queues with tooling rather than manual processes. Triage hardware failures found in qualification: isolate to component, drive RMA and vendor escalation, and feed failure patterns back into qualification gates. Run turn-up remotely by default, with on-site deployments of roughly a week per data hall as new halls reach facility availability, plus occasional overlapping-site engagements. Partner with network deployment, ICT, data center operations, and hardware teams during turn-up windows, and support incident response on newly deployed capacity. Ability to travel 20-30% of the time to data centers and lab environments as needed.
Requirements
Experience bringing up server or GPU fleets at scale, hundreds of nodes or more, and taking them all the way to production. Strong Linux background with hands-on experience using out-of-band management technologies such as BMC, IPMI, and Redfish. Experience automating hardware workflows using Python or Go rather than relying on manual processes. Hands-on data center experience including racking, cabling, hardware installation, troubleshooting, and component replacement. Ability to methodically troubleshoot issues across hardware, firmware, and software layers, isolating root causes before implementing fixes. Willingness and ability to travel during deployment and turn-up activities.
Nice to Have Skills & Experience
Kubernetes-based bare metal provisioning. Accelerator platform bring-up and validation (NVIDIA, AMD, or custom hardware). Burn-in and stress-testing framework design. DCIM and inventory management tooling experience.
Benefits & conditions
Benefit packages for this role will start on the 1st day of employment and include medical, dental, and vision insurance, as well as HSA, FSA, and DCFSA account options, and 401k retirement account access with employer matching. Employees in this role are also entitled to paid sick leave and/or other paid time off as provided by applicable law.
About the company
Our client is a leading AI infrastructure company building and operating large-scale compute environments that power next-generation AI workloads. Their teams design, deploy, and operate high-performance data center infrastructure at massive scale, with a focus on speed, reliability, and operational excellence.
Apply for this position
This job is hosted externally. Click below to view the full posting and apply.
Prepare application
- Draft this with your agent
- Open in Claude
- Open in ChatGPT
Good distractions
Talks and stories from around this role — technically off-topic, practically not.
Moments
Explore playlistsVideos
See allRelated articles
See all
7 Cloud Computing Trends Coming in 2025 for Developers
Highest Paying Tech Companies for Developers
Dev Digest 120 - Apple and peers
Best US AI Conferences for CTOs in 2026: Build vs. Buy, Vendor Evaluation, and Peer Intelligence