# Director, Software Engineering – Cloud Compute and Infrastructure at NVIDIA

NVIDIA is hiring a Director, Software Engineering – Cloud Compute and Infrastructure in Santa Clara, United States. It pays $320k-$489k a year and Level rates it Little AI ●○○○; you can [apply on Level](https://jobsbylevel.com/go/18287ef6-8ee2-4fd7-bbdb-d762b9946ef4).

AI Level 1, AI centrality 34 out of 100. US, CA, Santa Clara.

## Details

- Company: [NVIDIA](https://jobsbylevel.com/companies/nvidia)
- AI level: AI Level 1 (score 34 out of 100)
- Location: US, CA, Santa Clara
- Salary: $320k-$489k
- Posted: October 9, 2026
- Apply: https://jobsbylevel.com/go/18287ef6-8ee2-4fd7-bbdb-d762b9946ef4

## Description

NVIDIA is seeking an engineering director to lead the software teams behind our bare-metal datacenters and cloud compute infrastructure. The organization supports hundreds of megawatts of datacenter capacity already online, with more capacity coming. You will build the software that makes this growing physical infrastructure available as reliable, secure, and efficient compute services for NVIDIA engineering. Our team provides NVIDIA’s continuous integration (CI) infrastructure: the environment where hardware, firmware, drivers, networking, and system software are integrated and tested on their path to production. You will enable engineering teams to bring up preproduction systems, reproduce failures, validate changes, and move platforms toward production readiness. The fleet spans multiple hardware generations and maturity levels: x86 and Arm servers, GPUs, DPUs, high-speed networking, storage, and rack-scale, liquid-cooled systems. Platforms such as Grace Blackwell and Vera Rubin illustrate the breadth of compute and interconnect technology involved. This role combines hands-on systems judgment with leadership of the teams making that diversity manageable at scale. What you’ll be doing: Lead software engineering teams responsible for cloud compute, bare-metal fleet management, and CI infrastructure, owning architecture, implementation, deployment, and operations. Set the technical direction for compute control planes, resource provisioning, topology-aware placement, reservations, and capacity management across physical servers, virtual machines, and containers. Automate the bare-metal lifecycle: hardware discovery and inventory, server bring-up, imaging, firmware and driver configuration, out-of-band management, health checks, reprovisioning, and recovery. Build CI services that allocate the right hardware, provision repeatable test environments, record hardware and software configurations, collect diagnostics, and restore systems to a known state. Shorten the path from a hardware or software change to actionable validation results. Keep engineering services dependable while supporting evolving preproduction hardware and software. Define service-level objectives, isolate failures, improve observability, and turn incidents and recurring test-infrastructure failures into engineering fixes. Measure and improve hardware onboarding time, provisioning speed, CI queue time, productive fleet utilization, recovery time, and cost efficiency as capacity expands. Set engineering standards for design and code reviews, automated testing, secure development, release quality, and safe changes to production systems. Hire, coach, and retain engineers and engineering managers. Establish clear ownership, develop technical leaders, and build teams that deliver consistently over multiple release cycles. Partner with hardware, firmware, drivers, networking, storage, security, and validation teams to onboard platforms and diagnose failures across system boundaries. Translate engineering users’ needs into clear priorities and technical decisions. Work with datacenter operations and facilities teams to bring additional capacity online. Connect rack power, cooling, physical topology, and hardware-health telemetry to provisioning, placement, serviceability, and operational readiness. What we need to see: 15+ overall years of experience in software engineering, distributed systems, or cloud infrastructure, including 7+ years leading engineering teams and experience managing engineering managers. Direct engineering ownership of the underlying services of a public or private compute cloud, such as AWS EC2, Google Compute Engine, Azure Compute, OCI Compute, or a comparable infrastructure-as-a-service platform. Strong technical depth in distributed systems, Linux, virtualization, containers, and the networking and storage services that support large compute fleets. Experience building software and automation for bare-metal infrastructure, including server

The description is cut here. Read the full offer: https://jobsbylevel.com/jobs/director-software-engineering-cloud-compute-and-infrastructure-at-nvidia-ac7083

Source: https://jobsbylevel.com/jobs/director-software-engineering-cloud-compute-and-infrastructure-at-nvidia-ac7083

## Cite this page

Level. https://jobsbylevel.com/jobs/director-software-engineering-cloud-compute-and-infrastructure-at-nvidia-ac7083.

Get job alerts: https://jobsbylevel.com/newsletter
