# Cluster Datacenter Manager at Cerebras

Cerebras is hiring a Cluster Datacenter Manager. Level rates it Little AI ●○○○; you can [apply on Level](https://jobsbylevel.com/go/fa6df63c-41b5-4f84-9e8e-b92611bbad9b).

AI Level 1, AI centrality 36 out of 100. United States | Alabama.

## Details

- Company: [Cerebras](https://jobsbylevel.com/companies/cerebras)
- AI level: AI Level 1 (score 36 out of 100)
- Location: United States | Alabama
- Posted: October 8, 2026
- Apply: https://jobsbylevel.com/go/fa6df63c-41b5-4f84-9e8e-b92611bbad9b

## Description

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. This architecture allows Cerebras to deliver industry-leading training and inference speeds; over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation. Cerebras works with the leading model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras, to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. The Role Cerebras is seeking a hands-on Datacenter Cluster Manager to lead compute infrastructure operations and operational reliability across multiple data center sites. This leader manages technicians, engineers and contractors responsible for the health, availability, physical maintenance and lifecycle of high-performance compute and network infrastructure. The role requires strong technical judgment in server systems, networking, fiber connectivity and secure asset handling, alongside a working understanding of power, cooling and colocation dependencies. The Cluster Manager leads incident response, drives operational discipline, and ensures new deployments transition safely into sustained production support. Responsibilities Compute Systems Operations and Reliability Own day-to-day operational health and service readiness of compute clusters, server racks, supporting network infrastructure and associated physical systems across assigned sites. Lead hardware incident triage, fault isolation, break/fix, field-replaceable unit (FRU) replacement, diagnostic validation and escalation to systems engineering or vendors. Use monitoring, telemetry, logs and alerting to recognize degraded systems, prioritize impact, coordinate recovery and verify restoration of service. Oversee preventive maintenance, firmware or hardware change execution where authorized, spare parts readiness, and repeat-failure analysis. Partner with cluster ops , network and reliability engineering teams on root cause analysis, corrective actions and recurring fleet health issues. Network Infrastructure and Fiber Troubleshooting Apply practical understanding of data center network architecture, rack-level connectivity, switching, management networks and network redundancy to guide onsite diagnosis and escalation. Lead physical-layer troubleshooting of fiber and copper links, including patching, optics and transceivers, polarity, cleanliness, labeling and continuity testing using appropriate tools. Ensure structured cabling, fiber routing, documentation and change controls are maintained to engineering standards. Coordinate with network engineering to isolate physical connectivity faults versus configuration or software issues; execute approved remediation without assuming ownership of network design. Secure Media, Assets and Hardware Lifecycle Own physical asset accountability from inbound receiving, inspection, staging and inventory through deployment, repair, return and disposition. Enforce secure handling of data-bearing devices, including authorization, chain of custody, access controls, approved sanitization or destruction workflows and documented transfer. Ensure adherence to customer-specific controls, including two-person authorization and clean-in/clean-out procedures where applicable. Maintain accurate rack elevations, asset records, serial numbers, spares, repair histories and inventory reconciliation. People Leadership and Cluster Execution Lead, coach and evaluate technicians, engineers and contractor resources across multiple facilities; establish clear ownership, training and performance expectations. Plan staffing, shift coverage, on-call rotations and escalation readiness based on cluster demand, production service commitments and deployment schedules.

The description is cut here. Read the full offer: https://jobsbylevel.com/jobs/cluster-datacenter-manager-at-cerebras-d96ee5

Source: https://jobsbylevel.com/jobs/cluster-datacenter-manager-at-cerebras-d96ee5

## Cite this page

Level. https://jobsbylevel.com/jobs/cluster-datacenter-manager-at-cerebras-d96ee5.

Get job alerts: https://jobsbylevel.com/newsletter
