OpenAIRemote · San Francisco$180k-$260k1h ago
ScalewayPosted 4w ago
Head of Engineering GPU at Scaleway scores 65 out of 100 on AI centrality, which makes it AI Level 3 of 4 (Works on AI) on this board. The level measures how much of the work is AI, not seniority.
AI in this role
Lead GPU cloud support, HPC, and SRE teams to operate large-scale AI and HPC infrastructure at Scaleway.
WHY WE NEED YOU ?
As our GPU Cloud business continues to scale, we are strengthening our engineering leadership to support the deployment and operation of increasingly large and complex GPU clusters.
Your mission will be to lead our Support Engineering, HPC, and SRE teams, own key technology and architecture decisions, and ensure that our most strategic GPU infrastructure projects are successfully designed, delivered, and operated.
You will also play a critical role in validating the technical feasibility of commercial proposals and ensuring that the commitments we make to customers can be delivered reliably on our sovereign cloud infrastructure.
YOUR FUTURE TEAM
We work in a collaborative and international environment where the diversity of Scalers, combined with a strong culture of knowledge sharing, helps us bring ambitious projects to life.
You will lead an organization of 14 engineers across two squads, each managed by an Engineering Manager reporting directly to you.
As part of the broader GPU Cloud organization, reporting to the SVP GPU Cloud, you will work closely with GTM, Operations, Service Management, Product, and other engineering teams to build and operate large-scale AI and HPC infrastructure.
YOUR DAILY ROUTINE
Tasks
- Lead the Support Engineering, HPC, and SRE organizations, directly managing two Engineering Managers responsible for 14 engineers
- Own the technical strategy, architecture, and key technology choices for GPU Cloud infrastructure
- Review and validate the technical and service dimensions of strategic commercial proposals, ensuring commitments are realistic and deliverable
- Oversee the design, deployment, and operational readiness of new GPU clusters
- Drive the evolution of our cluster management, capacity management, automation, and operational capabilities
- Ensure the reliability, scalability, performance, and maintainability of our GPU infrastructure
- Provide technical leadership on complex AI and HPC infrastructure projects
- Build strong alignment between Engineering, GTM, Product, and Operations
- Develop the engineering organization through clear direction, effective delegation, coaching, and long-term team development
- Establish and maintain high standards of engineering rigor, operational excellence, and technical decision-making
ABOUT YOU
HARDSKILLS:
- 10+ years of experience in infrastructure engineering, including significant experience leading senior technical teams and managers
- Proven experience designing, deploying, or operating large-scale infrastructure and compute clusters
- Familiarity in GPU and/or HPC environments, ideally involving NVIDIA and AMD technologies
- Strong understanding of distributed infrastructure, cluster architecture, reliability, and production operations
- Experience with orchestration, provisioning, and observability technologies such as Kubernetes, Proxmox, Warewulf, Prometheus, and Grafana
- Experience with high-performance and distributed storage technologies such as Lustre, DDN, and VAST Data
- Ability to assess architectural trade-offs and translate complex technical constraints into clear engineering and business decisions
SOFT SKILLS:
- Strong leadership skills with the ability to lead experienced engineers and Engineering Managers
- Ability to navigate complex technical, organizational, and business situations
- High level of rigor and a strong sense of ownership
- Excellent organizational, prioritization, and planning skills
- Strong communication and stakeholder-management abilities
- Ability to synthesize complex engineering topics and communicate them clearly to both technical and non-technical stakeholders
- Comfortable making decisions in fast-moving environments with high technical and operational stakes
WHAT YOU WILL FIND AT SCALEWAY
- Hybrid work: We offer up to 3 days of remote work per week.
-
Offices: Our offices are spacious, dynamic workspaces with bold design, conveniently located near public transport. Most of our offices feature outdoor spaces (terraces) and bike parking facilities.
-
Dining: Our chef provides a healthy meal service at the headquarters, and breakfast is available across all our sites year-round. Scalers working from regional sites enjoy a Swile card for lunches.
-
Well-being commitments: Whether it’s access to a gym, daycare places, or discounted services for caring services, Scaleway is committed to supporting Scalers in maintaining a balanced life.
-
International environment: With dozens of nationalities, Scaleway offers a stimulating environment where English is as widely spoken as French.
-
Career & Mobility: Our managers value internal mobility, and opportunities to transition to other entities within the Iliad Group are accessible to all Scalers.
🚀 Why join the Scaleway adventure?
✔ A rich and diverse product offering: Scaleway offers over 100 public cloud products in IaaS, PaaS, and AI.
✔ A cutting-edge technical environment: Scaleway provides modern infrastructures, including high-performance bare metal servers, to tackle exciting technical challenges.
✔ Commitment to responsible cloud: Scaleway is dedicated to a more responsible cloud, with data centers powered solely by renewable energy since 2017, minimizing our ecological footprint and holding top-level certification.
🔜 THE NEXT STEPS …
-
HR discovery call (30 min)
-
Interview with to understand your technical skills and approach to the role (45 min)
-
Technical interview with CTO to validate your expertise (1h)
-
Interview with Management to deepen discussions and assess your fit with the team (45 min)
-
HR interview to tour our offices and meet your future colleagues
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- Tell me about a project where hpc was part of your work. What did you do?
- Tell me about a project where infrastructure management was part of your work. What did you do?
- Tell me about a project where leadership was part of your work. What did you do?
- What's a project where you used Gpu hands-on?
- Walk me through how you've used Sre in your day-to-day work.
Adapt your resume
- List these exact terms on your resume: Hpc, Infrastructure Management, Leadership, Gpu, and Sre. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Show where AI is part of your daily process, not a one-off project — this role expects it to be a running habit.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Similar roles
Software Engineering roles rated AI Level 3 at other companies.
Gong ioAustin | Chicago | New York City | Salt Lake City | San Francisco$142k-$205k4h ago
Solve IntelligenceNew York$100k-$220k5h ago
n8nRemote · Europe5h ago
SmartsheetRemote · -REMOTE, USA-$161k-$194k6h ago
CoreWeaveLivingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA$157k-$231k7h ago
ToastDublin, Ireland8h ago
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step







