Engineer, Supercomputing & Distributed Systems
AI in this role
About Krea
At Krea, we are building next-generation AI creative tools.
We're dedicated to making AI intuitive and controllable for creatives - our mission is to build tools that empower human creativity, not replace it. We believe AI is a new medium that allows us to express ourselves through various formats - text, images, video, sound, and even 3D. We're building better, smarter, and more controllable tools to harness this medium. We recently took this a step forward with the launch of Krea 2, our first foundation model, built completely from scratch for aesthetic diversity and stylistic control.
We've raised over $83M and are backed by world-class investors such as a16z, Bain Capital, and Abstract. We work full-time and in-person at our waterfront office in San Francisco. We care about creativity: our team includes musicians, designers, visual artists, and engineers.
Supercomputing / AI Infra at Krea
We build and operate the infrastructure for Krea's research and inference. Distributed training, 1000+ K8s GPU clusters, petabyte scale data pipelines, etc. We build a lot of this from scratch — custom distributed datastores, job orchestration systems, and streaming pipelines that replace tools like Kafka and Ray for modern AI workloads at scale.
Example projects:
Distributed data systems
Design multi-stage pipelines that turn petabytes of raw data into clean, annotated datasets
Run classification models on billions of images
Deploy and combine LLMs to caption massive multimedia data
GPU infrastructure
Manage distributed training and inference on 1000+ GPU Kubernetes clusters
Solve orchestration and scaling for large-scale GPU job processing
Scale workloads and research between clusters in multiple datacenters
Distributed training
Profile and optimize dataloaders streaming thousands of images per second
Profile and debug InfiniBand networking on huge training runs
Build fault tolerance systems for large-scale pretraining
Collaborate with researchers on evolving RL infrastructure
Applied ML pipelines
Find clean scenes in millions of videos using distributed shot-boundary detection
Customize and train models to filter billions of images for questions like "is this a screenshot?"
Build the systems that bridge raw cluster capacity and research output
Who we're looking for
Systems people. If you've read a blog post about InfiniBand debugging or building a custom distributed database and thought "I want to do that" — this is that team.
You'll spend your time working heavily with Python, Kubernetes, Torch, and data tools like DuckDB, Arrow, etc. It's OK if you don't have K8s or ML experience — the main thing we hire for is an intuition for distributed systems, and a great mental model of how systems interact and function under different conditions.
Strong candidates may have experience with…
Python, PyArrow, DuckDB, SQL, massive relational databases, PyTorch, Pandas, NumPy…
Kubernetes
Designing and implementing large-scale ETL systems
Fundamental knowledge of containerization, operating systems, file-systems, and networking
Distributed systems design
Distributed training systems (NCCL, InfiniBand, RDMA)
Streaming and event processing systems (Kafka, Pulsar, or similar)
PyTorch internals, custom dataloaders, and training infrastructure
What we offer
Team: Work alongside a world-class team building the future of AI creative tooling
Impact: Significant scope and company-wide impact
Competitive compensation: generous salary & equity packages
Health & wellness: 100% health & 99% dental/vision insurance premiums covered for employees, health FSA accounts, & long-term disability coverage
Time off: Flexible PTO policy
Financial planning: 401k with a 4% company-sponsored match
Meals in the office: breakfast, lunch, dinner - you name it, we'll cover it
Transit: Ubers covered to & from the office
Sponsorship: We're open to sponsoring international visas where we can (e.g., STEM OPT, OPT, H-1B, O-1, E-3).
And more!
Please note the above benefits & perks are for full-time employees
How we score this
Engineer, Supercomputing & Distributed Systems at Krea scores 96 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.
AI Level 4. Building AI systems is the job itself: without AI, the role would not exist.
- AI Level 480 to 100
- AI Level 360 to 79
- AI Level 240 to 59
- AI Level 10 to 39
Bands come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- What's a project where you used PyTorch hands-on?
- How would you decide a model or AI system is ready to ship?
- Tell me about a time a model underperformed in production. How did you find out, and what did you change?
Adapt your resume
- List these exact terms on your resume: PyTorch. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Get new AI jobs at AI Level 4+ by email
One email a week with the new AI jobs at AI Level 4+, each rated AI Level 1 to 4 for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Software Engineering roles rated AI Level 4 at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step