ML Researcher - Image / Video Diffusion
AI in this role
About Krea
At Krea, we are building next-generation AI creative tools.
We're dedicated to making AI intuitive and controllable for creatives - our mission is to build tools that empower human creativity, not replace it. We believe AI is a new medium that allows us to express ourselves through various formats - text, images, video, sound, and even 3D. We're building better, smarter, and more controllable tools to harness this medium. We recently took this a step forward with the launch of Krea 2, our first foundation model, built completely from scratch for aesthetic diversity and stylistic control.
We've raised over $83M and are backed by world-class investors such as a16z, Bain Capital, and Abstract. We work full-time and in-person at our waterfront office in San Francisco. We care about creativity: our team includes musicians, designers, visual artists, and engineers.
We're looking for an experienced Researcher with engineering skills who can work on large-scale image and video models training experiments, with experience training image models at scale.
Our culture
We work full-time and in-person at our North Beach office in San Francisco.
We believe that demonstrated interest in the creative space is key: our team includes musicians, designers, visual artists and more.
Fast iteration and execution speed. Bias towards action, agency, and independence.
What you'll do
Train diffusion models for image and video generation on large GPU clusters.
Fully optimize and profile large distributed training runs across model architectures, kernels, data loading, memory constraints, and communication.
Implement and improve various distributed training strategies including FSDP, CP, SP, TP, and EP.
Continuously improve model quality and reliability through data, model architecture, training pipeline, structuring experiments, and eval design.
Debug distributed training errors and implement fault tolerance solutions, identifying bad GPU, NVLink, Infiniband (IB) components as well as monitoring numerical errors and NCCL issues.
Ablate different architecture, attention, optimizer, data, and algorithmic choices to reliably improve efficiency and performance of our models.
What we're looking for
Proven track record in working with image or video models at scale (publications or open-source contributions a plus).
Strong proficiency in PyTorch and understanding of its inner workings.
Strong background in distributed training paradigms such as FSDP, CP, SP, USP, TP, and EP. Knowing how different parallelism strategies work together and their tradeoffs.
Experience in profiling and debugging large distributed training. Being comfortable with analyzing traces to identify bottlenecks and look for improvements.
Good knowledge of low precision training / inference in FP8, NVFP4, and MXFP8.
Solid understanding of diffusion model training pipeline across pretraining, midtraining, preference optimization, and reinforcement learning.
Keeping up with the developments in related fields such as LLM, VLM, representation learning, and robotics research.
Being comfortable working in a goal-oriented research environment.
Having good judgement around when one should explore different training strategies and when it's time to commit to a specific strategy to scale compute and data.
Comfortable working with underspecified goals. We expect every technical member to take an ambiguous research goal and break it down into concrete requirements, plans, experiment plan, and execution items.
Good research taste — bias towards simplicity and methods that scale well with compute, data, and minimal human supervision.
Ability to iterate rapidly, and propose creative research directions.
Be comfortable getting your hands dirty with data and designing custom data pipelines to improve data quality.
What we offer
Team: Work alongside a world-class team building the future of AI creative tooling
Impact: Significant scope and company-wide impact
Competitive compensation: generous salary & equity packages
Health & wellness: 100% health & 99% dental/vision insurance premiums covered for employees, health FSA accounts, & long-term disability coverage
Time off: Flexible PTO policy
Financial planning: 401k with a 4% company-sponsored match
Meals in the office: breakfast, lunch, dinner - you name it, we'll cover it
Transit: Ubers covered to & from the office
Sponsorship: We're open to sponsoring international visas where we can (e.g., STEM OPT, OPT, H-1B, O-1, E-3).
And more!
Please note the above benefits & perks are for full-time employees
How we score this
ML Researcher - Image / Video Diffusion at Krea scores 94 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.
AI Level 4. Building AI systems is the job itself: without AI, the role would not exist.
- AI Level 480 to 100
- AI Level 360 to 79
- AI Level 240 to 59
- AI Level 10 to 39
Bands come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- What's a project where you used PyTorch hands-on?
- How would you decide a model or AI system is ready to ship?
- Tell me about a time a model underperformed in production. How did you find out, and what did you change?
Adapt your resume
- List these exact terms on your resume: PyTorch. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Get new AI jobs at AI Level 4+ by email
One email a week with the new AI jobs at AI Level 4+, each rated AI Level 1 to 4 for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Research roles rated AI Level 4 at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step