Level

TzafonPosted 11mo ago

Member of Technical Staff - Foundations

Member of Technical Staff - Foundations at Tzafon scores 100 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.

Remote (San Francisco / Tel Aviv / Zurich)leadFullTime$200k-$500k

AI in this role

openaianthropicvllmpytorchjax
fine-tuning

Tzafon is a foundation model lab building scalable compute systems and advancing machine intelligence, with offices in San Francisco, Zurich & Tel Aviv. We’ve raised over $12m in funding to advance our mission of expanding the frontiers of machine intelligence.

We're a team of engineers and scientists with deep backgrounds in ML infrastructure & research. Founded by IOI and IMO medalists, PhDs, and alumni from leading tech companies, such as Google Deepmind, Character, and NVIDIA, we train models and build infrastructure for swarms of agents to automate work across real-world environments.

You'll work between our product and post-training teams to ship Large Action Models that actually work. Build evals, benchmarks, and fine-tuning pipelines. Define what good model behavior means and make it happen at scale.

What you'll do

  • Design and execute large scale training runs on our clusters

  • Build and optimize distributed training infrastructure across massive multi-node systems

  • Implement post-training pipelines at scale

  • Develop data pipelines that process and filter trillions of tokens for pre-training

  • Research and implement architectural improvements, scaling laws, and training optimizations

  • Debug training instabilities, loss spikes, and convergence issues in long-running jobs

  • Build tooling for cluster utilization, fault tolerance, and checkpoint management

  • Write custom CUDA/Triton kernels to optimize critical training operations (attention, normalization, activations)

  • Collaborate on research that advances the state of the art in foundation model training


We're looking for

  • Deep experience pre-training or post-training foundation models on large clusters

  • Expert-level at Python and ML frameworks (PyTorch, JAX, Torchtitan)

  • Strong systems skills: distributed training, FSDP/ZeRO, tensor parallelism, pipeline parallelism

  • Experience writing performant CUDA or Triton kernels for ML workloads

  • Track record of running stable multi-week training jobs and debugging distributed training failures

  • Understanding of cluster scheduling, networking bottlenecks, and GPU/TPU performance optimization

Preferred Experience

  • Trained foundation models at major AI labs (OpenAI, Anthropic, Google DeepMind, Meta, xAI, etc.)

  • Worked on large scale RL runs

  • Optimized critical training kernels (FlashAttention, fused optimizers, custom kernels)

  • Published research at top ML conferences (NeurIPS, ICML, ICLR)

  • Contributions to open source ML infrastructure (PyTorch, JAX, vLLM, etc.)

  • Experience with training data pipelines, data quality research, or synthetic data generation

Life at Tzafon

  • Full medical, dental, and vision coverage, plus 401(k) in the us

  • Office in SF, Zurich, and Tel Aviv

  • Early-stage equity in a future-defining company

Visa sponsorship: We do sponsor visas! However, we aren't able to successfully sponsor visas for every role and every candidate. But if we make you an offer, we will make every reasonable effort to get you a visa, and we retain an immigration lawyer to help with this.

Compensation starts at $200k-$500k + equity package, depending on experience & location.

We also offer a referral bonus of $5k for referral of successful hires (send to [email protected]).

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

Fine TuningOpenAIAnthropicvLLMPyTorchJax

Questions you could be asked

  1. Walk me through fine-tuning a model: what data did you use, and how did you check the result?
  2. Walk me through how you've used OpenAI in your day-to-day work.
  3. What are the limits of Anthropic that you've run into, and how did you work around them?
  4. What's a project where you used vLLM hands-on?
  5. Walk me through how you've used PyTorch in your day-to-day work.

Adapt your resume

  • List these exact terms on your resume: Fine Tuning, OpenAI, Anthropic, vLLM, and PyTorch. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want your resume actually rewritten for this job?

The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.

Similar roles

Other roles rated AI Level 4 at other companies.

More jobs at Tzafon

Related searches

Same AI level