Member of Technical Staff, Evals Platform
AI in this role
About Us:
Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that's building the future of enterprise AI.
We are seeking a Member of Technical Staff, Evals & Post-Training Product to help define how developers improve models on Fireworks. This role sits at a unique intersection of scalable system design, deep data science, and model quality.
You will build the infrastructure and workflows that connect evaluation and post-training. Our evaluation setup has grown past its original scope, and we need someone who can take it to the next stage, improving programmatic access and scale. You will work across backend systems, sandbox infrastructure, and user-facing surfaces to make it easier to author evals, understand results, and iterate quickly.
Key Responsibilities
Scale Eval Infrastructure: Take ownership of our internal eval setup and evolve it for the future. You will design systems to eliminate single-host coordination bottlenecks, resolve log-syncing latency, and build seamless programmatic access.
Benchmark Obsession & Reproduction: Act as a "metrics obsessive." Track state-of-the-art (SOTA) benchmarks, read the latest research papers, dig deep into data discrepancies, and insist on rigorously reproducing published results internally.
Pioneer New Benchmarks: Design and build entirely new benchmarks to measure model performance on complex, emerging, or domain-specific use cases.
Own Fine-Tuning Product Experiences: Build and improve user-facing workflows for post-training, including fine-tuning experiences across SFT, RFT, and related model-improvement capabilities.
Work Closely With Users: Partner with customers and internal stakeholders to understand evaluation and fine-tuning needs, triage issues, and convert bespoke workflows into productized, reusable solutions.
Minimum Requirements
Experience: 1–7 years of software engineering or data science experience (We are hiring at multiple levels for this role).
Strong System Design Skills: You know how to architect scalable, programmatic systems and transition legacy setups into robust infrastructure.
Sandbox Infrastructure: Hands-on experience building or working with sandbox environments and sandbox infrastructure for secure code execution and testing.
Analytical & Data Science Mindset: You possess a deep understanding of LLM evaluations, how to design them, and how to use the results to guide model improvement. You are meticulous about data and metrics.
Understanding of the GenAI Lifecycle: You understand the end-to-end workflow—from prompting a base model to curating a dataset, fine-tuning, and productionizing agents.
Preferred Qualifications
Experience: 3+ years of software engineering or applied data science experience.
Frameworks & Orchestration: Experience working with the Harbor framework or similar container registry and orchestration tools.
Public Writing & Analysis: A strong interest in discovering where different models excel and fall short, with a desire to write up and publish these insights publicly (e.g., technical blogs, whitepapers).
Inference & Hardware Knowledge: Interest in the hardware side of AI—understanding GPU constraints, inference optimization techniques, and how they relate to model performance.
Startup DNA: Experience in fast-paced environments where you own features end-to-end.
Why Fireworks?
Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving.
Build What’s Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally.
Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results.
Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation.
Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators.
How we score this
Member of Technical Staff, Evals Platform at Fireworks AI scores 97 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.
AI Level 4. Building AI systems is the job itself: without AI, the role would not exist.
- AI Level 480 to 100
- AI Level 360 to 79
- AI Level 240 to 59
- AI Level 10 to 39
Bands come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- Walk me through fine-tuning a model: what data did you use, and how did you check the result?
- Walk me through how you've used PyTorch in your day-to-day work.
- How would you decide a model or AI system is ready to ship?
- Tell me about a time a model underperformed in production. How did you find out, and what did you change?
Adapt your resume
- List these exact terms on your resume: Fine Tuning and PyTorch. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Get new remote AI jobs at AI Level 4+ by email
One email a week with the new remote AI jobs at AI Level 4+, each rated AI Level 1 to 4 for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Other roles rated AI Level 4 at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step