Level

Artificial Analysis

Member of Technical Staff (Hardware)

AI in this role

openaianthropichugging-facevllmgroq

Job Description – Member of Technical Staff (Hardware)

Location: San Francisco (on-site at our offices)

About Artificial Analysis

Artificial Analysis is the leading independent AI benchmarking company. We support labs, engineers and enterprises to understand AI capabilities and make critical decisions about their AI strategies. We are the go-to authority for understanding AI, from AI labs and enterprises to media, investors, and policymakers. Our benchmarks don’t just measure the cutting edge of AI, they are actively shaping the frontier.

Our benchmarks and analysis are trusted by hundreds of thousands of users and are the go-to reference for leading AI labs including OpenAI, Google, Meta, NVIDIA and Anthropic, and major publications including the Wall Street Journal, Bloomberg, the Financial Times and The Economist.

We are a team of 40+, on track to double by end of year, backed by Nat Friedman (GitHub, Meta), Daniel Gross (SSI, Meta), Andrew Ng (Google Brain, DeepLearning.ai, Amazon), Adam D’Angelo (Quora, Poe, OpenAI), Clem Delangue (Hugging Face) and other industry leaders.

The Opportunity

AI hardware is where the next decade of AI economics will be decided, and our hardware benchmarks are becoming the reference point for how the industry measures accelerators. We’re hiring into our hardware pillar to drive that coverage: benchmarking the GPUs, TPUs and custom silicon the AI industry runs on, and building the analysis that helps the industry understand them.

You’ll design and run performance benchmarks across accelerators and inference configurations, extend our hardware benchmarking stack, including AA-AgentPerf and our performance benchmarking suite, build cost and throughput models, and work directly with the leading chipmakers to benchmark their latest silicon. This is a technical, hands-on role at the intersection of silicon and AI, working directly with our founders and hardware pillar lead.

What You’ll Do

• Benchmark AI Accelerators: Design and execute performance benchmarking across GPUs, TPUs and custom inference silicon, measuring throughput, latency and price-performance the way the industry actually deploys

• Build Evaluation Methodology: Develop and maintain the frameworks that define how AI hardware performance and efficiency are measured, extending products like AA-AgentPerf, from tokens per second and time-to-first-token to total cost of ownership

• Analyze Inference Economics: Deeply understand how cost-per-token, utilization and hardware choice interact, and translate that into analysis the industry relies on for deployment and procurement decisions

• Partner with Chipmakers: Work with the top hardware companies in the world, from NVIDIA, AMD and Google to the leading new accelerator companies, at a deep technical level: benchmarking their latest silicon, shaping methodology together, and setting the standards the industry measures by

• Drive Strategic Analysis: Produce reports and data visualizations that communicate hardware performance and economics to technical and non-technical audiences

• Become AI-Native: Embrace an AI-native workflow, using cutting-edge AI tools to generate leverage in a fast-changing industry and maintain our competitive edge in AI benchmarking

What We’re Looking For

You should come from the world of AI accelerators.

Backgrounds include: engineering, product, performance or technical roles at AI accelerator and inference hardware companies (e.g. Cerebras, Groq, SambaNova, d-Matrix, Etched, MatX, Tenstorrent or similar), GPU and accelerator teams at larger players (NVIDIA, AMD, Google, Amazon, Qualcomm), or inference infrastructure companies working close to the silicon.

Required:

• 3+ years of professional experience, including at least 2 years at an AI chip company (e.g. NVIDIA, Cerebras, SambaNova, Groq or similar)

• Strong analytical and critical thinking skills

• Proficiency in Python and data analysis

• Deep familiarity with AI accelerators and the inference software stack (e.g. CUDA, TensorRT, vLLM or similar serving frameworks)

• Understanding of inference economics: cost-per-token, utilization, and the trade-offs that drive real deployment decisions

• Genuine, demonstrable interest and knowledge of Frontier AI. We want people who have informed opinions about where AI is heading, not just people who use AI tools

Why Artificial Analysis?

• Shape how AI gets built: The leading AI labs track our benchmarks and use them to guide their development priorities. Your work will directly influence the direction of AI.

• Become a world expert in AI: You will evaluate every major model, across every major capability, as they are released. Very few roles offer this breadth of exposure to frontier AI.

• Work with the most important players in AI: You’ll manage relationships with teams at the leading AI labs and major enterprises as a trusted, independent voice.

• Join at a defining moment: We’re 40+ people, on track to double by end of year, backed by some of the most connected investors in AI. The people who join now will shape the product, the team, and the strategy as we scale.

• Competitive compensation including equity

1

How we score this

Member of Technical Staff (Hardware) at Artificial Analysis scores 89 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.

Classification

AI Level 4. Building AI systems is the job itself: without AI, the role would not exist.

  1. AI Level 480 to 100
  2. AI Level 360 to 79
  3. AI Level 240 to 59
  4. AI Level 10 to 39

Bands come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

OpenAIAnthropicHugging FacevLLMGroq

Questions you could be asked

  1. What's a project where you used OpenAI hands-on?
  2. Walk me through how you've used Anthropic in your day-to-day work.
  3. What are the limits of Hugging Face that you've run into, and how did you work around them?
  4. What's a project where you used vLLM hands-on?
  5. Walk me through how you've used Groq in your day-to-day work.

Adapt your resume

  • List these exact terms on your resume: OpenAI, Anthropic, Hugging Face, vLLM, and Groq. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want your resume actually rewritten for this job?

The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.

Get new AI jobs at AI Level 4+ by email

One email a week with the new AI jobs at AI Level 4+, each rated AI Level 1 to 4 for how much AI is in the work. No recruiter spam, unsubscribe in one click.

Free. One email a week. Unsubscribe in one click.

Similar roles

Other roles rated AI Level 4 at other companies.

What kind of AI work fits you?

Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.

Find my next step

More jobs at Artificial Analysis

Related searches

Same AI level