Thinking Machines LabRemote · San Francisco$350k-$475k12h ago
EtchedPosted today
MTS - Agent Research at Etched scores 99 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.
AI in this role
About Etched
Etched is building hardware for frontier intelligence. We co-design chips, racks, software, and manufacturing to deliver best-in-class throughput and latency across both prefill and decode workloads. Our first products are heavily focused on inference. Backed by hundreds of millions from top-tier investors and staffed by leading engineers, Etched is redefining the infrastructure layer for the fastest growing industry in history.
Job Summary
As a research lead, you will set the research direction for this effort and spearhead the development of self-improving agents that reliably complete difficult, long-running engineering tasks. At Etched, this loop compounds: better agents help us build and optimize our systems faster, better systems expand our capacity to run more experiments, generating feedback and training data that further improve agent capabilities. Your job is to make this loop work in practice.
This requires strong research judgment. Where will better memory, context, or tooling unlock new capabilities? Which tasks benefit from structured workflows, and which require more open-ended autonomy? When do agent limitations call for changes to the harness, and when do they require custom models trained with supervised fine-tuning or reinforcement learning? You will set research priorities and design experiments that guide our investments in in-context learning, model training, and harness design.
You will work closely with engineers embedded across the company, with an initial focus on ASIC development, inference, and supercomputing. First, you will build the environments that let existing models do their best work through better tools, context, feedback, and in-context learning. Then, you will use the resulting trajectories and evaluations to develop specialized models where they can improve capability, reliability, or cost.
This is a hands-on research leadership role. You should be comfortable setting direction, writing and debugging code, and carrying an idea through to production. Research success means measurable improvements in the work Etched can accomplish, with disciplined use of compute.
Key Responsibilities
Own the research agenda for agentic systems across Etched. Identify the most important bottlenecks, prioritize promising approaches, and decide when to scale.
Work directly with domain engineers to understand real workflows, diagnose where agents fail, and turn those failures into research questions with concrete evaluations.
Develop systems for long-running tasks, including memory, context management, tool use, planning, and coordination between agents. Determine which approaches generalize across teams and which require domain-specific methods.
Build evaluations that measure correctness, reliability, and efficiency on real engineering tasks, and verify that gains transfer to production.
Train custom models using SFT, RL, and distillation, with the agent harness in the training loop. Use execution trajectories and failure analysis to guide training data and reward design.
Scale experiments to hundreds or thousands of concurrent agents while maintaining reproducibility, observability, and control over compute costs.
You may be a good fit if you have (Must-have qualifications)
Strong research taste and a record of turning ambiguous problems into clear hypotheses, decisive experiments, and useful systems.
Excellent engineering ability. You can move between research exploration, agent experimentation, low-level debugging, and production execution, and you take responsibility for systems working reliably.
Experience building agents for complex, multi-step tasks, with a practical understanding of context, memory, tools, evaluation, and failure recovery.
Hands-on experience with LLM post-training, including SFT and RL.
A track record of making meaningful progress under compute and data constraints. You care about the quality of the experiment and the cost of a successful outcome.
High agency and comfort with unfamiliar domains. You can find the problem, acquire the missing context, and drive the work forward.
The ability to set technical direction and work closely with domain experts to turn research into systems they trust and use.
Strong candidates may also have experience with (Nice-to-have qualifications)
ASIC design or verification, accelerator kernels, inference systems, or large-scale compute infrastructure.
Benefits
Medical, dental, and vision packages with generous premium coverage
$500 per month credit for waiving medical benefits
Housing subsidy of $2,500 per month for those living within walking distance of the office
Relocation support for those moving to San Jose (Santana Row)
Various wellness benefits covering fitness, mental health, and more
Daily lunch and dinner in our office
Unlimited compute budget subject to ROI justification
Base Compensation Range
$175K – $275K
Plus Significant Equity
How we’re different
Etched believes in the Bitter Lesson. We are the first inference-focused frontier AI system, betting early on transformer and transformer-like architectures and on increasing model sizes. Our addressable market is the entirety of inference, unlike many of our competitors.
We are a fully in-person team in San Jose (Santana Row), and greatly value engineering skills. We do not have boundaries between engineering and research, and we expect all of our technical staff to contribute to both and work across disciplines as needed.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- Walk me through fine-tuning a model: what data did you use, and how did you check the result?
- How would you decide a model or AI system is ready to ship?
- Tell me about a time a model underperformed in production. How did you find out, and what did you change?
Adapt your resume
- List these exact terms on your resume: Fine Tuning. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Similar roles
Research roles rated AI Level 4 at other companies.
TuringPalo Alto, California, United States; San Francisco, California, United States; Seattle, Washington, United States$250k-$400k12h ago
AmazonUS, MA, North Reading$159k-$215k1d
PerplexityBerlin1d
NVIDIAUS, CA, Santa Clara$38-$94/hr2d
MercorRemote · San Francisco$5000k2d
WaymoRemote · Mountain View, CA, USA; San Francisco, CA, USA; New York, NY, USA$213k-$263k2d
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step





