Head of Safety
Cognition is hiring a Head of Safety in San Francisco, United States. Level rates it ; you can apply on Level.
AI in this role
We are an applied AI lab building end-to-end software agents.
We're the makers of Devin, the first AI software engineer.
Our team is extremely talent-dense. Among our founding team, we have world-class competitive programmers, former founders, and leaders from companies at the cutting edge of AI including Scale AI, Palantir, Cursor, Waymo, Tesla, Lunchclub, Modal, Google DeepMind, and Nuro.
Building Devin is just the first step—our hardest challenges still lie ahead. If you’re excited to solve some of the world’s biggest problems and build AI that can reason on real-world tasks, apply to join us.
Role Mission
Devin is one of the most capable autonomous agents in production. It writes and runs code, uses tools, takes actions in real customer systems, and operates for hours without a human in the loop. That makes it one of the most important places in the industry to get safety right, and one of the few where safety work ships to millions of developers rather than staying in a paper. You will build Cognition's safety function from the ground up: the research agenda, the evaluation and red-teaming program, the deployment policies, and the team. You will work directly with the researchers training our models and the engineers building the agent harness. This is a hands-on role for someone who has done serious safety or alignment work at a frontier lab and wants to apply it where agents actually operate.
What You'll Accomplish
Own safety end to end: Set the safety strategy for Devin and the models behind it, and own the outcomes.
Build the evaluation program: Design evals and red-teaming for agentic risk: unsafe actions, prompt injection, data exfiltration, sandbox escape, reward hacking, and misuse. Make them part of every model and product release.
Shape training and the agent harness: Partner with pre-training, post-training, and agent teams so safety is built into how models are trained and how Devin plans, acts, and asks for help.
Set deployment policy: Define what Devin is and is not allowed to do, how permissions and oversight work, and how we handle the gray areas, in a way that holds up with enterprise customers.
Build the team and the external voice: Hire and lead safety researchers and engineers. Represent Cognition's safety work with customers, policymakers, and the broader research community.
Exceptional Candidates Have Demonstrated
Frontier lab safety or alignment experience: Hands-on work on safety, alignment, evaluations, or red-teaming at a frontier AI lab. You have shipped safety work that affected real models or products.
Agentic systems depth: You understand how agents fail in practice: tool misuse, specification gaming, long-horizon drift, adversarial inputs. You have ideas about how to measure and mitigate it.
Research credibility: A track record of published or widely used safety research, evals, or methods. An advanced degree in Computer Science, Machine Learning, or a related field is a plus.
Technical fluency: Comfortable in Python and in the training and inference stack. You can read the code, run the evals, and argue with researchers on the details.
Builder, not reviewer: You have started or scaled a safety function and prefer shipping mitigations to writing memos about them.
Judgment under uncertainty: You can make clear calls on deployment risk with incomplete information and explain them to engineers, executives, and customers.
Equal Opportunity
Cognition is an equal opportunity employer. We do not discriminate on the basis of race, color, religion, sex, sexual orientation, gender identity, national origin, age, disability, veteran status, or any other protected characteristic under applicable law. We are committed to providing reasonable accommodations for candidates with disabilities throughout the hiring process - please let us know if you need any.
How we rate this
Head of Safety at Cognition rates 88 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.
Builds AI. The job is building AI systems.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- How do you decide when an AI agent can act on its own versus asking for approval first?
- Walk me through how you've used Cursor in your day-to-day work.
- What are the limits of Devin that you've run into, and how did you work around them?
- How would you decide a model or AI system is ready to ship?
- Tell me about a time a model underperformed in production. How did you find out, and what did you change?
Adapt your resume
- List these exact terms on your resume: AI agents, Cursor, and Devin. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them. A tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped. This role is judged on the AI system itself, not the tools around it.
Want an expert to read your CV for this job?
Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.
Get new AI jobs (Builds AI ●●●●) by email
One email a week with the new AI jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Other roles that build AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step