Level

Anduril Industries

Staff AI Test and Evaluation Engineer

Anduril Industries is hiring a Staff AI Test and Evaluation Engineer . It pays $253k-$336k a year and Level rates it ; you can apply on Level.

AI in this role

ml-opsai-evaluationcomputer-vision

Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century’s most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril’s family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years.

ABOUT THE TEAM

Discovery is Anduril's team for taking the newest problems across domains- space, missile systems, air, sensor capability, autonomy, and cyber- and proving out what is worth solving. We build the models, run the tests, and carry what works to the point where a program can pick it up. Discovery works alongside Air Defense, Space, Intelligence, Cyber, GNC, Hardware, and every other group at Anduril to incubate the solutions to the hardest problems.

ABOUT THE JOB

Discovery is building an expeditionary force: a team of engineers who want to solve difficult problems and who navigate unfamiliar territory as an operating standard. As a Discovery engineer, you pick up a concept nobody has proven and build the analysis or prototype that tests it to get to an answer. Whether you spend your time working on hard problems on our autonomy stack, acoustic sensor analysis, or battle space management, we tackle every challenge the same way: by questioning assumptions and building our way to the answer.

We are focused on taking AI models from research into production—onto classified platforms and edge hardware where they have to perform reliably in the real world. As our model portfolio and classified work grow, rigorous, repeatable evaluation of how these models actually perform has become mission-critical.

WHAT YOU'LL DO

  • Develop Test Scenarios Simulation Environments: Develop comprehensive test scenarios, and simulation environments to assess agentic AI performance in classified simulations.
  • Validate AI on Classified & Edge Platforms: Lead the integration and validation of agentic AI systems onto classified platforms and edge hardware.
  • Build Reusable Evaluation Pipelines: Build reusable evaluation pipelines, automated test harnesses, and monitoring dashboards for continuous validation.
  • Define Actionable Metrics: Figure out what good metrics look like for AI models and traditional models—establishing systematic, historic capture of performance rather than one-shot, deployment-specific measurement.
  • Partner Cross-Functionally: Partner with cross-functional teams to define requirements, document test results, and contribute to AI model cards and deployment readiness reviews.
  • Own End-to-End Evaluation: Architect, build, and maintain the evaluation and validation infrastructure for our agentic AI and ML models, from unit-level model checks through full-system, scenario-based assessment.
  • Troubleshoot and Debug: Analyze and resolve issues uncovered in evaluation and in deployment, ensuring reliability and operational success across every release.

REQUIRED QUALIFICATIONS

  • Technical Expertise: Bachelor's or Master's degree in Computer Science, Machine Learning, Software Engineering, Computer Engineering, Electrical Engineering, or a related technical field.
  • Programming Proficiency: At least 12+ years of hands-on experience writing production-grade code, with strong Python skills for building evaluation pipelines, test harnesses, and tooling.
  • Evaluation Ownership Mindset: Demonstrated experience owning evaluation or automated testing for production ML or software systems—test scenarios, metrics, regression suites, and CI infrastructure.
  • AI/ML Evaluation Experience: Hands-on experience designing evaluation methodologies for AI or machine learning models—defining metrics, building benchmarks, and assessing model behavior against real-world scenarios.
  • Simulation Experience: Experience designing or working extensively with simulation environments to exercise model and system behavior.
  • Metrics Discipline: Working knowledge of how to define, capture, and reason about performance metrics for AI models and traditional models, including systematic historic capture over time.
  • Systems-Level Thinking: Ability to navigate and contribute to complex systems and established codebases.
  • Program Ownership: Comfort operating between technical program management and software engineering—defining requirements, coordinating across teams, and documenting results.
  • Real-World Impact: Passion for building the evaluation infrastructure that proves AI works—directly influencing mission-critical outcomes.
  • Security Clearance: Must be eligible for a US security clearance.

PREFERRED QUALIFICATIONS

  • Prior Title Background: Experience as an ML Test & Evaluation Engineer, SDET for ML systems, ML/Evaluation Engineer, Simulation Engineer, or Technical Program Manager for AI/ML.
  • Agentic AI Evaluation: Experience designing test methodologies for agentic AI systems—tasking, decision-making, and scenario-based behavior validation.
  • Classified / Edge Deployment: Familiarity with validating and deploying AI models onto classified platforms, edge hardware, or resource-constrained environments.
  • Model Cards & Readiness Reviews: Experience contributing to AI model cards, deployment readiness reviews, or similar model-governance and release-gating practices.
  • Aerospace/Defense T&E: Familiarity with test and evaluation practices in aerospace or defense, including qualification testing, range operations, or operational assessment.
  • MLOps & Monitoring: Experience with MLOps tooling, monitoring dashboards, and continuous validation pipelines for ML models in production.
  • Programming Skills: Additional experience with Go, C++, or scripting for test automation and tooling.
  • Growth into Broader Ownership: Interest in growing from T&E ownership into deeper ownership of the AI evaluation and deployment stack over time.
US Salary Range$253,000—$336,000 USD

The salary range for this role is an estimate based on a wide range of compensation factors, inclusive of base salary only. Actual salary offer may vary based on (but not limited to) work experience, education and/or training, critical skills, and/or business considerations. Highly competitive equity grants are included in the majority of full time offers; and are considered part of Anduril's total compensation package. Additionally, Anduril offers top-tier benefits for full-time employees, including: 

 

Benefits

At Anduril, we invest in our people. Our comprehensive, competitive benefits package (available at little to no cost to employees) ensures you’re supported in health, recovery, and whatever comes next. For more information, Explore Our Benefits.

 

Protecting Yourself from Recruitment Scams

Anduril is committed to maintaining the integrity of our Talent acquisition process and the security of our candidates. We've observed a rise in sophisticated phishing and fraudulent schemes where individuals impersonate Anduril representatives, luring job seekers with false interviews or job offers. These scammers often attempt to extract payment or sensitive personal information.

To ensure your safety and help you navigate your job search with confidence, please keep the following critical points in mind:

  • No Financial Requests: Anduril will never solicit payment or demand personal financial details (such as banking information, credit card numbers, or social security numbers) at any stage of our hiring process. Our legitimate recruitment is entirely free for candidates.

  • Please always verify communications:
    • Direct from Anduril: If you receive an email from one of our recruiters, it will only come from an @anduril.com address.
    • Via Agency Partner: If contacted by a recruiting agency for an Anduril role, their email will clearly identify their agency. If you suspect any suspicious activity, please verify the agency's authenticity by reaching out to contact@anduril.com. 
  • Exercise Caution with Unsolicited Outreach: If you receive any communication that appears suspicious, contains grammatical errors, or makes unusual requests, do not engage. Always confirm the sender's email domain is @anduril.com before providing any personal information or clicking on links.

  • What to Do If You Suspect Fraud: Should you encounter any questionable or fraudulent outreach claiming to be from Anduril, please report it immediately to contact@anduril.com. Your proactive caution is invaluable in protecting your personal information and upholding the security and trustworthiness of our recruitment efforts.

 

Data Privacy

To view Anduril's candidate data privacy policy, please visit https://anduril.com/applicant-privacy-notice/. 

 

By submitting your application, you consent to Anduril Industries using a third-party service provider to conduct pre-employment risk, integrity, and due diligence screening and assessing potential risks as part of your application process. This third-party service provider provides risk-intelligence services that may include analysis of sanctions and watchlists, adverse media, public-record information, and other lawful open-source or commercial data sources. This third-party service provider does not act as a consumer reporting agency. Use of this provider helps to ensure compliance with applicable laws and protect technology, intellectual property, and organizational security.

How we rate this

Staff AI Test and Evaluation Engineer at Anduril Industries rates 92 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.

Classification

Builds AI. The job is building AI systems.

  1. ●●●● Builds AI80 to 100
  2. ●●●○ Works on AI60 to 79
  3. ●●○○ Uses AI40 to 59
  4. ●○○○ Little AI0 to 39

Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

ML OpsAI EvaluationComputer vision

Questions you could be asked

  1. How do you monitor a model once it's live, and how do you know it needs retraining?
  2. How do you decide that one model's output is better than another's for a given task?
  3. Walk me through a computer vision problem you solved, from raw data to a deployed model.
  4. How would you decide a model or AI system is ready to ship?
  5. Tell me about a time a model underperformed in production. How did you find out, and what did you change?

Adapt your resume

  • List these exact terms on your resume: ML Ops, AI Evaluation, and Computer vision. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
  • Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.

Want an expert to read your CV for this job?

Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.

Get new software engineering jobs (Builds AI ●●●●) by email

One email a week with the new software engineering jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.

Free. One email a week. Unsubscribe in one click.

Similar roles

Software Engineering roles that build AI, at other companies.

Version 1

London, Birmingham, Manchester, Newcastle upon Tyne, Edinburgh, Belfast, England, United Kingdom1h ago

What kind of AI work fits you?

Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.

Find my next step

More jobs at Anduril Industries

Related searches

Same AI level