# Staff AI Test and Evaluation Engineer  at Anduril Industries

Anduril Industries is hiring a Staff AI Test and Evaluation Engineer . It pays $253k-$336k a year and Level rates it Builds AI ●●●●; you can [apply on Level](https://jobsbylevel.com/go/e59dd6cb-a361-44ba-b308-60f7c8990f59).

AI Level 4, AI centrality 92 out of 100. Washington, District of Columbia, United States.

## Details

- Company: [Anduril Industries](https://jobsbylevel.com/companies/anduril-industries)
- AI level: AI Level 4 (score 92 out of 100)
- Location: Washington, District of Columbia, United States
- Salary: $253k-$336k
- Posted: October 8, 2026
- Apply: https://jobsbylevel.com/go/e59dd6cb-a361-44ba-b308-60f7c8990f59

## Description

Anduril Industries is a defense technology company with a mission to transform U.S. and allied military capabilities with advanced technology. By bringing the expertise, technology, and business model of the 21st century’s most innovative companies to the defense industry, Anduril is changing how military systems are designed, built and sold. Anduril’s family of systems is powered by Lattice OS, an AI-powered operating system that turns thousands of data streams into a realtime, 3D command and control center. As the world enters an era of strategic competition, Anduril is committed to bringing cutting-edge autonomy, AI, computer vision, sensor fusion, and networking technology to the military in months, not years. ABOUT THE TEAM Discovery is Anduril's team for taking the newest problems across domains- space, missile systems, air, sensor capability, autonomy, and cyber- and proving out what is worth solving. We build the models, run the tests, and carry what works to the point where a program can pick it up. Discovery works alongside Air Defense, Space, Intelligence, Cyber, GNC, Hardware, and every other group at Anduril to incubate the solutions to the hardest problems. ABOUT THE JOB Discovery is building an expeditionary force: a team of engineers who want to solve difficult problems and who navigate unfamiliar territory as an operating standard. As a Discovery engineer, you pick up a concept nobody has proven and build the analysis or prototype that tests it to get to an answer. Whether you spend your time working on hard problems on our autonomy stack, acoustic sensor analysis, or battle space management, we tackle every challenge the same way: by questioning assumptions and building our way to the answer. We are focused on taking AI models from research into production—onto classified platforms and edge hardware where they have to perform reliably in the real world. As our model portfolio and classified work grow, rigorous, repeatable evaluation of how these models actually perform has become mission-critical. WHAT YOU'LL DO Develop Test Scenarios Simulation Environments: Develop comprehensive test scenarios, and simulation environments to assess agentic AI performance in classified simulations. Validate AI on Classified & Edge Platforms: Lead the integration and validation of agentic AI systems onto classified platforms and edge hardware. Build Reusable Evaluation Pipelines: Build reusable evaluation pipelines, automated test harnesses, and monitoring dashboards for continuous validation. Define Actionable Metrics: Figure out what good metrics look like for AI models and traditional models—establishing systematic, historic capture of performance rather than one-shot, deployment-specific measurement. Partner Cross-Functionally: Partner with cross-functional teams to define requirements, document test results, and contribute to AI model cards and deployment readiness reviews. Own End-to-End Evaluation: Architect, build, and maintain the evaluation and validation infrastructure for our agentic AI and ML models, from unit-level model checks through full-system, scenario-based assessment. Troubleshoot and Debug: Analyze and resolve issues uncovered in evaluation and in deployment, ensuring reliability and operational success across every release. REQUIRED QUALIFICATIONS Technical Expertise: Bachelor's or Master's degree in Computer Science, Machine Learning, Software Engineering, Computer Engineering, Electrical Engineering, or a related technical field. Programming Proficiency: At least 12+ years of hands-on experience writing production-grade code, with strong Python skills for building evaluation pipelines, test harnesses, and tooling. Evaluation Ownership Mindset: Demonstrated experience owning evaluation or automated testing for production ML or software systems—test scenarios, metrics, regression suites, and CI infrastructure. AI/ML Evaluation Experience: Hands-on experience designing evaluation methodologies for AI or

The description is cut here. Read the full offer: https://jobsbylevel.com/jobs/staff-ai-test-and-evaluation-engineer-at-anduril-industries-edd151

Source: https://jobsbylevel.com/jobs/staff-ai-test-and-evaluation-engineer-at-anduril-industries-edd151

## Cite this page

Level. https://jobsbylevel.com/jobs/staff-ai-test-and-evaluation-engineer-at-anduril-industries-edd151.

Get job alerts: https://jobsbylevel.com/newsletter
