Principal AI/ML Engineer
AI in this role
Key Responsibilities:
AI Architecture & Technical Leadership
- Define and lead the technical architecture for enterprise-scale AI and ML platforms.
- Design scalable, resilient, and reusable AI systems capable of supporting mission-critical workloads.
- Establish architectural standards, engineering patterns, and best practices for AI deployment and operations.
- Drive technical decisions around model serving, inference optimization, agent architectures, orchestration frameworks, observability, and AI infrastructure.
Productize AI Research
- Partner closely with AI researchers to transform cutting-edge prototypes into production-grade solutions.
- Lead efforts to operationalize advanced AI capabilities across areas such as:
- Large Language Models (LLMs)
- Trustworthy and Responsible AI
- Agentic AI Systems
- Establish repeatable pathways that accelerate innovation-to-production cycles.
- Ensure production solutions maintain scientific rigor while meeting enterprise engineering standards.
- Bridge the gap between research breakthroughs and sustainable business value.
Engineering Excellence & Scalability
- Solve the organization's most complex AI engineering and scalability challenges.
- Design systems that operate reliably at enterprise scale while balancing performance, latency, governance, security, and cost.
- Drive adoption of MLOps, LLMOps, and AI platform engineering best practices.
- Improve the robustness, maintainability, observability, and operational readiness of our AI products.
- Identify and eliminate architectural bottlenecks that impact scale, reliability, or client experience.
- Raise standards through coaching, architecture reviews, design guidance, and technical leadership.
Production Reliability & Operational Leadership
- Own the operational excellence, reliability, performance and availability of our products.
- Lead technical response and resolution efforts for complex production incidents, performance degradation, model failures, and system outages.
- Serve as the senior technical escalation point for the team's most challenging production challenges.
- Establish best practices for AI system monitoring, observability, alerting, incident management, capacity planning, and service-level objectives (SLOs).
- Mentor and lead junior engineers in troubleshooting, root cause analysis, operational decision-making, and incident response.
- Drive post-incident reviews focused on learning, continuous improvement, and long-term corrective actions.
- Develop operational processes that ensure AI solutions remain secure, scalable, performant, and reliable for business-critical use cases.
- Partner with product, infrastructure, security, and support teams to proactively identify operational risks and continuously improve service reliability.
Mentorship & Thought Leadership
- Mentor AI and ML engineers within the team.
- Foster a culture of technical excellence and operational ownership where engineers are accountable not only for building systems, but also for running and supporting them successfully in production.
- Represent our team as a thought leader in scalable AI deployment, operational excellence, and responsible AI practices.
Required Qualifications
- 10+ years of experience in software engineering, machine learning engineering, AI engineering, or related technical disciplines.
- Deep expertise designing, deploying, and supporting large-scale AI and ML systems in production environments.
- Demonstrated success leading complex technical initiatives from concept through deployment and ongoing operations.
- Strong knowledge of software architecture, reliability engineering, observability, ML Ops, DevOps, and cloud technologies.
- Proven ability to mentor engineers and lead teams through highly complex technical and operational challenges.
Preferred Qualifications
- Experience with foundation models, Large Language Models, and agentic AI architectures.
- Experience deploying agentic AI systems and multi-agent workflows.
- Experience with Trustworthy AI, Responsible AI, AI governance, or model risk management frameworks.
- Experience optimizing large-scale inference systems and AI infrastructure.
- Experience working in highly regulated environments and mission-critical production systems.
What You'll Gain
This role offers a unique opportunity to operate at the forefront of applied artificial intelligence and help bridge world-class research with real-world impact.
You will:
- Work directly with world-class AI researchers on breakthrough technologies and next-generation AI capabilities.
- Own a critical position in the pipeline that transforms cutting-edge research into client value.
- Tackle some of the most difficult AI engineering, scalability, and operational challenges in the industry.
- Build AI capabilities that deliver meaningful business outcomes for clients.
- Develop deep expertise in operating advanced AI systems at scale while collaborating with leaders across research, product, and engineering.
Special Factors
Sponsorship
Vanguard is not offering visa sponsorship for this position.About Vanguard
At Vanguard, we don't just have a mission—we're on a mission.
To work for the long-term financial wellbeing of our clients. To lead through product and services that transform our clients' lives. To learn and develop our skills as individuals and as a team. From Malvern to Melbourne, our mission drives us forward and inspires us to be our best.
How We Work
Vanguard has implemented a hybrid working model for the majority of our crew members, designed to capture the benefits of enhanced flexibility while enabling in-person learning, collaboration, and connection. We believe our mission-driven and highly collaborative culture is a critical enabler to support long-term client outcomes and enrich the employee experience.
How we rate this
Principal AI/ML Engineer at Vanguard rates 95 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.
Builds AI. The job is building AI systems.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- How do you monitor a model once it's live, and how do you know it needs retraining?
- How do you think about the risk of an AI system in this kind of role failing silently?
- Tell me about a research question you investigated. What did you find?
- How would you decide a model or AI system is ready to ship?
- Tell me about a time a model underperformed in production. How did you find out, and what did you change?
Adapt your resume
- List these exact terms on your resume: ML Ops, AI Safety, and AI Research. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want an expert to read your CV for this job?
Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.
Get a free CV reviewGet new machine learning engineer jobs (Builds AI ●●●●) by email
One email a week with the new machine learning engineer jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Data roles that build AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step