MLOps / AI Operations Engineer
AI in this role
Job Description & Summary
The opportunity
Industrialize AI delivery through automated deployment, evaluation operations, observability, reliability engineering and transparent consumption management.
What you will be doing
· Build CI/CD pipelines for AI services, prompts, agent configurations, infrastructure and evaluation assets.
· Automate environment provisioning, testing, deployment, rollback and release evidence.
· Implement tracing, logging, model and agent monitoring, alerts and operational dashboards.
· Operationalize evaluation thresholds, incident handling and continuous-improvement loops.
· Monitor latency, capacity, token usage, infrastructure consumption and cost drivers.
· Define runbooks, service ownership and production support handover.
What we need from you
· 4+ years in DevOps, platform engineering, ML engineering, SRE or cloud operations.
· Strong automation, containers, cloud services, observability and Infrastructure as Code capability.
· Experience deploying or operating ML, generative AI or distributed application workloads.
· Understanding of release controls, reliability, security and cost optimization.
Relevant AI technologies and tooling
· Hands-on experience with GitHub Actions, Azure DevOps, GitLab CI or equivalent, plus Infrastructure as Code using Terraform, Bicep or comparable tooling.
· Strong container and orchestration capability using Docker and Kubernetes, together with experience deploying AI or agent services across cloud and hybrid environments.
· Experience operating model and prompt assets, agent configurations, evaluation datasets and release evidence using MLflow, platform-native registries or equivalent lifecycle tooling.
· Practical implementation of agent tracing and observability using OpenTelemetry and tools such as LangSmith, MLflow, Langfuse, Azure Monitor, Prometheus or Grafana.
· Ability to monitor model and agent quality, tool failures, retrieval performance, latency, token usage, cost, capacity and workflow-level service indicators.
· Experience with progressive delivery, rollback, secrets management, vulnerability scanning, incident response and reliability practices for non-deterministic AI systems.
Measures of success
· Deployment frequency and success rate
· Mean time to detect and restore
· Evaluation and monitoring coverage
· Service reliability and latency
· Cost and consumption transparency
Key interfaces
· Other members of the AI Transformation & Agentic Systems Practice
· PwC sector, functional, cloud, cyber, risk, Responsible AI and change specialists
· Client business owners, product owners, technology teams and operational users
· Technology alliance and implementation partners where relevant
Contribution to the practice
· Support proposals, client workshops and market development appropriate to seniority.
· Contribute reusable methods, patterns, code, assets and lessons learned.
· Coach colleagues and participate in the capability’s continuous learning agenda.
· Uphold PwC quality, independence, confidentiality and risk-management requirements.
#LI-BS1 #LI-Hybrid
How we score this
MLOps / AI Operations Engineer at PwC scores 10 out of 100 on AI centrality, which makes it AI Level 1 of 4 (Little AI) on this board. The level measures how much of the work is AI, not seniority.
AI Level 1. The work itself involves no AI, or AI only appears as scenery, such as a company tagline.
- AI Level 480 to 100
- AI Level 360 to 79
- AI Level 240 to 59
- AI Level 10 to 39
Bands come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- How do you monitor a model once it's live, and how do you know it needs retraining?
- How do you think about the risk of an AI system in this kind of role failing silently?
- What are the limits of Mlflow that you've run into, and how did you work around them?
Adapt your resume
- List these exact terms on your resume: Ml Ops, AI Safety, and Mlflow. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Get new AI jobs by email
One email a week with the new AI jobs, each rated AI Level 1 to 4 for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Operations roles rated AI Level 1 at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step