Mistral AISingapore5h ago
UPSPosted 12mo ago
GCP Infrastructure Engineer - Google Cloud, Terraform, Python, Bash, GKE, CI/CD at UPS scores 96 out of 100 on AI centrality, which makes it AI Level 4 of 4 (Builds AI) on this board. The level measures how much of the work is AI, not seniority.
AI in this role
Before you apply to a job, select your language preference from the options available at the top right of this page.
Explore your next opportunity at a Fortune Global 500 organization. Envision innovative possibilities, experience our rewarding culture, and work with talented teams that help you become better every day. We know what it takes to lead UPS into tomorrow—people with a unique combination of skill + passion. If you have the qualities and drive to lead yourself or teams, there are roles ready to cultivate your skills and take you to the next level.
Job Description:
Job Summary:
We are seeking a highly skilled GCP Infrastructure Engineer to design, build, and manage the cloud infrastructure that powers Generative AI (GenAI) applications at scale. In this role, you will leverage Google Cloud Platform (GCP) Vertex AI, IBM Watsonx, and containerization technologies such as Docker and Kubernetes (GKE) to deliver secure, scalable, and high-performance AI solutions. You will own the end-to-end infrastructure lifecycle — from design and provisioning to automation, monitoring, and optimization — while enabling data scientists and ML engineers to seamlessly deploy and operate GenAI workloads.
Key Responsibilities:
Cloud Infrastructure & Platform Engineering
Design, provision, and maintain scalable, secure, and cost-efficient infrastructure for GenAI applications on GCP.
Deploy and manage containerized workloads using Docker and Kubernetes (GKE).
Configure and optimize Vertex AI and IBM Watsonx platforms for training, fine-tuning, and serving LLMs and other generative models.
Implement high-performance GPU/TPU clusters to support distributed training and large-scale inference.
Ensure business continuity through backup, disaster recovery, and multi-region deployments.
Automation & Reliability
Develop and maintain Infrastructure as Code (IaC) templates with Terraform, or Cloud Deployment Manager.
Adopt GitOps practices (Flux) for infrastructure lifecycle management.
Build and optimize CI/CD pipelines for data pipelines, model workflows, and GenAI applications.
Apply SRE principles (SLIs, SLOs, SLAs) to guarantee platform reliability and uptime.
Security, Governance & Compliance
Embed DevSecOps best practices across the infrastructure lifecycle, including policy-as-code, vulnerability scanning, and secrets management.
Enforce identity and access management (IAM), network segmentation, and data encryption in compliance with standards (HIPAA, SOX, GDPR, FedRAMP).
Collaborate with enterprise security and compliance teams to implement governance frameworks for GenAI platforms.
Monitoring, Observability & Cost Optimization
Implement observability stacks (Prometheus, Grafana, Cloud Monitoring, Datadog) for both infra health and ML-specific metrics (model drift, data anomalies).
Define KPIs to monitor system health, performance, and adoption across AI workloads.
Optimize cloud cost efficiency for GPU/TPU-intensive workloads using autoscaling, preemptible instances, and utilization monitoring.
Collaboration & Enablement
Partner with data scientists, ML engineers, and software teams to streamline GenAI application development and deployment.
Provide onboarding, documentation, and reusable templates to enable faster adoption of AI infrastructure.
Stay current with the latest advancements in GenAI, cloud-native infrastructure, and container orchestration.
Required Education
Bachelor’s or master’s degree in computer science, Software Engineering, or a related field.
Required Experience
8+ years of experience in cloud infrastructure engineering, DevOps, or platform engineering.
Experience with GenAI use cases (chatbots, content generation, code assistants, etc.).
Strong hands-on expertise with Google Cloud Platform (GCP), especially Vertex AI.
Experience with IBM Watsonx for AI application deployment and management.
Proven skills in Docker, Kubernetes (GKE), and container orchestration at scale.
Proficiency in Python, Bash, or other relevant scripting languages.
Strong understanding of cloud networking, IAM, and security best practices.
Experience with CI/CD tools (GitHub Actions, GitLab CI, Jenkins) and IaC tools (Terraform, Pulumi, Ansible, Deployment Manager).
Familiarity with data pipelines and integration tools (Dataflow, Apache Beam, Pub/Sub, Kafka).
Excellent problem-solving, debugging, and communication skills.
Preferred Experience
Experience in MLOps practices for model deployment, monitoring, and retraining.
Exposure to multi-cloud or hybrid cloud environments (GCP, AWS, Azure, on-prem).
Hands-on experience with feature stores (Vertex AI Feature Store, Feast) and ML observability tools (EvidentlyAI, Fiddler).
Knowledge of distributed training frameworks (Horovod, DeepSpeed, PyTorch Distributed).
Contributions to open-source projects in infrastructure, MLOps, or GenAI.
Experience managing infrastructure in regulated industries.
Preferred Certifications:
Google Cloud Certified - Professional Cloud Architect
Google Cloud Certified - Machine Learning Engineer
Certified Kubernetes Administrator (CKA) or Certified Kubernetes Application Developer (CKAD)
IBM Certified Watsonx Generative AI Engineer – Associate
IBM Certified Solution Architect - Cloud Pak for Data
Other relevant certifications in AI, Machine Learning, or Cloud-Native technologies.
Employee Type:
UPS is committed to providing a workplace free of discrimination, harassment, and retaliation.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- Walk me through fine-tuning a model: what data did you use, and how did you check the result?
- How do you monitor a model once it's live, and how do you know it needs retraining?
- What are the limits of Vertex AI that you've run into, and how did you work around them?
- What's a project where you used PyTorch hands-on?
- How would you decide a model or AI system is ready to ship?
Adapt your resume
- List these exact terms on your resume: Fine Tuning, Ml Ops, Vertex AI, and PyTorch. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped — this role is judged on the AI system itself, not the tools around it.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Get new AI jobs at AI Level 4+ by email
One email a week with the new AI jobs at AI Level 4+, each rated AI Level 1 to 4 for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Software Engineering roles rated AI Level 4 at other companies.
Epic GamesVancouver,British Columbia,CanadaCAD 274k-CAD 402k5h ago
SmartsheetRemote · Bangalore, INDIA13h ago
OpenAIRemote · San Francisco$266k-$445k17h ago
Anduril IndustriesWaltham, Massachusetts, United States$191k-$253k18h ago
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step





