# Senior Machine Learning Platform Engineer — Model Hosting & MLOps at HP

AI Level 4, AI centrality 90 out of 100. Spring, Texas, United States of America.

## Details

- Company: [HP](https://jobsbylevel.com/companies/hp)
- AI level: AI Level 4 (score 90 out of 100)
- Location: Spring, Texas, United States of America
- Salary: $147k-$231k
- Posted: October 6, 2026
- Apply: https://jobsbylevel.com/go/1dbb275e-f023-41eb-947c-b15f05a5b70a

## Description

Senior Machine Learning Platform Engineer — Model Hosting & MLOps Description - Role summary We are hiring a Senior Machine Learning Platform Engineer to build the infrastructure and engineering workflows that take custom models from development into reliable production use. This is a hands-on role for someone who has hosted LLMs on GPU infrastructure, developed the cloud platform around model serving, and built MLOps pipelines that make releases repeatable and observable. You will work with model developers, application engineers, security, and cloud platform teams to turn a model artifact into a secure, scalable inference service. The work spans serving architecture, infrastructure as code, deployment automation, model lifecycle management, and production operations across AWS and Azure, with potential integration into on-premises environments. You will also help shape agent workflows that use hosted models, so serving choices support the needs of multi-step applications. Success means teams can deploy, update, monitor, and troubleshoot models through clear, reusable platform patterns. What you will do Design and build hosting for custom ML and AI models across AWS and Azure, with particular focus on GPU-backed LLM inference and real-time endpoints; support batch inference where appropriate. Package models and their dependencies into reproducible serving workloads; choose and implement suitable managed services, containers, or Kubernetes-based patterns based on throughput, latency, security, reliability, and cost. Provision and configure Kubernetes clusters or other suitable hosting infrastructure, including compute, storage, API access, identity and access controls, secrets, networking integration, observability, and environment configuration. Help design secure connections and deployment patterns between cloud and on-premises environments as hosting needs evolve. Develop infrastructure as code and deployment automation so model-hosting environments can be provisioned, reviewed, promoted, and maintained consistently. Build MLOps workflows for model registration, versioning, validation, release, rollback, and retirement. Connect training or model preparation to deployment through automated pipelines and appropriate quality gates. Partner with application teams to design and prototype agent workflows, including model and tool orchestration, state handling, failure recovery, and evaluation. Translate agent workload patterns into hosting decisions about model selection, context length, concurrency, latency, cost, tool access, and end-to-end tracing. Establish production monitoring for service health, latency, throughput, errors, GPU and other resource use, and model behavior. Share responsibility for diagnosing incidents and improving capacity, reliability, and cost. Create reusable deployment templates, reference architectures, documentation, and onboarding paths that help other teams ship models safely. Partner with model and application teams on practical tradeoffs such as online versus batch inference, managed versus self-hosted serving, scaling, evaluation, data handling, and operational ownership. Required experience Hands-on experience hosting LLM inference on GPU infrastructure in a production environment. You can explain what you personally built, how models reached production, and how you managed throughput, latency, utilization, reliability, and cost. Experience building the surrounding model-serving platform for custom models, such as inference runtimes, deployment patterns, endpoint access, scaling, and operational tooling. Strong software engineering skills, especially Python, with experience building services, automation, and maintainable production code. Experience developing infrastructure for ML workloads across AWS and Azure, with deep hands-on delivery in at least one and practical ability to work in the other. You have used infrastructure as code such as Terraform or an equivalent tool. Experience

The description is cut here. Read the full offer: https://jobsbylevel.com/jobs/senior-machine-learning-platform-engineer-model-hosting-mlops-at-hp-cb171e

Source: https://jobsbylevel.com/jobs/senior-machine-learning-platform-engineer-model-hosting-mlops-at-hp-cb171e

## Cite this page

Level. https://jobsbylevel.com/jobs/senior-machine-learning-platform-engineer-model-hosting-mlops-at-hp-cb171e.

Get job alerts: https://jobsbylevel.com/newsletter
