Site Reliability Engineer, Provider Operations
OpenRouter is hiring a Site Reliability Engineer, Provider Operations for a remote role open to applicants in United States. Level rates it ; you can apply on Level.
AI in this role
Site Reliability Engineer to own the operational health, monitoring, and failover infrastructure for AI inference model providers.
About OpenRouter
OpenRouter is the leading AI routing and infrastructure layer that enterprises use to access, manage, and optimize the best large language models across providers—without lock-in, capacity constraints, or unnecessary cost. We power the most advanced AI teams in the world by giving them the flexibility to move fast, scale confidently, and stay future-proof as models evolve.
As enterprise adoption of AI accelerates, OpenRouter sits at the center of how organizations operationalize LLMs across research, product, and production workloads.
About the Role
OpenRouter routes almost a billion requests and more than 20 trillion tokens a day, across 80+ providers and thousands of endpoints. Every one of those providers can degrade, rate-limit, change behavior, or go down without warning. Our customers count on us to absorb that chaos so their apps never notice.
We're hiring our first AI Inference SRE to own the operational health of our provider supply. You'll make sure every endpoint we route to is fast, correct, and available, and that we detect and route around problems before customers do. You'll sit on the Provider Operations team, reporting to the Provider Operations Manager.
What You'll Do
Provider health and observability. Build and own monitoring for every provider and endpoint: latency, throughput, error rates, uptime, and output correctness. Set SLOs per provider tier and alert on them.
Detection and failover. Improve how quickly we detect degraded endpoints, and work with the routing team so traffic shifts away from them automatically.
Incident response. Own on-call for provider incidents: triage, mitigate, communicate with providers, run postmortems, and drive follow-ups to closure.
Provider accountability. Turn telemetry into scorecards and SLO reporting that providers act on, and be the technical escalation point when a provider's endpoint is misbehaving.
Quality regression detection. Build continuous canaries and evals that catch silent regressions (quantization changes, broken tool calling, truncated streams, pricing or usage-reporting mismatches), not just outright downtime.
Automate the toil. Replace manual provider-ops work (disabling endpoints, capacity changes, deprecations, rate-limit tuning) with safe, auditable tooling.
Capacity and launch readiness. Build tooling to load-test endpoints before big launches so day-zero traffic doesn't take them down.
About You
4+ years in SRE, production engineering, or infrastructure roles running high-traffic, customer-facing systems.
Strong with observability tooling and practice: metrics, tracing, logs, SLOs/error budgets, alerting that is always actionable.
Capable software engineer who prefers writing tools over executing runbooks. TypeScript and/or Python.
Experienced with distributed systems failure modes: timeouts, retries, backpressure, partial outages, noisy neighbors.
Calm, clear incident commander who communicates well with external partners under pressure.
Understands, or is eager to learn deeply, how LLM inference is served: streaming, tool calling, prompt caching, throughput/latency tradeoffs, and how provider APIs differ.
Nice to Have
Experience at an inference provider, model lab, GPU cloud, or API gateway/CDN company.
Experience with our stack: TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, Vercel.
Background in routing, load balancing, or traffic management systems.
Experience with evals or synthetic monitoring for ML systems.
How we rate this
Site Reliability Engineer, Provider Operations at OpenRouter rates 85 out of 100 for how much of the daily work is AI. That makes it Builds AI (AI Level 4 of 4). The level is about AI in the job, not seniority.
Builds AI. The job is building AI systems.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- Tell me about a project where site reliability engineering was part of your work. What did you do?
- Tell me about a project where observability was part of your work. What did you do?
- Tell me about a project where incident response was part of your work. What did you do?
- Tell me about a project where infrastructure was part of your work. What did you do?
- Tell me about a project where load testing was part of your work. What did you do?
Adapt your resume
- List these exact terms on your resume: Site Reliability Engineering, Observability, Incident Response, Infrastructure, and Load Testing. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them. A tool named with nothing behind it rarely survives a human read.
- Lead with what you built, trained or shipped. This role is judged on the AI system itself, not the tools around it.
Want an expert to read your CV for this job?
Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.
Get new remote software engineering jobs (Builds AI ●●●●) by email
One email a week with the new remote software engineering jobs (Builds AI ●●●●), each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Software Engineering roles that build AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step