# Senior Software Engineer - Reliability at Atlan

Atlan is hiring a Senior Software Engineer - Reliability for a remote role open to applicants in India. Level rates it Works on AI ●●●○; you can [apply on Level](https://jobsbylevel.com/go/c8f009bd-eb6e-41c6-b1e2-cd499fbc63bc).

AI Level 3, AI centrality 69 out of 100. Remote (India).

## Details

- Company: [Atlan](https://jobsbylevel.com/companies/atlan)
- AI level: AI Level 3 (score 69 out of 100)
- Location: Remote (India)
- Posted: October 8, 2026
- Apply: https://jobsbylevel.com/go/c8f009bd-eb6e-41c6-b1e2-cd499fbc63bc

## Description

Who We Are Atlan is building the context layer for enterprise AI. Enterprises are pouring money into AI and most of it dies in production because the AI does not understand the business context around the data. That is the problem Atlan solves. Gartner has named context the defining issue in enterprise AI, and calls context graphs the essential infrastructure for AI agents, naming Atlan one of three vendors already building it. We are also the only vendor named a Leader across all four major Gartner and Forrester evaluations for data catalogs, data governance and metadata management, the foundation the context layer runs on. Come build the infrastructure that AI runs on. The Team You'll join the Reliability team, which runs a multi-agent AI SRE platform that is already live in production. Today, agents investigate incidents, remediate them, and resolve networking tickets across every tenant we run. A small core team built it and shipped it fast. The scope has now outgrown them. Be clear on what this role is not. It is not an incident-command SRE seat. Our view is simple: if a human is fixing production by hand, the system has failed. It is also not a greenfield charter. You'll join a mature, opinionated codebase with documented architecture decisions and eval-gated pull requests, and you'll make it better. Why now: the platform works, and the hardest problems are wide open. Making agent investigations accurate enough that engineers trust them without re-checking is a genuine frontier problem in agentic reliability. Other Atlan teams are lining up to put their own reliability metrics on the platform. What you build in the next year decides how far autonomous operations can go at Atlan. What You Will Do Make investigation agents right, not just fast. Push root-cause accuracy toward the point where engineers trust the answer without re-checking it. Treat every wrong diagnosis as a class of failure to remove, not a one-off bug to patch. Grow auto-remediation coverage, safely. Take remediation from a handful of playbooks to dozens. Each one earns autonomy in stages: dry run before execute, human-approved before autonomous. Own the evals that decide when an agent can be trusted. Build and run fault-injection benchmarks and eval harnesses. Set pass marks before the run, report results with honest denominators, and validate the harness itself, not just the agent. Design the gates that let an agent write to production. Fail-closed checks, kill switches, approval flows and blast-radius limits. You start from the question "what happens when the agent is wrong?" Build new agents where the toil says they belong. Every agent maps to a category of toil it removes, and each one you ship should make the next one cheaper to build. Build the platform for other teams. Clean interfaces, guardrails and safe defaults, so teams beyond reliability can trust the platform on their own production. Your hiring manager owns cross-team alignment. You earn adoption through what you build. Ship the unglamorous fix when it matters. When something is on fire, you stop the bleeding first, then come back and remove the class. What Makes You a Match You've lived operational toil. You've carried a pager, owned incidents end to end, or worked a support or escalation queue long enough to know which pain is worth removing and why. The toil was yours, not something you read about. You remove classes, not tasks. You can point to a recurring operational problem you made disappear, explain the category behind it, and tell us why you picked that one. Not a faster runbook. Gone. You start from the problem and measure by adoption. When you describe something you built, you reach on your own for who used it and what changed, with real numbers and real denominators. You won't call one success a rate. You can also name something you built that didn't get adopted, and what you learned. AI has changed how you work, structurally. You've rebuilt a core part of your own work end to

The description is cut here. Read the full offer: https://jobsbylevel.com/jobs/senior-software-engineer-reliability-at-atlan-00ce0d

Source: https://jobsbylevel.com/jobs/senior-software-engineer-reliability-at-atlan-00ce0d

## Cite this page

Level. https://jobsbylevel.com/jobs/senior-software-engineer-reliability-at-atlan-00ce0d.

Get job alerts: https://jobsbylevel.com/newsletter
