Level

JobgetherPosted today

L1

Lead Software Engineer, Cloud Site Reliability (SRE)

Lead Software Engineer, Cloud Site Reliability (SRE) at Jobgether scores 30 out of 100 on AI centrality, which makes it a Level 1 role on this board.

Remote (India)leadFull-time

AI in this role

Lead SRE to manage cloud reliability, Azure infrastructure, and advance AIOps and self-healing capabilities.

azureakskubernetesdockerdatadog
site-reliabilityincident-managementinfrastructure-as-codeobservabilityaiops

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Software Engineer, Cloud Site Reliability (SRE) based in India.

This role leads cloud reliability and 24x7 site reliability operations for critical technology environments, with a strong focus on Azure infrastructure and cloud-native platforms. You will take ownership of major incidents, drive operational excellence, and ensure high availability and SLA adherence across production systems. The position combines hands-on cloud engineering with observability, automation, incident management, and reliability improvements. You will work extensively with Azure, AKS, Kubernetes, Docker, Datadog, and infrastructure-as-code technologies to build resilient and scalable environments. The role also provides an opportunity to advance proactive monitoring, anomaly detection, AIOps, and self-healing capabilities. As a technical leader, you will mentor engineers, collaborate with development teams, and communicate operational performance to stakeholders and leadership.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

Site ReliabilityIncident ManagementInfrastructure As CodeObservabilityAiopsAzureAksKubernetes

Questions you could be asked

  1. Tell me about a project where site reliability was part of your work. What did you do?
  2. Tell me about a project where incident management was part of your work. What did you do?
  3. Tell me about a project where infrastructure as code was part of your work. What did you do?
  4. Tell me about a project where observability was part of your work. What did you do?
  5. Tell me about a project where aiops was part of your work. What did you do?

Adapt your resume

  • List these exact terms on your resume: Site Reliability, Incident Management, Infrastructure As Code, Observability, and Aiops. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.

Want your resume actually rewritten for this job?

The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.

Similar roles

Software Engineering roles rated Level 1 at other companies.

More jobs at Jobgether

Related searches

Same AI level