# Information Technology Lead at Zendesk

Zendesk is hiring an Information Technology Lead in Pune, India. Level rates it Little AI ●○○○; you can [apply on Level](https://jobsbylevel.com/go/77a55568-9f5b-4f3e-b376-f5eccefa5398).

AI Level 1, AI centrality 35 out of 100. Pune, India.

## Details

- Company: [Zendesk](https://jobsbylevel.com/companies/zendesk)
- AI level: AI Level 1 (score 35 out of 100)
- Location: Pune, India
- Posted: October 8, 2026
- Apply: https://jobsbylevel.com/go/77a55568-9f5b-4f3e-b376-f5eccefa5398

## Description

Job Description Job Title: Observability & Monitoring Lead Location: Pune, India Love making complex systems feel simple and reliable? We’re looking for an Observability & Monitoring Engineer who is equal parts builder and detective—someone who instruments services end-to-end, shines a light on blind spots, and turns noise into actionable signals. You’ll help us evolve a modern RunOps capability that improves reliability, reduces toil, and elevates the employee experience across Zendesk. About Zendesk At Zendesk, we believe outstanding customer and employee experiences start with great service and resilient platforms. We lead with empathy, innovate with purpose, and celebrate diversity and inclusion in everything we do. Join our global team and help us build an observability practice that others want to copy. The Role: We are seeking a highly experienced Observability & Monitoring Lead with 6+ years of industry experience to spearhead our SRE and observability strategy. You will be responsible for maturing our monitoring ecosystem, driving operational excellence, and reducing system downtime through proactive detection and incident management. You will bridge the gap between engineering and operations, ensuring that observability data directly translates into improved reliability, SLA compliance, and reduced MTTR. You will design and operate the telemetry backbone for our internal platforms and business-critical applications. This role spans metrics, logs, traces, synthetics, RUM, and event correlation—instrumenting services, building dashboards, tuning alerts, and partnering with Incident/Problem/Change to drive measurable reliability outcomes. What You’ll Do Design the observability stack: Define and implement standards for metrics, logs, traces, and profiling (e.g., OpenTelemetry collectors, exporters, and context propagation). Instrument what matters: Establish golden signals, SLIs/SLOs, and health checks for priority services; automate baselining and anomaly detection. Build actionable visibility: Create executive and on-call views (dashboards, service health, dependency maps) for Apps, Network, Collaboration tools, HRIS, and integrations. Engineer signal > noise: Develop alerting policy as code; reduce false positives; implement suppression, deduplication, and auto- Security & compliance: Ensure monitoring data is handled per policy; implement role-based access and guardrails for sensitive logs/metrics Incident Reduction: Develop and execute strategies to systematically reduce the volume of high-priority incidents through trend analysis and proactive stability improvements. Incident Automation: Lead the engineering of automated response workflows and self-healing capabilities to minimize manual intervention and accelerate resolution times. Strategy & Architecture: Design and coordinate with engineering teams for comprehensive observability strategies, including logging, metrics, and tracing across distributed systems. Tooling Enhancements : End-to-end management of observability stacks (Data Dog) and its integration with ITSM tools (Ticketing Tools) and applications. Seek assistance from product and engineering team for any tool enhancement use cases to build a robust observability platform integrated with the ticketing tools. Application Onboarding: Responsible and accountable for onboarding relevant observability and monitoring necessities for new applications which are onboarded and handed over to operations. Incident & SLA Management: Oversee the operational lifecycle of monitoring tickets and Lead the monthly analysis and reporting of performance metrics (Eg - Mean time to detect, Mean time to restore,Mean time to Respond/Accept, First touch resolution) and ensure SLA compliance . Continual Service Improvement: Liaise with technical and level 2 teams to improve and mature the current monitoring state of applications in scope. Process Automation: Drive the automation of incident reporting, alert correlation rules,

The description is cut here. Read the full offer: https://jobsbylevel.com/jobs/information-technology-lead-at-zendesk-72e668

Source: https://jobsbylevel.com/jobs/information-technology-lead-at-zendesk-72e668

## Cite this page

Level. https://jobsbylevel.com/jobs/information-technology-lead-at-zendesk-72e668.

Get job alerts: https://jobsbylevel.com/newsletter
