Level

Cisco

Site Reliability Engineer

AI in this role

Operate and scale reliable cloud infrastructure supporting large-scale data and ML workloads at Cisco ThousandEyes.

apache-airflowaws-emrsparkhadoopamazon-eksterraformpythonaws
site-reliability-engineeringinfrastructure-as-codefinopscloud-computingautomation

Site Reliability Engineer — NADP Platform 

Location 

Bangalore, India — Hybrid 

Team 

Cisco ThousandEyes — Network Assurance Data Platform / Cloud & SRE Engineering 

 

Meet the Team 

The Network Assurance Data Platform team within Cisco ThousandEyes is responsible for building, operating, and scaling the core data infrastructure that powers large-scale network assurance, analytics, and intelligence capabilities. 

The team manages mission-critical cloud infrastructure, big data workflows, ML platform components, and reliability engineering practices across AWS environments. We focus on improving platform reliability, scalability, performance, automation, and cost efficiency while enabling engineering and data teams to deliver business-critical capabilities at scale. 

As a Site Reliability Engineer in the NADP team, you will work on highly scalable infrastructure supporting data pipelines, ML workloads, cloud-native services, and cost-optimized AWS operations. 

 

Your Impact 

  • Own and operate scalable, reliable, and cost-efficient infrastructure for the ThousandEyes Network Assurance Data Platform. 
  • Manage, optimize, and improve large-scale data workflows using Apache Airflow, AWS EMR, Spark, and Hadoop-based processing platforms. 
  • Operate and improve Amazon EKS environments supporting containerized services, ML workloads, and production-grade platform components. 
  • Build and maintain infrastructure automation using Terraform and other infrastructure-as-code practices. 
  • Develop Python-based automation, integrations, operational tooling, reporting, and reliability improvements. 
  • Drive FinOps practices across NADP infrastructure, including cost visibility, cost allocation, forecasting, anomaly detection, optimization, and governance. 
  • Partner with data engineering, ML, platform, finance, and product teams to improve reliability, performance, scalability, and cost efficiency. 
  • Identify infrastructure bottlenecks, performance issues, inefficient workloads, and cost optimization opportunities. 
  • Improve observability, alerting, incident response, and operational readiness across data and ML platforms. 
  • Support capacity planning, right-sizing, autoscaling, storage optimization, and workload efficiency across AWS services. 
  • Lead technical discussions, influence design decisions, and guide teams toward reliable and cost-conscious architecture. 
  • Provide senior-level technical leadership, mentorship, and operational guidance to engineers across the team. 
  • Drive continuous improvement in platform reliability, automation, deployment practices, and operational excellence. 

 

Minimum Qualifications 

  • Bachelor’s degree or higher in Engineering, Computer Science, or equivalent practical experience. 
  • 8–10 years of relevant experience in Site Reliability Engineering, DevOps, Cloud Infrastructure, Platform Engineering, Data Infrastructure, or Production Engineering. 
  • Strong hands-on experience operating production infrastructure on AWS. 
  • Strong experience with Apache Airflow for workflow orchestration, pipeline operations, scheduling, monitoring, and troubleshooting. 
  • Hands-on experience with AWS EMR, Spark, Hadoop, or similar large-scale data processing platforms. 
  • Strong experience operating Amazon EKS or Kubernetes-based environments in production. 
  • Experience supporting containerized workloads, preferably including ML workloads or data platform services. 
  • Strong hands-on experience with Terraform and infrastructure-as-code practices. 
  • Strong Python programming or scripting experience for automation, integrations, operational tooling, and infrastructure workflows. 
  • Practical experience with cloud cost optimization, FinOps, AWS cost analysis, tagging, budgeting, forecasting, and cost governance. 
  • Strong understanding of Linux systems, networking, distributed systems, and production troubleshooting. 
  • Experience with observability tools such as CloudWatch, Prometheus, Grafana, Splunk, OpenSearch, Datadog, or similar platforms. 
  • Experience with incident management, production support, root cause analysis, reliability improvements, and operational excellence. 
  • Ability to analyze infrastructure, performance, and cost data and convert findings into clear technical recommendations. 
  • Strong communication skills with the ability to collaborate across data, ML, platform, finance, and engineering teams. 
  • Ability to operate independently, drive initiatives end to end, and provide technical leadership in a fast-paced environment. 

 

Preferred Qualifications 

  • Experience working with large-scale SaaS platforms or high-volume data infrastructure. 
  • Experience with ML infrastructure, model execution platforms, batch processing, or data pipeline reliability. 
  • Experience optimizing EMR, Spark, Airflow, EKS, storage, and compute workloads for performance and cost. 
  • Experience with AWS services such as EC2, S3, RDS, IAM, VPC, CloudWatch, OpenSearch, Lambda, ElastiCache, and related cloud-native services. 
  • Experience with AWS Savings Plans, Reserved Instances, Spot adoption, Graviton migration, storage lifecycle management, and workload right-sizing. 
  • Experience with cloud cost management tools such as AWS Cost Explorer, AWS CUR, Cloudability, CloudHealth, Kubecost, or similar platforms. 
  • Experience driving FinOps programs, cost reviews, stakeholder reporting, OKR tracking, and executive-level updates. 
  • Experience with CI/CD systems, GitHub workflows, Atlantis, or similar deployment automation platforms. 
  • Experience with Puppet, Ansible, Helm, Argo CD, or other configuration and deployment management tools. 
  • FinOps certification or equivalent hands-on cloud financial management experience is a plus. 
  • Ability to influence engineering teams toward cost-aware, scalable, and reliable design patterns. 

 

What Success Looks Like 

  • NADP data and ML infrastructure is reliable, scalable, performant, and cost-efficient. 
  • Airflow, EMR, Spark, and EKS workloads are operated with strong observability, automation, and production maturity. 
  • AWS cost visibility, forecasting, and governance are improved across NADP-owned infrastructure. 
  • Cloud wastage is reduced through right-sizing, automation, storage optimization, and workload efficiency improvements. 
  • Infrastructure is managed through production-grade Terraform and automation practices. 
  • Engineering and data teams have better visibility into platform health, cost drivers, risks, and optimization opportunities. 
  • Cross-functional stakeholders receive clear updates on reliability, performance, cost trends, and execution progress. 
  • The team continuously improves operational excellence, incident response, and platform reliability. 

 

Role Summary 

We are looking for a Site Reliability Engineer with 8+ years of experience and strong hands-on expertise in AWS, Apache Airflow, EMR, Spark, EKS, Terraform, Python, and FinOps. This role is ideal for someone who enjoys operating large-scale data and ML infrastructure, improving reliability, automating operational workflows, and driving cloud cost efficiency across complex production systems. 


Why Cisco? 

At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.

Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere. 

We are Cisco, and our power starts with you. 


Disclaimer

To ensure that we hire the best talent in the right way, we follow a strict hiring process and recently, Cisco has been made aware of fraudulent recruiters claiming to be from the company. Please be advised that any communication from Cisco about careers will:

  • be in direct response to an application you have submitted through the company career site
  • begin with screening or an interview
  • originate from a Cisco email address, and
  • be conducted across email, phone, or WebEx

 

Cisco will never make a job offer without conducting an interview process or ask you for money in any way. If you have been requested to apply for a role or have received an offer from a site other than https://careers.cisco.com or cisco.wd5.myworkday.com, do not provide any personal identifying information, including your Aadhaar or other personal identifying number, birth certificate, banking information, driver's license, or passport.


If you are the target of a recruiting scam, consider filing a report with your local law enforcement authorities. Cisco bears no responsibility, and cannot be held liable, for any claims, damages, expenses, or other inconvenience resulting from or in any way connected to recruiting scams.


How we rate this

Site Reliability Engineer at Cisco rates 30 out of 100 for how much of the daily work is AI. That makes it Little AI (AI Level 1 of 4). The level is about AI in the job, not seniority.

Classification

Little AI. AI is not part of the work.

  1. ●●●● Builds AI80 to 100
  2. ●●●○ Works on AI60 to 79
  3. ●●○○ Uses AI40 to 59
  4. ●○○○ Little AI0 to 39

Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.

Prepare for this job

A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.

Skills and AI tools this role asks for

Site Reliability EngineeringInfrastructure As CodeFinopsCloud ComputingAutomationApache AirflowAWS EmrSpark

Questions you could be asked

  1. Tell me about a project where site reliability engineering was part of your work. What did you do?
  2. Tell me about a project where infrastructure as code was part of your work. What did you do?
  3. Tell me about a project where finops was part of your work. What did you do?
  4. Tell me about a project where cloud computing was part of your work. What did you do?
  5. Tell me about a project where automation was part of your work. What did you do?

Adapt your resume

  • List these exact terms on your resume: Site Reliability Engineering, Infrastructure As Code, Finops, Cloud Computing, and Automation. An applicant tracking system matches the wording, not the idea.
  • Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.

Want an expert to read your CV for this job?

Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.

Get a free CV review

Get new AI jobs by email

One email a week with the new AI jobs, each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.

Free. One email a week. Unsubscribe in one click.

Similar roles

Software Engineering roles that involve little AI, at other companies.

What kind of AI work fits you?

Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.

Find my next step

More jobs at Cisco

Cisco

$216k-$281kSan Jose, California, US4d

Related searches

Same AI level