Site Reliability Engineer (DevSecOps), Production Engineering - ThousandEyes
Cisco is hiring a Site Reliability Engineer (DevSecOps), Production Engineering - ThousandEyes in Krakow, Poland. Level rates it ; you can apply on Level.
AI in this role
Design and manage large-scale distributed cloud systems and infrastructure to support high availability and reliability for ThousandEyes.
Please note that we operate a hybrid working model, and this role will require you to work in one of our Engineering hubs 2 days a week.
Cisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital experiences across every network—even those beyond their ownership. Leveraging AI and an unparalleled set of cloud, internet, and enterprise network telemetry data, ThousandEyes enables IT teams to proactively detect, diagnose, and resolve issues before they impact end-user experiences.
ThousandEyes is deeply integrated across Cisco's extensive technology portfolio, supporting customers in scaling deployments while offering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios.
Your Impact
We are seeking a skilled Site Reliability Engineer (DevSecOps) in Production Engineering with a strong background in SaaS, operations and security. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating directly with application development teams to enhance the reliability, performance, and security of our platform.
Responsibilities
- Collaborate with software engineers to optimize architecture and services for availability, latency, performance, and reliability using cloud-native tools.
- Design and implement scalable operations tooling to support platform growth and scaling across multiple regions.
- Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure.
- Participate in and improve our 24x7 incident response and on-call rotation.
- Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability.
- Automate production operations to provide guardrails and continuous platform operation.
- Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing.
- Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform.
- Identify and provide solutions to common obstacles hindering operational excellence across engineering teams.
- Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform.
- Play a key role in the ThousandEyes platform by leveraging scale testing, additional environments, and working with application teams to improve system reliability.
- Manage a rapidly growing infrastructure capable of handling substantial daily data volumes, emphasizing operations/infrastructure/everything as code.
Minimum Qualifications
- Hands-on experience deploying, operating, and troubleshooting containerized workloads in production Kubernetes environments.
- Professional experience diagnosing and administrating Linux/Unix systems, including process management, file systems, and networking protocols (TCP/IP, DNS, HTTP).
- Professional experience developing automation tooling, operational scripts, or backend services using Python or Go.
- Practical experience building hardened container images and integrating automated security scanning tools (SAST, DAST, or container vulnerability scanners) into CI/CD pipelines.
Preferred Qualifications
- Familiarity with best practices for operating a large-scale, highly available enterprise platform.
- 3+ years of experience in a related role.
- Excellent communication and documentation skills.
- Strong sense of ownership, drive, and attention to detail.
Why Cisco?
At Cisco, we’re revolutionizing how data and infrastructure connect and protect organizations in the AI era – and beyond. We’ve been innovating fearlessly for 40 years to create solutions that power how humans and technology work together across the physical and digital worlds. These solutions provide customers with unparalleled security, visibility, and insights across the entire digital footprint.
Fueled by the depth and breadth of our technology, we experiment and create meaningful solutions. Add to that our worldwide network of doers and experts, and you’ll see that the opportunities to grow and build are limitless. We work as a team, collaborating with empathy to make really big things happen on a global scale. Because our solutions are everywhere, our impact is everywhere.
We are Cisco, and our power starts with you.
How we rate this
Site Reliability Engineer (DevSecOps), Production Engineering - ThousandEyes at Cisco rates 25 out of 100 for how much of the daily work is AI. That makes it Little AI (AI Level 1 of 4). The level is about AI in the job, not seniority.
Little AI. AI is not part of the work.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- Tell me about a project where site reliability engineering was part of your work. What did you do?
- Tell me about a project where devsecops was part of your work. What did you do?
- Tell me about a project where distributed systems was part of your work. What did you do?
- Tell me about a project where automation was part of your work. What did you do?
- Tell me about a project where cloud native was part of your work. What did you do?
Adapt your resume
- List these exact terms on your resume: Site Reliability Engineering, Devsecops, Distributed Systems, Automation, and Cloud Native. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
Want an expert to read your CV for this job?
Free. Send your CV and the role you want next. We reply by email within 2 to 4 business days.
Get a free CV reviewGet new remote cybersecurity jobs by email
One email a week with the new remote cybersecurity jobs, each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Software Engineering roles that involve little AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step