Senior Software Engineer, Reliability
AI in this role
Senior Software Engineer on the Reliability team to build robust infrastructure, monitoring, and fault-tolerant systems for large-scale platforms.
Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators.
At Roblox, we’re building the tools and platform that empower our community to bring any experience that they can imagine to life. Our vision is to reimagine the way people come together, from anywhere in the world, and on any device. We’re on a mission to connect a billion people with optimism and civility, and looking for amazing talent to help us get there.
A career at Roblox means you’ll be working to shape the future of human interaction, solving unique technical challenges at scale, and helping to create safer, more civil shared experiences for everyone.
Are you a seasoned engineer with a passion for reliability and scalability? We’re looking for exceptional Software Engineers to join the Reliability team at Roblox. In this pivotal role, you will drive the evolution of our systems, ensuring they meet the highest standards of performance, reliability, and efficiency. You’ll collaborate with cross-functional teams to build robust infrastructure that supports our growth. If you have a track record of solving complex technical challenges, we want to hear from you. Join us in shaping the future of our platform and delivering unparalleled value to our users.
At Roblox, our vision is to achieve 1 billion daily active users. We believe this engineer will be instrumental in driving us towards that ambitious goal.
You Will:
- Create libraries that promote fault-tolerance and resilience– like retries, circuit breakers, and adaptive concurrency limits.
- Build, automate and standardize process automation to create a "golden path" of tooling and platform support that powers the fundamental Roblox ecosystem.
- Create tooling that provides production guardrails, for example evaluating release candidate capacity with load testing tooling before deploying to production.
- Create performance monitoring services and observability towards understanding capacity issues and platform degradations.
- Create tooling that monitors production services and their changes, like generalized canarying services with alerting.
You Have:
- Experience: you have a BS degree (or equivalent professional experience) in Computer Science or related engineering field with at least 4+ years of experience with added advantage working in the Site Reliability space in SRE or Software Engineering
- Passion for systems: You have experience and good habits around building software and tools and getting them adopted. Your system's focus informs a view of code needing to be deeply reliable.
You Are:
- A Partner: You know that the best tools integrate broadly with the tooling ecosystem. You approach partners and processes with curiosity and seek to understand a problem deeply before you start coding.
- A Coder: you have experience writing common programming languages ( Go, C#, Java…).
- Self-organized: you're excited about getting in front of complex problems, organizing your work by any means possible; overcome emergent issues and contributing to long-running projects as a part of the team.
- Problem Solver: you ask the right questions to solve issues within your expertise and you use data to test your theories.
- Planner - You have experience in large project lifecycles. You have experienced working in sprints, breaking down complex tasks into milestones, and reporting status to keep project scheduling accurate.
For roles that are based at our headquarters in San Mateo, CA: The starting base pay for this position is as shown below. The actual base pay is dependent upon a variety of job-related factors such as professional background, training, work experience, location, business needs and market demand. Therefore, in some circumstances, the actual salary could fall outside of this expected range. This pay range is subject to change and may be modified in the future. All full-time employees are also eligible for equity compensation and for benefits as described on this page.
Annual Salary Range$243,290—$295,250 USDRoles that are based in an office are onsite Tuesday, Wednesday, and Thursday, with optional presence on Monday and Friday (unless otherwise noted).
Roblox provides equal employment opportunities to all employees and applicants for employment and prohibits discrimination and harassment of any type without regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws. Roblox also provides reasonable accommodations to candidates with qualifying disabilities or religious beliefs during the recruiting process.
For US based roles only, please note the Company may not be able to employ candidates for this role who have United States work authorization related to certain U.S. visa categories, or support future H-1B sponsorship at this time.
How we rate this
Senior Software Engineer, Reliability at Roblox rates 0 out of 100 for how much of the daily work is AI. That makes it Little AI (AI Level 1 of 4). The level is about AI in the job, not seniority.
Little AI. AI is not part of the work.
- ●●●● Builds AI80 to 100
- ●●●○ Works on AI60 to 79
- ●●○○ Uses AI40 to 59
- ●○○○ Little AI0 to 39
Levels come from how often the tools, models and workflows of the role are named in the posting itself. Open the description and count.
Prepare for this job
A free preview built only from this posting: what it asks for, what you could be asked in an interview, and how to adjust your resume.
Skills and AI tools this role asks for
Questions you could be asked
- Tell me about a project where sre was part of your work. What did you do?
- Tell me about a project where reliability was part of your work. What did you do?
- Tell me about a project where scalability was part of your work. What did you do?
- Tell me about a project where observability was part of your work. What did you do?
- Tell me about a project where monitoring was part of your work. What did you do?
Adapt your resume
- List these exact terms on your resume: Sre, Reliability, Scalability, Observability, and Monitoring. An applicant tracking system matches the wording, not the idea.
- Attach one line of real, concrete experience to at least one of them — a tool named with nothing behind it rarely survives a human read.
Want your resume actually rewritten for this job?
The free preview above is everything we have today. A full resume rewrite is not live yet and has no price set. Join the waitlist and we will email you if we open it.
Get new AI jobs by email
One email a week with the new AI jobs, each rated for how much AI is in the work. No recruiter spam, unsubscribe in one click.
Free. One email a week. Unsubscribe in one click.
Similar roles
Software Engineering roles that involve little AI, at other companies.
What kind of AI work fits you?
Answer 12 practical questions in about three minutes. Get a simple profile, the work it points to, and live roles to explore next.
Find my next step