Level

NubankPosted 1mo ago

L1

Senior Reliability Engineer — Regulatory Reliability

Senior Reliability Engineer — Regulatory Reliability at Nubank scores 36 out of 100 on AI centrality, which makes it a Level 1 role on this board.

Remote (Ciudad de México)seniorFullTime

AI in this role

claudecursorclaude-code
About Nu

Nu is the leading digital bank in Latin America, serving 140 million customers across Brazil, Mexico, and Colombia. The company has been leading an industry transformation by leveraging data and proprietary technology to develop innovative products and services.

Guided by its mission to fight complexity and empower people, Nu caters to customers’ complete financial journey, promoting financial access and advancement with responsible lending and transparency. The company is powered by an efficient and scalable business model that combines low cost to serve with growing returns.

Nu’s impact has been recognized in multiple awards, including Time 100 Most Influential Companies, Fast Company’s Most Innovative Companies, and Forbes World’s Best Banks.

Visit our Institutional Page

About the role

We are looking for a Senior Reliability Engineer to drive reliability, resilience, and operational excellence for Nubank's regulatory infrastructure. In this role, you will own critical reliability outcomes for systems that support real-time interbank transfers, regulatory integrations and key regulatory platforms. You will work across infrastructure, software, networking, security, and operations to keep mission-critical services available, observable, and resilient under demanding regulatory and operational constraints.

What you'll do

  • Own end-to-end reliability for IT regulatory infrastructure services, including connectivity, transaction processing flows, and settlement-related operational workflows.

  • Define, measure, and continuously improve SLIs, SLOs, and operational health indicators for critical transactions and platforms.

  • Lead incident response for different severity production events, coordinate recovery, and drive high-quality postmortems with concrete follow-through on corrective actions.

  • Design and evolve observability for the platform, with emphasis on tracing, queue health, signature validation, infrastructure signals, and early detection of degraded regulatory links or transaction bottlenecks.

  • Plan and execute disaster recovery and business continuity exercises, including failover validation, contingency readiness, and RTO/RPO verification.

  • Reduce operational toil through automation, tooling, and improved runbooks for recurring failure and operational procedures.

  • Drive capacity planning and performance engineering to ensure the platforms can safely absorb peak transaction volumes and evolving business demand.

  • Partner closely with security, networking, middleware, and software engineering teams to improve platform hardening, resilience, and change management safety.

  • Participate in an on-call rotation.

What we're looking for

  • Experience with SRE, DevOps, or production engineering experience operating mission-critical, high-availability systems.

  • Hands-on experience with IT infrastructure platforms, virtualized and hyperconverged environments such as VxRail, and physical production infrastructure.

  • Hands-on experience with Linux systems, including performance tuning, kernel parameters, and security hardening.

  • Track record leading incident response and postmortem processes for customer-impacting services.

  • Solid knowledge of networking fundamentals: TCP/IP and routing.

  • Proficiency in at least one scripting or programming language (Python, Shell scripting) for automation and tooling.

  • Experience with observability stacks (Prometheus, Grafana) and distributed tracing.

  • AWS infrastructure experience across services like EC2, S3, CloudWatch, KMS, EKS/ECS, VPCs, RDS/DynamoDB, and SQS/SNS.

  • IT Regulatory audit experience, including evidence gathering, control validation, audit prep, and remediation follow-through.

  • Experience with infrastructure-as-code tools (like Puppet) and Git-based workflows.

  • Working knowledge of AI-assisted engineering tools (Claude, Cursor, Claude Code).

  • Strong written and verbal communication skills in Spanish and English.


Location for this opportunity

  • Mexico City, Mexico

Our Benefits

  • Chance of earning equity at Nu

  • Extended maternity and paternity leaves

  • Health and life insurance

  • Dental and Vision Insurance

  • NuCare - Our mental health and wellness assistance program

  • Nucleo - Our learning platform of courses

  • NuLanguage - Our language learning program

  • Holiday Bonus ("Aguinaldo") of 30 days of pay per year

  • 17 days of paid vacation with 25% vacation bonus

  • Gym partnership

  • Food card

  • Work-from-home Allowance

  • Parental Consultancy

  • Relocation Assistance Package, if applicable

Work Model for this Role

Our recruitment process may involve the use of artificial intelligence–enabled tools, such as automated interview transcription and analysis, to support the evaluation process. Artificial intelligence is not used to make final hiring decisions; all decisions are made by human reviewers.

L1Get new remote AI jobs by email

One email a week with the new remote AI jobs, each rated Level 1 to 4 for how much AI is in the work. No recruiter spam, unsubscribe in one click.

Free. One email a week. Unsubscribe in one click.

Similar roles

Software Engineering roles rated Level 1 at other companies.

More jobs at Nubank

Related searches

Same AI level