Senior Manager, Site Reliability Engineering (SRE)
Nium · India
Nium provides global infrastructure for real-time cross-border payments. We were founded on the mission to deliver the global payments infrastructure of tomorrow, today. Our platform enables banks, fintechs, and global businesses to move money instantly, everywhere. Co-headquartered in San Francisco and Singapore with offices in 14 markets worldwide, we are entering one of the most exciting chapters in our journey. In March 2026, we delivered the largest month in our 11-year history with record revenue, record volumes, and EBITDA profitability. Today, Nium moves nearly $60B in payments annually, almost entirely for enterprises, while continuing to strengthen an already healthy balance sheet. It is an incredible time to join us, and we are only just getting started. Our payout network spans 190+ countries and 100 currencies, with 100 + corridors in real time. We power seamless transfers to accounts, wallets, and cards, support local collections in 35 markets, and as a principal card issuer on Visa, Mastercard, Discover, and UATP, Nium issues over 50 million card tokens every year. Backed by regulatory licenses in 40+ markets, we make it simple for our partners to onboard, integrate, and scale globally. This scale and innovation have earned us recognition as one of CNBC’s World’s Top Fintech Companies 2025, winner of Best Cross-Border Payments Solution at the PayTech Awards, and inclusion in FXC Intelligence’s Top 100 Cross-Border Payments Companies list. In 2024, we raised US$50 million in Series E funding at a US$1.4 billion valuation to accelerate network expansion, product innovation, and talent growth. With the B2B payments market projected to hit US$175 trillion by 2030, Nium offers ambitious builders the chance to shape the future of global money movement with the scale of a leader and the energy of a high-growth company.
Responsibilities : Lead, mentor, and grow a team of SREs and reliability engineers across multiple time zones, setting clear goals, career paths, and performance expectations.
Own Nium's reliability strategy — defining and driving SLIs/SLOs/error budgets across critical payment, card issuance, and compliance services.
Drive incident management end-to-end: on-call structure, escalation paths, major incident response, and blameless postmortems that produce durable fixes, not just tickets.
Partner with product engineering leaders to embed reliability, scalability, and operational readiness into the software development lifecycle from design through launch.
Build and scale observability (metrics, logging, tracing, alerting) so that issues are detected and diagnosed before they impact customers or partner banks.
Lead capacity planning and performance engineering for systems processing high-volume, real-time financial transactions across a global, multi-region infrastructure.
Champion automation and self-healing systems to reduce toil, eliminate manual runbooks, and improve mean-time-to-detect and mean-time-to-resolve.
Own disaster recovery, business continuity, and chaos engineering practices, running regular game days to validate resilience assumptions.
Collaborate with Security and Compliance teams to ensure infrastructure practices meet regulatory and audit requirements (PCI-DSS, SOC 2, ISO 27001, and regional financial regulations).
Manage the reliability budget: tooling investments, cloud cost/performance trade-offs, and staffing plans, in partnership with finance and engineering leadership.
Represent SRE in executive reviews, translating technical risk and system health into business-relevant reporting for leadership and the board.
Establish and continuously refine production readiness reviews, runbooks, and operational standards across all engineering teams.
Requirements: 10+ years of experience in software engineering, infrastructure, or site reliability engineering, with 4+ years in a people-management or technical leadership role leading SRE/DevOps/Infrastructure teams.
Proven track record operating and scaling production systems for a high-availability, transaction-heavy platform — fintech, payments, banking, or e-commerce experience strongly preferred.
Deep hands-on expertise with cloud infrastructure (AWS) Kubernetes, container orchestration, and infrastructure-as-code (Terraform, CloudFormation, or similar).
Strong background in observability stacks (Prometheus, Grafana, Datadog, ELK/OpenSearch, or equivalent) and building alerting that reduces noise while catching real issues.
Demonstrated experience defining and operationalizing SLOs/error budgets, and using them to drive engineering prioritization.
Solid understanding of distributed systems, databases, caching, messaging queues, and API-driven microservice architectures at scale.
Experience leading major incident response for critical, customer-facing systems, including postmortem processes that drive real change.
Familiarity with security and compliance frameworks relevant to financial services (PCI-DSS, SOC 2, ISO 27001) and how they shape infrastructure and access practices.
Excellent communication skills — able to translate technical reliability concepts for engineering peers, product leaders, and executives alike.
A pragmatic, metrics-driven approach to engineering decisions: you measure before you optimize, and you test assumptions in production rather than in theory.
Experience with scripting/automation (Python, Go, or Bash) and CI/CD pipelines (Jenkins, ArgoCD, GitLab CI, or similar).
Nice to Have
Experience operating systems subject to real-time payment rails, card networks, or banking core integrations.
Prior experience running SRE or infrastructure functions through hypergrowth or rapid international expansion.
Exposure to FinOps practices and cloud cost optimization at scale.
Experience building or scaling an SRE function from the ground up within an existing engineering organization.