Site Reliability Engineer (SRE)
NOVACARD · Remote job · posted Sep 14, 2026
Open to candidates worldwide
More remote jobs open to candidates in India, Vietnam, Indonesia, Pakistan, Nigeria, Bangladesh, Malaysia, Philippines, Brazil, Colombia, Kenya, Ghana, Egypt, South Africa, Mexico, Argentina, Turkey, Ukraine, Sri Lanka and Nepal
You apply on the company's own site. We never charge to apply.
What this role actually asks for
Extracted by RemoteHuntMust have
- •Ensure stability, performance, and fault tolerance
- •Develop and maintain infrastructure automation
- •Improve observability
- •Monitor system health and respond to incidents
- •Perform root cause analysis
- •Define and manage SLIs, SLOs, and Error Budgets
Nice to have
- •Collaborate with development teams
- •Integrate and monitor external vendor systems
Tools and technologies
The full posting
At NOVACARD , we’re redefining how people use credit. We are the first interest-free and no-annual-fee credit card in Mexico , designed to simplify personal finances and give users complete control - all from a mobile app. With NOVACARD, users can access up to $200,000 MXN in credit , only pay when they use it, and manage everything digitally in under 5 minutes.
Our mission is to empower people to make smarter financial decisions by offering flexibility, transparency, and the freedom they need to reach their goals. Simple finances, big goals.
About the Role
We’re looking for a Site Reliability Engineer (SRE) to ensure the stability, performance, and reliability of our critical production systems. You’ll work at the intersection of development and operations — building automation tools, improving observability, and preventing incidents before they occur.
Key Responsibilities
Ensure the stability, performance, and fault tolerance of production systems. Develop and maintain infrastructure automation and observability tools. Monitor system health, respond to incidents, and perform root cause analysis (RCA).
Collaborate with development teams to improve scalability and reliability of services. Define and manage SLIs , SLOs , and Error Budgets . Lead incident response: organize recovery, document RCA, and run blameless post-mortems.
Configure and administer Grafana and Zabbix , design insightful dashboards, and fine-tune alerting. Integrate and monitor external vendor systems, collaborating with vendor technical support when needed.
Similar remote jobs
- Senior SRE/ML Ops Engineer - PerformanceVibe
- Staff Software Engineer - Databases SRE | UK | RemoteGrafanalabs
- Site Reliability Engineer (Contract, Rotation-Based)Invisibletech
- Database Support Engineer (APAC)Supabase
- Senior Site Reliability Engineer (SRE)Mirantis
- Incident/Change Governance & Control (ICGC) Engineer | RemoteGSB Solutions
Is this one actually worth your time?
RemoteHunt scores every remote job 0–100 against your own resume, so you apply to the handful that fit instead of the hundred that don't. Free plan, no card required.