Senior Site Reliability Specialist II
Everbridge · United States · posted Aug 12, 2026
What this role actually asks for
Extracted by RemoteHuntMust have
- •Design and operate complex production systems
- •Cloud infrastructure and cloud-native architectures
- •Distributed systems and container platforms
- •Infrastructure as Code and automation
- •CI/CD and software delivery practices
- •Observability, monitoring, logging, and telemetry
- •Incident response and operational excellence
Nice to have
- •Experience in regulated environments (FedRAMP, DoD, IL4/IL5, SOC 2, ISO 27001)
Tools and technologies
The full posting
At Everbridge, reliability isn’t just about uptime—it’s about ensuring that critical systems are available when they matter most. Every improvement you make helps organizations deliver life-saving communications and maintain operations during emergencies. As a Senior Site Reliability Engineer II, you’ll do more than operate infrastructure. You’ll improve the resiliency of our engineering organization by building reliable platforms, eliminating operational toil, mentoring engineers, and helping teams design systems that are secure, scalable, and resilient by default. This is a highly collaborative technical leadership role for an engineer who enjoys solving systemic problems, influencing architecture, and enabling others to build reliable software. What you'll do:: Build platform capabilities that enable engineering teams to deliver reliable software safely and efficiently. Lead complex technical initiatives spanning cloud infrastructure, Kubernetes, observability, automation, networking, and developer platforms. Influence engineering decisions through technical expertise, collaboration, and data. Help engineering teams become increasingly self-sufficient through coaching, automation, and well-designed platform capabilities. Continuously improve operational excellence by reducing complexity, eliminating manual work, and strengthening engineering practices. Design and implement solutions that improve the availability, scalability, performance, and resilience of our platform. Build automation that eliminates repetitive operational work and reduces engineering toil. Improve observability, monitoring, alerting, and operational readiness across the organization. Use production data, reliability metrics, and engineering judgment to identify systemic improvements. Lead complex cross-functional engineering initiatives from design through production. Partner with architects, software engineers, security, product, and platform teams to build resilient systems from the beginning. Review architectures and designs with a focus on reliability, scalability, recoverability, and operational excellence. Establish and evolve engineering standards, best practices, and operational readiness guidance. Partner directly with engineering teams to improve the reliability of the services they own. Coach teams on observability, incident response, disaster recovery, capacity planning, and production readiness. Help teams adopt SLOs, error budgets, meaningful alerting, and engineering practices that improve customer outcomes. Make the right engineering decisions easier through automation, paved roads, and self-service capabilities. Participate in an on-call rotation supporting critical production systems. Lead the technical response during high-severity incidents. Facilitate blameless post-incident reviews that focus on learning and systemic improvement. Drive corrective actions through completion and measure their effectiveness over time. Share knowledge through documentation, technical design reviews, and collaborative problem solving. Foster a culture of ownership, continuous improvement, operational excellence, and customer focus. What you'll bring:: Experience with: Designing and operating complex production systems. Cloud infrastructure and cloud-native architectures. Distributed systems and container platforms. Infrastructure as Code and automation. CI/CD and software delivery practices. Observability, monitoring, logging, and telemetry. Incident response and operational excellence. Reliability engineering principles including SLOs, SLIs, capacity planning, and performance optimization. Writing software or automation using one or more modern programming languages. Linux and networking fundamentals. Experience working within regulated environments such as FedRAMP, DoD, IL4/IL5, SOC 2, or ISO 27001 is a plus.
Is this one actually worth your time?
RemoteHunt scores every remote job 0–100 against your own resume, so you apply to the handful that fit instead of the hundred that don't. Free plan, no card required.