SonicJobs Logo
Left arrow iconBack to search

Senior Site Reliability Engineer

Inspire
Posted 2 days ago, valid for 13 days
Location

Atlanta, GA, US

Salary

Competitive

Contract type

Full Time

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.

Sonic Summary

info
  • Inspire Brands is seeking two Senior Site Reliability Engineers to enhance the reliability of high-traffic digital platforms.
  • Candidates should have a minimum of 5 years of experience in Site Reliability Engineering or related fields, along with 2 years of experience with Kubernetes.
  • The role involves defining SLIs, SLOs, and implementing observability solutions while leading incident management and automation efforts.
  • Strong programming skills in languages such as Python or Go, along with expertise in cloud platforms like AWS or Azure, are required.
  • The position is based in Atlanta with an expected on-site presence of 80%, but the salary details are not specified.

Inspire Brands is hiring two Senior Site Reliability Engineers to help build and scale reliable, resilient, and observable systems supporting high-traffic, customer-facing digital platforms. These role blends software engineering, systems thinking, and operational excellence to reduce toil, prevent incidents, and improve system reliability at scale.

The ideal candidate has hands-on experience applying and implementing SRE principles — not just supporting production systems, but engineering reliability into them.

RESPONSIBILITIES

Reliability Engineering

  • Define and manage SLIs, SLOs, and Error Budgets for critical services
  • Drive production readiness reviews and reliability requirements into architecture and design
  • Perform capacity planning, failure mode analysis, and dependency risk assessments
  • Identify systemic reliability risks and drive remediation before they cause customer impact

Observability

  • Design monitoring, alerting, logging, and tracing solutions using modern observability tooling
  • Improve signal-to-noise ratio and reduce alert fatigue
  • Build dashboards and telemetry that reflect true service health, not just infrastructure metrics

Incident Management

  • Lead technical response for high-severity incidents
  • Drive blameless postmortems and root cause analysis focused on systemic fixes
  • Continuously improve detection, response, and recovery processes
  • Participate in an on-call rotation

Automation & Toil Reduction

  • Identify and eliminate manual, repetitive operational work through automation
  • Build self-healing systems, tooling, and scripts to reduce human intervention
  • Improve CI/CD pipelines and deployment safety (canary, rollback, blue-green)
  • Support Infrastructure as Code (Terraform, Bicep, or similar)

Performance & Scalability

  • Conduct load testing, performance benchmarking, and bottleneck analysis
  • Partner with engineering to design systems for horizontal scalability and fault tolerance

Collaboration & Culture

  • Partner with engineering teams to implement resiliency patterns (circuit breakers, retries, graceful degradation, rate limiting)
  • Mentor engineers on SRE best practices
  • Promote a culture of engineering-driven reliability over reactive operations

EDUCATION AND EXPERIENCE QUALIFICATIONS

Required Qualifications 

  • 5+ years experience in Site Reliability Engineering, Software Engineering, or Platform Engineering
  • 2+ years experience with Kubernetes and containerized workloads
  • 4-year degree in Computer Science or related field

Preferred Qualifications 

  • Experience with chaos engineering or resiliency testing
  • Experience with high-volume, high-availability transactional systems
  • Experience with AI-assisted observability or operational automation
  • Experience making meaningful contributions to internal SRE tooling, frameworks, or platforms

REQUIRED KNOWLEDGE, SKILLS, OR ABILITIES

  • Strong programming/scripting skills (Python, Go, Java, or Node.js)
  • Demonstrated experience defining and operating against SLOs/Error Budgets
  • Strong skills in leading incident response and root cause analysis for production systems
  • Solid understanding of distributed systems and microservices architecture
  • Deep knowledge and expertise in at least one major cloud platform (Azure, AWS, or GCP)
  • Expertise with observability platforms and monitoring strategy

This position is based in our Atlanta Support Center, with an expected on-site presence of 80%.


 

Inspire is a multi-brand restaurant company whose portfolio includes more than 33,300 Arby’s, Baskin-Robbins, Buffalo Wild Wings, Dunkin’, Jimmy John’s, and SONIC restaurants worldwide. We’re made up of some of the world’s most iconic restaurant brands, but we’re much more than just a restaurant company. We’re a team of hundreds of thousands who individually and collectively are changing the way people eat, drink, and gather around the table. We know that food is much more than a staple—it’s an experience. At Inspire, that’s our purpose: to ignite and nourish flavorful experiences.



Learn more about this Employer on their Career Site

Apply now in a few quick clicks

By applying, a Sonicjobs account will be created for you. Sonicjobs's Privacy Policy and Terms & Conditions will apply.

SonicJobs' Terms & Conditions and Privacy Policy also apply.