We’re building a world of health around every individual — shaping a more connected, convenient and compassionate health experience. At CVS Health®, you’ll be surrounded by passionate colleagues who care deeply, innovate with purpose, hold ourselves accountable and prioritize safety and quality in everything we do. Join us and be part of something bigger – helping to simplify health care one person, one family and one community at a time.
Position Summary:
About the Team
Our SRE team is the execution engine behind the reliability, availability, and performance of distributed store technology powering thousands of retail and pharmacy locations nationwide. We operate across pharmacy platforms, Point of Sale (POS) systems, handheld devices, store servers, dispensing systems, and edge computing infrastructure in hybrid cloud and on-premises environments deployed at fleet scale.
Our engineering philosophy is grounded in five pillars: Detection, Prevention, Recovery, Learning Loops, and Developer Experience (DevX).
Our operating principle is what we call the reliability covenant: our success is not measured by how many incidents we respond to, it is measured by how much reliability capability we transfer to the engineering teams we serve. The goal is development teams that carry reliability ownership independently, not teams that rely on SRE to keep their services running. If you are drawn to building capability that outlasts your direct involvement, this team is built for that purpose.
We track operational toil as an engineering metric, not as a permanent operational reality. Engineers are expected to identify recurring manual work, eliminate it through automation, and document the reduction. Toil accumulation is treated as a reliability risk and a capacity cost.
About the Role
As a Software Engineer — SRE, you are a practitioner-level contributor focused on building and running reliable distributed systems. You work within assigned services and domains implementing observability, improving alerting quality, responding to incidents, writing automation, and contributing to the reliability programs that run across the organization.
Scope:Â Service and task level, you execute with direction and grow toward autonomous ownership.
The environment you are joining
This role exists inside an active SRE transformation. Many of the systems you will monitor, the processes you will contribute to, and the toolchains you will use are being built or significantly improved in parallel with the day-to-day operational work. You will contribute to defining processes as much as following them. Comfort with ambiguity and a bias toward building, not just operating is essential to success in this role. Engineers who thrive here find that environment energizing, not frustrating.
The operating environment includes an edge computing fleet deployed directly inside store locations, unattended nodes that cannot be reached by on-site SRE engineers. This means a deployment or configuration change that goes wrong can simultaneously affect thousands of locations. You will develop a fleet operations mindset alongside a service reliability mindset: blast radius is geographic, not just functional.
What You Will Do
Detection & Observability
- Implement alerting and dashboards using Prometheus, Grafana, Loki, Jaeger, and OpenTelemetry to provide visibility into distributed retail and pharmacy systems
- Contribute to Service Level Indicator (SLI) and Service Level Objective (SLO) definition for assigned services under Senior(s) guidance; learn error budget mechanics and how burn rate translates to patient and customer impact
- Build and maintain service dashboards covering the golden signals: availability, latency, error rate, and saturation anchored to Critical User Journey (CUJ) outcomes, not just infrastructure metrics
- Participate in alert tuning exercises; document false-positive patterns and noise sources with enough specificity to enable SSE-level remediation
- Learn the principles of anomaly-based detection: understand the difference between threshold alerting and time-series baseline deviation, and how ML-generated signals differ from rule-based alerts
Prevention & Reliability Engineering
- Participate in Production Readiness Reviews (PRR); execute assigned checklist items, contribute findings, and understand the rationale behind each gate
- Write unit and integration tests for SRE tooling and automation scripts at production quality
- Follow and actively improve existing operational runbooks; flag gaps, missing failure modes, outdated steps, and ambiguous procedures for remediation, don't just identify, propose the fix
- Support performance testing and reliability audits under SSE direction; develop hands-on methodology exposure alongside execution skills
- Apply standard reliability engineering patterns: circuit breakers, retries with exponential backoff, timeouts, and bulkhead isolation etc in code and configuration contributions
- Build awareness of fleet-scale deployment risk: understand cohort-based rollout strategies, deployment blast radius concepts, and why configuration drift is a first-order reliability concern in unattended device fleets
Incident Response & Recovery
- Participate in on-call rotations; escalate promptly per defined escalation paths, speed of escalation is as important as technical accuracy at this level
- Execute validated runbooks during incidents; document all actions, decisions, and timeline entries in real time within incident management systems
- Triage and categorize incoming alerts; route to appropriate domain owners with context, an initial hypothesis, and a timeline of observed signals
- Attend post-incident reviews; capture assigned action items and drive them to closure not just acknowledgment
- Build operational familiarity with the edge computing fleet: Kubernetes clusters at store level, store servers, networking components, and the connectivity failure modes specific to unattended infrastructure
Learning Loops & Continuous Improvement
Learning Loops at this level is not documentation, it is closing the loop. The measure of a good postmortem contribution is not a well-written timeline. It is a subsequent monitoring or alerting change that reduces the recurrence of a specific incident type.
- Attend and actively contribute to postmortems; document timeline details, contributing factors, and observations that would otherwise be lost
- Apply postmortem findings directly: after each postmortem, own at least one monitoring, alerting, or runbook improvement that reduces the likelihood or detection time of the same incident class recurring
- Maintain assigned runbooks and operational documentation in the team's Confluence or wiki space; review for accuracy quarterly and after every related incident
- Complete a structured SRE learning path: SLO fundamentals, incident command, chaos engineering concepts, and retail/pharmacy domain knowledge for the systems you operate
- Track personal MTTD (Mean Time to Detect) and MTTR (Mean Time to Restore) trends across your on-call rotations; use the data to identify your own skill gaps and improvement targets
Developer Experience & Automation (DevX)
- Eliminate toil: write automation scripts in Python, Bash, or Go to remove recurring manual operational tasks. For every significant manual task you perform more than twice, your default question should be "why isn't this automated yet?" and your default action should be to fix it
- Maintain a personal record of toil eliminated per quarter, volume of manual touchpoints removed, time recovered, and error modes eliminated
- Contribute to CI/CD pipeline improvements using GitHub Actions, ArgoCD, and Helm; apply GitOps practices for configuration and deployment automation
- Participate in DORA metrics baseline activities; track deployment frequency and lead time for assigned services as indicators of delivery health
- Collaborate with development teams on instrumentation and observability integration during feature development, reliability should be designed in, not added after deployment
- Learn and apply containerization and cloud-native deployment patterns using Kubernetes, Helm, and infrastructure-as-code tools such as Terraform or Ansible
​
Required Qualifications
- 2+ years of experience in SRE, DevOps, platform engineering, or related technology roles with production systems responsibility
- 2+ years of experience delivering software in large-scale distributed environments with hands-on application of reliability and resilience concepts
- 1+ year of experience with at least one programming language at production quality: Python, Go, Bash, or Java
- 1+ year of hands-on cloud platform experience: AWS, Microsoft Azure, or Google Cloud Platform (GCP)
- Practical experience with observability and monitoring tools such as Prometheus, Grafana, ELK, Splunk, Datadog, or Dynatrace
- Foundational understanding of containerization and orchestration with Kubernetes and Docker. Experience with AI-assisted tooling and development.
- Comfort contributing to processes that are being defined, not just processes that already exist, ability to operate productively in an environment under active transformation
- Strong written and verbal communication skills; ability to engage both technical and non-technical stakeholders in incident and postmortem contexts
Preferred Qualifications
- Experience supporting retail, pharmacy, healthcare, or other distributed-fleet systems at scale
- Exposure to incident management, change management, and problem management processes (ITIL familiarity a plus)
- Familiarity with time-series anomaly detection concepts; understanding of how ML-generated signals differ from rule-based alerting
- Experience with CI/CD pipeline tooling: GitHub Actions, Jenkins, ArgoCD, or CircleCI
- Exposure to microservices architecture, service mesh (Istio, Linkerd), and cloud-native distributed systems patterns
- Basic experience with infrastructure-as-code: Terraform, Ansible, or Pulumi
- Experience with fleet-scale or edge computing deployments where unattended node management is a reliability concern
What Success Looks Like at 6 Months
- You are on-call as a primary/secondary responder, escalating with speed and context, and executing runbooks without hand-holding
- You own and maintain dashboards and SLOs for at least two assigned services, with documented alert quality improvements (fewer false positives, faster detection of real issues)
- You have participated in at least three postmortems and closed every action item assigned to you, with at least one resulting in a monitoring or alerting improvement that reduced recurrence of that incident type
- You have automated at least three recurring manual tasks and documented the before/after reduction in operational toil
- Your peers describe you as someone who closes loops. not someone who documents problems and waits for someone else to fix them
Education:
- Bachelor's degree in Computer Science, Engineering, or a related field — or equivalent practical experience
Anticipated Weekly Hours
40Time Type
Full timePay Range
The typical pay range for this role is:
$72,100.00 - $158,620.00This pay range represents the base hourly rate or base annual full-time salary for all positions in the job grade within which this position falls. Â The actual base salary offer will depend on a variety of factors including experience, education, geography and other relevant factors. Â This position is eligible for a CVS Health bonus, commission or short-term incentive program in addition to the base pay range listed above.Â
Â
Our people fuel our future. Our teams reflect the customers, patients, members and communities we serve and we are committed to fostering a workplace where every colleague feels valued and that they belong.
Great benefits for great people
We take pride in offering a comprehensive and competitive mix of pay and benefits that reflects our commitment to our colleagues and their families.
Additional details about available benefits are provided during the application process and on Benefits Moments.
Qualified applicants with arrest or conviction records will be considered for employment in accordance with all federal, state and local laws.
Learn more about this Employer on their Career Site
