Back to Jobs
Garnerhealth
Development 10h ago

Staff Site Reliability Engineer

Garnerhealth
United StatesUnited States
Full-time
$241,000 - $270,000
Senior-Level

Job Description

Key Skills Required

Master these to land this role

DevOps1h 38mFree Trial ✨
Start 10-Day Free Trial
Python2h 41mFree Trial ✨
Start 10-Day Free Trial
AI EngineerKubernetesAWS

Want to know if you're a match for this job?

Calculate My Match Score

About the role:

We are seeking an exceptional Staff Site Reliability Engineer to own the reliability strategy for the cloud infrastructure powering Garner’s products and AI/ML workloads. This role sits on our Platform Engineering team. As the most senior reliability voice in the organization, you will set the technical direction for how Garner defines, measures, and upholds production quality, architecting the SLO framework, incident response program, and automation standards that every engineering team builds on. Because our systems directly influence health outcomes for millions of patients, maintaining the highest standards of production quality is imperative. This is an automation-first role: you will use AI tools to continuously convert manual operational work into monitored, hands-free processes, and build the platform that lets every Garner engineer do the same.

What you will do:

  • Own the Reliability Strategy: Architect and own the end-to-end reliability, performance, and resilience of Garner’s cloud environments (AWS, Kubernetes), including those powering AI/ML workloads; design the SLO framework our critical services are measured against and lead the technical decision-making that keeps us ahead of scale
  • Lead the Incident Response Program: Set the standard for how Garner responds to incidents: serve in and level up the on-call rotation, lead response for the most complex escalations, drive deep-dive root cause analysis, and build the review culture that sees corrective actions through to resolution
  • Own Observability: Architect the monitoring, alerting, and observability platform that lets us detect and resolve issues before users feel them, and that lets stakeholders quickly identify the health of every team’s products
  • Translate Ambiguity: Take high-level, ambiguous scaling and reliability requirements and transform them into well-defined, automated, and composable infrastructure-as-code deliverables (Terraform); proactively identify and implement cost-efficiency and performance gains across the stack to maximize cloud ROI
  • Course-Correct Technical Direction: Proactively identify when infrastructure workflows or technical paths are inefficient or fragile and redirect efforts to ensure the highest ROI for the engineering function, paying down impactful tech debt and using AI tools and automation to convert repetitive operational work into hands-free, monitored processes
  • Multiply the Engineering Team: Build and own the deployment and observability standards that empower the broader engineering team to ship AI features faster and more reliably; mentor engineers across the organization and provide high-quality feedback that raises the bar for operational rigor and discipline
  • Uphold Security & Compliance: Ensure our infrastructure and operations meet Garner’s security and HIPAA compliance obligations, and lead rigorous review of infrastructure changes so platform work meets the same standards as our customer-facing products

The ideal candidate has:

  • 7+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role
  • Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred), with a track record of architecting reliability for systems at scale
  • Experience designing an organization’s reliability practice (SLO frameworks, observability platforms, incident response programs, and blameless post-incident reviews) and the judgment to know when to build vs. buy
  • Strong Python or Go skills applied to infrastructure automation (Kubernetes API experience a plus)
  • Track record driving cloud cost-efficiency and performance optimization across compute, storage, and networking
  • Mentorship experience and the ability to set technical direction as the senior reliability voice
  • Excellent communication skills—able to make complex reliability concepts land with both technical and non-technical stakeholders
  • Fluency with AI tools (e.g., Claude) applied to real engineering and operations workflows, or strong motivation to build it fast
  • Experience supporting AI/ML or data-intensive workloads in production is a plus
  • Experience operating in a security-conscious or regulated environment (HIPAA, SOC 2) is a plus
  • A desire to be a part of a high-performing, mission-driven team that operates with intense urgency, a strong sense of individual accountability, and a commitment to authentic feedback

Technologies we use:

  • AWS, Kubernetes, Terraform, Istio, Python, Go, TypeScript, Postgres, NATS, Datadog, GitLab

This is a unique opportunity to join a fast-growing company in a transformative role, helping shape the future of healthcare.

How would you rate this job post?

See what other professionals think about this role.

banner

Garnerhealth is a forward-thinking healthcare technology company dedicated to revolutionizing the way medical professionals and patients interact with health data. With a strong foundation in data analytics and cutting-edge technology, Garnerhealth aims to provide innovative solutions that enhance patient outcomes, streamline clinical workflows, and improve the overall quality of care. The company's mission is to empower healthcare providers with actionable insights, facilitating informed decision-making and personalized medicine. By harnessing the power of artificial intelligence, machine learning, and data science, Garnerhealth is poised to make a significant impact in the healthcare industry, driving progress and excellence in medical care.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More
Staff Site Reliability Engineer at Garnerhealth | HireSkys