Back to Jobs
Garnerhealth
Development 4h ago

Senior Site Reliability Engineer

Garnerhealth
United StatesUnited States
Full-time
$191,000 - $226,000
Senior-Level

Job Description

Key Skills Required

Master these to land this role

Python2h 41mFree Trial ✨
Start 10-Day Free Trial
DevOps1h 38mFree Trial ✨
Start 10-Day Free Trial
AWSKubernetesTerraform

Want to know if you're a match for this job?

Calculate My Match Score

We are seeking a Senior Site Reliability Engineer to own the reliability, performance, and resilience of the cloud infrastructure powering Garner’s products and AI/ML workloads. This role sits on our Platform Engineering team. You will run the machine: defining and upholding SLOs, leading incident response, and driving the automation and standards that let every Garner engineer ship faster and more reliably. Because our systems directly influence health outcomes for millions of patients, maintaining the highest standards of production quality is imperative. This is an automation-first role: you will use AI tools to continuously convert manual operational work into monitored, hands-free processes, so the role gets more leveraged as you build.

What you will do:

  • Run the Machine: Own the end-to-end reliability, performance, and resilience of Garner’s cloud environments (AWS, Kubernetes), including those powering AI/ML workloads; define, measure, and uphold SLOs across our critical services
  • Lead Incident Response: Serve in the on-call rotation, lead incident response, and drive deep-dive root cause analysis, seeing corrective actions through to resolution and rigorously reviewing infrastructure changes
  • Own Observability: Build and maintain the monitoring, alerting, and observability systems that let us detect and resolve issues before users feel them
  • Scale & Optimize: Translate ambiguous, high-performance scaling requirements into well-defined, automated, and composable infrastructure-as-code deliverables (Terraform); proactively identify and implement cost-efficiency and performance gains across the stack to maximize cloud ROI
  • Automate Away Toil: Pay down impactful tech debt and reduce operational toil, using AI tools and automation to convert repetitive operational work into hands-free, monitored processes, and holding our internal platform to the same rigorous standards as our customer-facing products
  • Enable Engineering: Build and maintain the deployment and observability standards that empower the broader engineering team to ship AI features faster and more reliably; communicate complex cloud and reliability concepts clearly to technical and non-technical stakeholders
  • Uphold Security & Compliance: Ensure our infrastructure and operations meet Garner’s security and HIPAA compliance obligations

The ideal candidate has:

  • 4+ years of hands-on experience operating production cloud infrastructure at scale in an SRE, DevOps, or platform engineering role
  • Deep expertise with Kubernetes and Terraform in a cloud-first environment (AWS preferred)
  • A strong track record with production observability: defining SLOs, building monitoring and alerting, and leading incident response and blameless post-incident reviews
  • Strong software engineering fundamentals in Python or Go, applied to infrastructure automation (experience with Kubernetes APIs a plus)
  • Experience driving cloud cost-efficiency and performance optimization across compute, storage, and networking
  • Experience supporting AI/ML or data-intensive workloads in production is a plus
  • Experience operating in a security-conscious or regulated environment (HIPAA, SOC 2) is a plus
  • Fluency with AI tools (e.g., Claude) applied to real engineering and operations workflows, or strong motivation to build it fast
  • A desire to be a part of a high-performing, mission-driven team that operates with intense urgency, a strong sense of individual accountability, and a commitment to authentic feedback

Technologies we use:

  • AWS, Kubernetes, Terraform, Istio, Python, Go, TypeScript, Postgres, NATS, Datadog, GitLab

This is a unique opportunity to join a fast-growing company in a transformative role, helping shape the future of healthcare.

How would you rate this job post?

See what other professionals think about this role.

banner

Garnerhealth is a forward-thinking healthcare technology company dedicated to revolutionizing the way medical professionals and patients interact with health data. With a strong foundation in data analytics and cutting-edge technology, Garnerhealth aims to provide innovative solutions that enhance patient outcomes, streamline clinical workflows, and improve the overall quality of care. The company's mission is to empower healthcare providers with actionable insights, facilitating informed decision-making and personalized medicine. By harnessing the power of artificial intelligence, machine learning, and data science, Garnerhealth is poised to make a significant impact in the healthcare industry, driving progress and excellence in medical care.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More
Senior Site Reliability Engineer at Garnerhealth | HireSkys