Back to Jobs
Development Just now

Release Engineer (SRE)

United StatesUnited States
Full-time
Not Disclosed
Senior-Level

Job Description

Key Skills Required

Master these to land this role

BackendBestseller 🔥
Learn in 18 Hours
DevOpsBestseller 🔥
Learn in 63 Hours
QA EngineerBestseller 🔥
Learn in 10 Hours
CybersecurityAutomation Engineer

Want to know if you're a match for this job?

Calculate My Match Score

About the Role

We're looking for a Release Engineer (SRE) to join our Release Engineering team (part of EngOps) — a production-operations expert who brings an SRE mindset to how Supabase ships and runs, making deploys safe, observable, and recoverable at scale.

Release Engineering's scope has grown well beyond build-and-ship: we increasingly own the operational reliability of the systems that deploy and run Supabase. In this role you'll treat our deployment pipelines, pre-production signal, and the control plane itself as production systems — with SLOs, error budgets, and on-call ownership — and you'll be the person teams lean on when reliability is on the line.

This is not a "gatekeeper" role. You'll make the reliable path the easy path: standardising how we deploy, instrumenting what we ship, and ensuring that when something breaks, we detect it quickly and recover quickly.

What You'll Be Responsible For

In this role, you'll:

  • Own the reliability of Supabase's deployment and release systems, and the control plane they run on, against clear SLOs and error budgets
  • Turn pre-production into a trustworthy signal — standardizing and instrumenting today's fragmented, ad-hoc deployment workflows
  • Drive disaster-recovery readiness, including making environments reproducibly deployable from scratch (untangling undocumented secrets, unclear configuration ownership, and circular service dependencies)
  • Build and operate health and SLO monitoring for critical user flows, using synthetic testing to catch regressions before customers do
  • Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents — which account for a large share of our incident load
  • Participate in on-call, lead blameless postmortems, and turn findings into runbooks, alerting, and automation that remove toil
  • Improve deployment observability and auditability — a clear record of what shipped where, when, and by whom
  • Document operational procedures — break-glass paths, access models, and runbooks — so reliability knowledge isn't tribal

Reliability & Operations

  • Define and track SLAs, SLOs, error budgets, and DORA delivery metrics — with meaningful alerting over noise
  • Ensure deployments fail fast and safely when health checks degrade
  • Harden access and break-glass workflows (e.g. scoped self-service) so the right people can act in an incident without unsafe workarounds
  • Partner with product engineering and platform teams to align release practices with reliability and availability targets

You Might Be a Good Fit If You

  • Have 5+ years in SRE, production operations, platform engineering, or release engineering
  • Have operated production systems at scale and carried on-call for them
  • Are fluent in SLAs, SLOs, error budgets, DORA metrics, and operational KPIs — and the observability tooling behind them (Prometheus, Grafana, Alertmanager, or similar)
  • Have led incident response with tooling like incident.io (or PagerDuty / Opsgenie), run blameless postmortems, and driven down MTTD/MTTR
  • Operate confidently on AWS (multiple accounts, IAM, VPC) in production
  • Are comfortable with infrastructure-as-code (Pulumi, Terraform) and Kubernetes
  • Script and automate to eliminate toil rather than absorb it
  • Communicate clearly with both infrastructure specialists and product engineers
  • Thrive in async, globally distributed teams
  • Are comfortable navigating ambiguity and iterating toward better systems over time

How would you rate this job post?

See what other professionals think about this role.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More