Release Engineer (SRE)
United StatesJob Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
About the Role
We're looking for a Release Engineer (SRE) to join our Release Engineering team (part of EngOps) — a production-operations expert who brings an SRE mindset to how Supabase ships and runs, making deploys safe, observable, and recoverable at scale.
Release Engineering's scope has grown well beyond build-and-ship: we increasingly own the operational reliability of the systems that deploy and run Supabase. In this role you'll treat our deployment pipelines, pre-production signal, and the control plane itself as production systems — with SLOs, error budgets, and on-call ownership — and you'll be the person teams lean on when reliability is on the line.
This is not a "gatekeeper" role. You'll make the reliable path the easy path: standardising how we deploy, instrumenting what we ship, and ensuring that when something breaks, we detect it quickly and recover quickly.
What You'll Be Responsible For
In this role, you'll:
- Own the reliability of Supabase's deployment and release systems, and the control plane they run on, against clear SLOs and error budgets
- Turn pre-production into a trustworthy signal — standardizing and instrumenting today's fragmented, ad-hoc deployment workflows
- Drive disaster-recovery readiness, including making environments reproducibly deployable from scratch (untangling undocumented secrets, unclear configuration ownership, and circular service dependencies)
- Build and operate health and SLO monitoring for critical user flows, using synthetic testing to catch regressions before customers do
- Reduce mean-time-to-detect and mean-time-to-recover for deploy-related incidents — which account for a large share of our incident load
- Participate in on-call, lead blameless postmortems, and turn findings into runbooks, alerting, and automation that remove toil
- Improve deployment observability and auditability — a clear record of what shipped where, when, and by whom
- Document operational procedures — break-glass paths, access models, and runbooks — so reliability knowledge isn't tribal
Reliability & Operations
- Define and track SLAs, SLOs, error budgets, and DORA delivery metrics — with meaningful alerting over noise
- Ensure deployments fail fast and safely when health checks degrade
- Harden access and break-glass workflows (e.g. scoped self-service) so the right people can act in an incident without unsafe workarounds
- Partner with product engineering and platform teams to align release practices with reliability and availability targets
You Might Be a Good Fit If You
- Have 5+ years in SRE, production operations, platform engineering, or release engineering
- Have operated production systems at scale and carried on-call for them
- Are fluent in SLAs, SLOs, error budgets, DORA metrics, and operational KPIs — and the observability tooling behind them (Prometheus, Grafana, Alertmanager, or similar)
- Have led incident response with tooling like incident.io (or PagerDuty / Opsgenie), run blameless postmortems, and driven down MTTD/MTTR
- Operate confidently on AWS (multiple accounts, IAM, VPC) in production
- Are comfortable with infrastructure-as-code (Pulumi, Terraform) and Kubernetes
- Script and automate to eliminate toil rather than absorb it
- Communicate clearly with both infrastructure specialists and product engineers
- Thrive in async, globally distributed teams
- Are comfortable navigating ambiguity and iterating toward better systems over time
How would you rate this job post?
See what other professionals think about this role.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.