Site Reliability Engineer III (DBA) at Backblaze
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Backblaze is the object storage leader in the open cloud movement, fueling customer success with cloud storage built purposefully to unlock budgets, unburden administrators, and unleash innovators.
Founded in 2007, Backblaze scaled the business with less than $3 million in outside funding until 2021, when it went public on the Nasdaq stock exchange. Today, Backblaze generates over $100M in revenue and is the leading specialized storage cloud, managing over three billion gigabytes of data storage for 500K+ customers in 175+ countries.
About the Role:
Individuals fulfilling this role will be responsible for ensuring the stability, scalability, and reliability of our production database systems, primarily Vitess (distributed MySQL) and Cassandra, alongside the rest of our production services and infrastructure. This role carries the same on-call, incident response, and service ownership expectations as other SRE IIIs, with database systems serving as the area of deepest technical ownership.
Because our SRE Database Engineering function is new, this role will also help establish its operational foundation by designing database architecture and developing the runbooks, escalation guidance, procedures, and training materials that our Level 1 and Level 2 SRE Database Engineers will use as they onboard.
What You'll Do:
- Database Architecture & Administration:
- Design, deploy, and own highly available database architecture for Vitess (distributed MySQL) and Cassandra
- Establish and document operational procedures, runbooks, and escalation guidance for Level 1 and Level 2 SRE Database Engineers
- Optimize database performance through query tuning, indexing strategies, schema design, and capacity planning
- Own database backup, recovery, replication, and disaster recovery strategies
- Perform and validate disaster recovery testing and database recovery procedures
- Drive database security, access control, patching, hardening, and compliance practices
- Partner with the DBA and Data Infrastructure teams on resharding, capacity planning, replication, and architecture decisions for sharded MySQL environments
- Service Reliability & Operations:
- Support the availability and durability of critical services across production environments
- Monitor service health using SLIs, SLOs, error budgets, monitoring, logging, and alerting platforms
- Partner with service owners to define and improve SLIs, SLOs, error budget policies, and alerting
- Participate in on-call rotations, incident response, root cause analysis, and post-incident reviews
- Serve as an escalation point for complex database production incidents
- Follow established ITIL/OSS processes including incident, change, problem, and capacity management
- Take ownership of operational issues and drive projects from problem discovery through resolution
- Automation & Tooling:
- Develop automation for common operational and database administration tasks to reduce manual intervention and operational toil
- Contribute to monitoring, logging, and alerting frameworks including Prometheus, Grafana, Catchpoint, and ELK
- Help integrate operational runbooks and incident response workflows with FireHydrant
- Work with CI/CD pipelines, configuration management, and infrastructure as code tools including Terraform, Ansible, and Jenkins
- Develop scripts using Bash, Python, Go, or similar technologies to improve reliability and operational efficiency
- Operate and troubleshoot containerized production environments using Kubernetes and Docker
- Work within Kubernetes and Vitess environments using technologies such as kubectl, mysqlsh, and Vitess keyspaces
- Project Management:
- Lead Production Readiness Reviews (PRRs) for functionality being handed off from engineering partner teams
- Support the operational readiness of new database-backed services before they enter production
- Build training plans, onboarding materials, and technical documentation for new Level 1 and Level 2 SRE Database Engineers
- Partner with Engineering, Product, Operations, and DBA/Data Infrastructure teams on reliability initiatives
- Assist with capacity planning, disaster recovery exercises, database migrations, and infrastructure projects
- Work with vendors and service providers to troubleshoot service issues and track SLA performance
- Identify opportunities for automation and process efficiency
- Incident Response:
- Respond to and resolve production database, infrastructure, and service incidents
- Troubleshoot and escalate database, Linux, networking, application, and infrastructure issues as needed
- Participate in the on-call rotation and serve as an escalation point for database-related incidents
- Lead or contribute to root cause analysis and post-incident reviews
- Identify recurring issues and develop long-term corrective actions to improve reliability
- What We Value:
- A proactive mindset with a can-do attitude
- Someone who can work independently, take ownership, and drive complex technical problems through resolution
- Someone who steps up, supports teammates, mentors others, and shares knowledge freely
- Strong problem-solving skills and a willingness to learn new technologies
- Curiosity, reliability, and a desire to improve the reliability and scalability of production systems
Required Qualifications:
- 6–8 years of experience in site reliability engineering, systems engineering, infrastructure operations, database engineering, or similar roles, with meaningful experience supporting production database systems.
- Deep hands-on experience with MySQL and distributed or sharded database systems.
- Experience with Vitess in a production environment strongly preferred.
- Experience administering and supporting NoSQL databases such as Cassandra.
- Experience designing high-availability database architecture, replication topology, backup strategies, and disaster recovery processes.
- Strong SQL skills, including query performance analysis, indexing, schema design, and troubleshooting.
- Solid Linux systems administration and troubleshooting skills.
- Experience with security-focused operations including patching, system hardening, access controls, and vulnerability remediation.
- Strong understanding of service reliability concepts including monitoring, alerting, incident response, root cause analysis, SLIs, SLOs, and error budgets.
- Experience working with containers and orchestration platforms including Kubernetes and Docker.
- Comfortable operating in Kubernetes and Vitess environments using tools such as kubectl, mysqlsh, and Vitess keyspaces.
- Experience with infrastructure and configuration management technologies including Terraform, Ansible, Jenkins, and HashiCorp products such as Vault and Nomad.
- Proficiency in at least one scripting language such as Python, Bash, or Go.
- Experience establishing operational procedures, runbooks, documentation, and escalation processes.
- Experience mentoring, training, or helping onboard engineers into complex technical environments.
- Experience in SaaS, cloud services, service provider, or large-scale distributed systems environments preferred.
- Experience with AWS, GCP, Azure, or similar cloud platforms preferred.
- Familiarity with ITIL/OSS practices and SLA/SLO management preferred.
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent professional experience.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
Train and Evaluate AI Agents in CAD Environments (Freelance)
Mindrift
United StatesTrain and Evaluate AI Agents in CAD Environments (Freelance)
Mindrift
United StatesSenior Engineer – DoD/U.S. Navy Energetics Facility Design and Construction
Eastern Research Group
United StatesSecurity Operations Specialist
HiddenLayer
United StatesMore Openings at Backblaze
Explore Top Companies in this Space
MinIO
Cloud Storage / Data Infrastructure / Enterprise Software
NetApp
Cloud Data Infrastructure / Hybrid Cloud Storage & Data Management / Software-Defined Storage & CloudOps
Qumulo
Enterprise Software / Cloud Computing / Data Storage / AI Infrastructure
BlueFlag Security
Enterprise Software / Cybersecurity / Developer Tools / AI Governance
Backblaze
View Company ProfileBackblaze is a leading, independent cloud storage and data backup platform fundamentally engineered to offer enterprise-grade performance at a fraction of the cost of legacy hyperscalers like AWS. Founded in 2007 and operating as a remote-first, publicly traded company (NASDAQ: BLZE) originally rooted in San Mateo, California, the company completely disrupts the complex and expensive cloud infrastructure market. Under the hood, Backblaze provides incredibly resilient, S3-compatible object storage (B2 Cloud Storage) and unlimited computer backup solutions, aggressively stripping away unpredictable egress fees, hidden storage tiers, and vendor lock-in. Their primary target audience spans aggressive tech startups, massive media & entertainment studios requiring high-throughput data pipelines, and IT leaders who desperately need highly scalable, ransomware-resilient infrastructure. What sets Backblaze apart in the heavily monopolized cloud ecosystem is its radical transparency and specialized hardware/software architecture—allowing them to actively manage over an exabyte of data for more than 500,000 global customers while consistently delivering cloud storage at just one-fifth the cost of Amazon S3.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.


