Back to Jobs
United States
Development Just now
Senior Site Reliability Engineer
United StatesFull-time
$175-185k
Senior-Level
Be the first applicant! 🚀
Job Description
Key Skills Required
Master these to land this role
DevOpsBestseller 🔥
Learn in 63 HoursBackendBestseller 🔥
Learn in 18 HoursJavaSpring BootKubernetes
Want to know if you're a match for this job?
About the role:
As the Senior Site Reliability Engineer, you will lead Branch’s effort to achieve greater reliability, performance, scalability, capacity and observability of our platform through automation and tooling. You will also participate in and improve the software development and deployment life cycles as well as develop and implement technical best practices.
Responsibilities include, but are not limited to:
- Partner with Developers to produce high-performing and robust services through rigorous testing and release procedures
- Design infrastructure, monitoring, processes, and standards for systems and applications
- Support services through design, development, load testing, and launch phases
- Develop, measure, and monitor key performance and service level indicators including availability, latency, and overall system health
- Define and establish SLIs, SLOs, and error budgets with service owners, and drive adoption across platform teams
- Profile and optimize platform performance, resilience, and efficiency, including latency, throughput, and capacity planning under load
- Participate in incident response and root cause analysis
- Remediate tasks and develop preventative and automated measures to meet SLAs/SLOs/SLIs
- Manage monitoring services utilized by applications
Qualifications (required):
- Bachelor's degree in an appropriate engineering discipline or equivalent experience required
- 3+ years experience in site reliability engineering
- Strong hands-on experience building and operating Java / Spring Boot services in production
- Experience with Terraform, Go, Java, Gradle, Docker, OpenTelemetry and Kubernetes
Qualifications (preferred):
- Experience in scripting using Python and Bash
- Experience developing and maintaining Kubernetes Operators
- Exposure to Google Pub/Sub, Redis, Prometheus, Grafana, Google Spanner and MySQL
- Experience with JVM performance tuning and profiling
- Experience working on GCP platforms
How would you rate this job post?
See what other professionals think about this role.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.