Site Reliability Engineer
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
TherapyNotes is the go-to superhero for behavioral health Practice Management and EHR software! Our top-notch SaaS solution handles scheduling, billing, documenting, telehealth, and more so clinicians can focus on awesome patient care.
We're a dynamic team of pros who love to innovate and push the envelope, keeping our software cutting-edge. Join us, and let's revolutionize behavioral health software together while making a real difference!
The Role
We are seeking a Site Reliability Engineer to improve the reliability and operability of the production services and shared platforms supporting our growing 24×7 SaaS environment. In this role, you will apply software and systems engineering practices to improve availability, performance, scalability, resilience, observability, incident response, and operational automation. You will partner with software development, infrastructure, database, security, and other technology teams to establish measurable reliability goals, reduce operational toil, and ensure services are supportable throughout their lifecycle. If you are passionate about building reliable systems, solving complex production problems, and driving continuous improvement, we want to hear from you.
Requirements
- BS degree in Information Systems, Engineering, or equivalent experience.
- 5+ years of engineering experience in Systems Engineering, Cloud or Platform Engineering, DevOps, Software Engineering, and/or SRE.
- Experience designing and operating production systems using cloud-based compute, storage, networking, and containerization technologies; Azure and Kubernetes preferred.
- Strong Linux systems and networking fundamentals, with experience troubleshooting complex distributed systems in production.
- Expertise with an observability platform; Datadog experience strongly preferred. Experience with Prometheus, Grafana, New Relic, or equivalent platforms is also valuable.
- Experience with scripting and operational automation using tools such as Bash, PowerShell, or Python, along with infrastructure-as-code and configuration-management practices.
- Experience participating in production on-call rotations, incident response, root cause analysis, and post-incident improvement.
- Experience working in Agile/DevOps environments and operating production services using ITSM practices where applicable.
- Prior software development experience—or experience investigating application behavior through code, logs, and distributed traces—is a plus.
Responsibilities
- Own and continuously improve how we use Datadog to make reliability visible and actionable across metrics, logs, traces, dashboards, monitors, alerts, and service-level views.
- Design, implement, and maintain high-availability, high-throughput, data- and compute-intensive critical systems supporting a growing 24×7 SaaS platform.
- Partner with service owners to define and improve reliability through meaningful SLIs, SLOs, error budgets, actionable alerting, and operational-readiness practices.
- Participate in and help drive incident management for production events, serving as an incident commander or technical responder as needed. Coordinate triage, service restoration, escalation, communication, incident documentation, root cause analysis, and completion of corrective actions.
- Partner with development teams to investigate issues across the infrastructure and application layers using metrics, logs, distributed traces, and code-level context.
- Improve deployment safety and service resilience through automated validation, recovery and rollback capabilities, reliability testing, and analysis of system failure modes.
- Partner with other technical leaders to ensure all newly introduced systems are supportable and maintainable by both development and operations.
- Provide escalated technical guidance and support to other technology teams throughout the organization.
- Provide on-call coverage for production support and other duties as required.
- Ensure supported systems and operational activities comply with organizational security, HIPAA, and operating policies.
- Identify and eliminate repetitive operational toil using Bash, PowerShell, Python, or Ansible. Manage infrastructure as code using Terraform/OpenTofu and configuration automation using Ansible.
Benefits
- Competitive salary - $110,000-$150,000
- Employer-sponsored health, dental, vision, life, and disability insurance
- Retirement plan with company contribution
- Annual company profit sharing
- Personal development/training budget
- Open, collaborative work environment
- Extensive 2-week onboarding plan
- Comprehensive mentorship program
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at TherapyNotes
Explore Top Companies in this Space
TherapyNotes
View Company ProfileTherapyNotes is a leading provider of practice management software for mental health professionals. The company offers a comprehensive platform that streamlines clinical, administrative, and billing tasks, allowing therapists to focus on what matters most - providing exceptional patient care. With TherapyNotes, practitioners can efficiently manage patient records, scheduling, and insurance claims, while also accessing a range of tools and features designed to support their unique needs. From solo practitioners to large groups, TherapyNotes is dedicated to helping mental health professionals achieve their full potential and deliver outstanding care to their clients. By leveraging cutting-edge technology and a deep understanding of the mental health industry, TherapyNotes has established itself as a trusted partner for therapists and a driving force behind the evolution of modern mental health care.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.
