Site Reliability Engineering Team Leader (60% IC, 40% Leadership) at Remote
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Remote’s SRE team exists so that our engineers can move quickly and our customers get a product that stays up. The team owns Kubernetes, AWS, PostgreSQL, CI infrastructure, our observability stack, and the reliability practices that sit on top of all of it.
We are looking for a Team Leader to run that team. This is a 60% Individual Contributor (IC), 40% leadership role. You will own the career development of your reports, steer the team’s focus using judgment against the company goals, and you will be the spokesperson for the team across engineering. You will also stay close enough to the technical work to set direction with credibility and to know when something is going wrong before it is escalated to you.
Reliability practice at Remote is maturing rather than mature. Our SLO framework is live on its first few teams and needs to reach the rest; there is real work to do on how we balance operational load against project delivery. If you want a team where the foundations are in place and the interesting problems are still open, this is that team.
What you bring
People leadership
- You have led an SRE, infrastructure, or platform engineering team, and you have owned your reports’ growth, performance, and career progression rather than just their sprints.
- You coach both craft and the soft skills; you can point to people who grew because of it.
- You handle underperformance directly and early, with clarity and empathy.
- You have hired engineers, and you can tell the difference between a good interview and a good engineer.
- You read team dynamics well and you resolve conflict rather than routing around it.
- You have a natural talent fostering commitment to the goals of the company.
Technical depth
- A hands-on background in site reliability, DevOps, or cloud infrastructure engineering, deep enough that you can review your team’s work, challenge a design, and be taken seriously in an incident.
- Kubernetes in production, including the operational reality of it rather than the happy path.
- AWS at meaningful scale.
- Hands-on AI building, enablement, scaling AI infrastructure.
- Solid observability (o11y) practices and principles.
- Infrastructure as code with Terraform.
- CI/CD systems such as GitLab CI, GitHub Actions, or Jenkins.
- Docker and shell scripting.
- You have run a reliability practice: incident response, on-call, SLOs and error budgets, and the discipline of turning incidents into changes that stick.
- Understanding and history of working in regulated environments.
Ways of working
- You prioritize exceptionally well when operational load and project work compete, and you protect your team’s focus without dropping the operational commitment.
- You write clearly. Remote is fully distributed and async, so most of your leadership will happen in writing.
- You build relationships across teams. A lot of SRE’s value comes from being the team others bring problems to early.
Nice to have
- Working knowledge of a backend language, ideally Elixir, or otherwise Java, Clojure, Node.js, Python, or similar.
- Depth in modern observability: OpenTelemetry, distributed tracing, and tools such as Honeycomb.
- Database operations experience, particularly PostgreSQL or Aurora performance, connection pool health, and query tuning.
- Running and configuring Linux systems outside a cloud environment.
- Security capability from both a defensive and an offensive standpoint.
- Cloud cost management and FinOps.
- Experience growing a team from a small base, including building the hiring bar as you go.
Key Responsibilities
Your people
- The full career lifecycle of your reports: onboarding, feedback, performance assessment, progression, and hiring.
- Team health, dynamics, and the retrospective habit that keeps them honest.
- Being the team’s spokesperson to the rest of engineering and to senior leadership.
Delivery
- The SRE goals: what the team commits to, in what order, and why.
- The support rotation and on-call model.
The platform
- Remote’s core infrastructure: Kubernetes, AWS, PostgreSQL, DNS, and TLS, CI infrastructure.
- The reliability practice: SLOs, error budgets, incident response, and the observability stack.
- The partnership with our Security team on threats, patching, and infrastructure controls, including our audit and compliance obligations.
- The vendor relationships that sit behind the platform, including renewals and commercial conversations with support from your Director.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Remote
Site Reliability Engineering Team Leader (60% IC, 40% Leadership) at Remote
Remote
United StatesBack-End Marketing & Client Success Coordinator (Part-Time to Full-Time)
Remote
South AfricaStrategic Alliance Principal, UK & Ireland
Remote
United KingdomSenior Executive Assistant - Business Operations
Remote
PhilippinesExplore Top Companies in this Space
Clipboard
Software as a Service (SaaS), Technology
DesignFiles
Software as a Service (SaaS) and Design Technology
Pushpay
Software as a Service (SaaS), Faith-based Technology
HighLevel
Software as a Service (SaaS), Technology
Remote
View Company ProfileRemote is a company that specializes in enabling remote work for businesses and individuals. With the shift towards a more flexible and distributed workforce, Remote provides innovative solutions to facilitate collaboration, communication, and productivity. The company's platform and tools empower teams to work seamlessly from anywhere, at any time, and on any device. By offering a range of features such as virtual offices, remote team management, and secure data storage, Remote enables companies to unlock new levels of efficiency, creativity, and growth. Whether it's a startup, a scale-up, or an enterprise, Remote's cutting-edge technology and expertise help organizations navigate the complexities of remote work, ensuring that they can stay competitive, agile, and successful in an ever-changing business landscape. With a strong focus on user experience, security, and customer support, Remote is committed to helping businesses thrive in a remote-first world.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.
