Back to Jobs
PostHog
Development 2h ago

Site Reliability Engineer (SRE) - ClickHouse & AWS Infrastructure

PostHog
United KingdomUnited Kingdom
Full-time
Not Disclosed
Senior-Level

Job Description

Key Skills Required

Master these to land this role

BackendBestseller 🔥
Learn in 18 Hours
DevOpsBestseller 🔥
Learn in 63 Hours
AWSClickHouseAutomation Engineer

Want to know if you're a match for this job?

Calculate My Match Score

What you'll be doing

We run one of the largest self-managed ClickHouse installations on AWS, at petabyte scale, and we’re actively preparing it for the next 10–50× of growth. This role sits at the centre of that effort.

You won’t be in a typical “keep the lights on” SRE role. The work is about turning a fast-growing, stateful system into a predictable, well-automated platform. (provisioning, scaling, rebalancing, recovery)
That means reducing operational stress, designing safe automation for data-heavy workloads, and building the tooling and patterns that let the system scale without scaling human effort.

You’ll work on the kind of problems that only show up at large scale (petabytes of data, thousands of cores, constant ingestion).

  • Managing large fleets of EC2-based VMs, disks, and networking for data-intensive workloads
  • Improving operational tooling around deploys, schema changes, backups, restores, and incident response
  • Working closely with ClickHouse engineers to turn database-level needs into infra-level solutions
  • Reducing operational load by identifying repeat pain points and eliminating them through code and self-healing automation
  • Participating in on-call and incident response, with a strong focus on making incidents rarer over time
  • You’ll have room to design and automate, not just respond to alerts.

You should join this team if you like deep ownership of production systems, and are not afraid of working with stateful infrastructure.

Requirements

  • Prior experience with ClickHouse or other OLAP databases
  • Strong experience operating production infrastructure on AWS
  • Hands-on experience with VM-based systems (EC2), not just managed PaaS
  • Experience automating infrastructure using tools like Terraform, Ansible, or similar
  • Solid understanding of Linux systems (disk, memory, networking, failure modes)
  • Experience supporting stateful systems (databases, queues, storage systems, etc.)
  • Ability to debug and reason about performance and reliability issues in production
  • You’re comfortable owning systems end-to-end, including on-call responsibilities

You don’t need to be a ClickHouse expert on day one. We’ll teach you the database internals, but you do need to enjoy owning complex infrastructure.

How would you rate this job post?

See what other professionals think about this role.

banner

PostHog is a powerful, open-source product operating system designed to help software engineers and product teams build better products faster. Founded in 2020 by Y Combinator alumni James Hawkins and Tim Glaser, the company has completely disrupted the traditional product analytics market by offering a single, unified suite of developer tools. Under the hood, PostHog provides deep product analytics, session recording, feature flagging, A/B testing, and user data pipelines (CDP) all within one platform that can be seamlessly self-hosted or deployed in the cloud. Their primary target audience spans developers, product managers, and agile engineering teams at hyper-growth startups and enterprises who want full control over their user data without relying on a fragmented stack of third-party SaaS tools. What sets PostHog apart in the crowded data landscape is its transparent, open-source architecture, its incredibly vibrant developer community, and its ability to consolidate the entire modern product stack into one workflow—giving teams the ultimate power to understand user behavior and ship winning features with absolute confidence.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More