Site Reliability Engineer (SRE)
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Who we're looking for
We’re looking for people (US - Central/Eastern Timezone) that like deep ownership of production systems, people that are not afraid of working with stateful infrastructure and love working in AWS, VMs, automation, and making messy systems reliable.
In general we seek SRE’s who are:
Enthusiastic drivers. We need proactive people that can fully own projects and get them done, and know to get help when needed. "Are we there yet?" is the wrong question.
Optimistic problem solvers. Things get hard here sometimes, whether it's scaling, shipping complex products, handling a stream of support requests, or trying to ship something that touches multiple teams. We need people who won't get disheartened, and will collaborate, iterate, and ship their way out of anything.
Grown ups. We’re an international bunch of weirdos, but one thing unites us: everyone is kind, considerate, and professional towards each other. This isn't about age or experience, it's about being low-ego, flexible, and respectful.
Genuine builders. PostHog is full of people who just love building stuff, people who would still be building software even if there wasn't a paycheck at the end. If this sounds like you, we should talk.
What you'll be doing
You won’t be in a typical “keep the lights on” SRE role. The work is about turning a fast-growing, stateful system into a predictable, well-automated platform (provisioning, scaling, rebalancing, recovery). That means reducing operational stress, designing safe automation for traffic-heavy workloads, and building the tooling and patterns that let the system scale without scaling human effort.
You'll work on the kind of problems that only show up at large scale (petabytes of data, thousands of cores, constant ingestion) across a multi-region, multi-account AWS platform running many services on Kubernetes.
Operating EKS clusters across several environments with Karpenter autoscaling, Cilium networking, and ArgoCD-driven GitOps deployments
Managing and evolving a multi AWS account organization, provisioning, networking, access control, and cross-account connectivity
Maintaining the Terraform/Terragrunt IaC platform - modules, automated plan-on-PR / apply-on-merge pipelines, and safe patterns for shared infrastructure
Improving operational tooling around deploys, schema changes, backups, restores, and incident response
Reducing operational load by identifying repeat pain points and eliminating them through code and self-healing automation
Optimizing cloud spend as you go
Participating in on-call and incident response, with a strong focus on making incidents rarer over time
You'll have room to design and automate, not just respond to alerts. You should join this team if you like deep ownership of production systems and enjoy building the platform layer that everything else runs on.
Requirements
Deep hands-on experience with Kubernetes in production (EKS preferred). You've debugged node pressure, networking issues, and deployment failures at scale (thousands of nodes)
Strong experience operating production infrastructure on AWS. Not just one account, but understanding organizational boundaries, IAM, and networking between many
Experience automating infrastructure using Terraform or Terragrunt at scale, including module design and state management
Solid understanding of Linux systems (disk, memory, networking, failure modes)
Experience supporting stateful systems (databases, queues, storage systems, etc.)
Ability to debug and reason about performance and reliability issues in production
You're comfortable owning systems end-to-end, including on-call responsibilities
You don't need to be an expert in every system we run on day one. But you do need to enjoy owning complex infrastructure and learning how the pieces fit together.
Nice to have
Experience with GitOps workflows (ArgoCD) and CI/CD pipelines (GitHub Actions)
Experience with building AI agent-enabled base-level infra services for teams that move fast
Familiarity with multi-region infrastructure and the consistency/availability tradeoffs that come with it
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at PostHog
Explore Top Companies in this Space
LeoLabs
Aerospace / SpaceTech / Data Analytics
Aristo Sourcing
Human Resources
Jetbrains
Software Development Tools and Technologies
HireDigital
Human Resources, Recruitment Technology
PostHog
View Company ProfilePostHog is a powerful, open-source product operating system designed to help software engineers and product teams build better products faster. Founded in 2020 by Y Combinator alumni James Hawkins and Tim Glaser, the company has completely disrupted the traditional product analytics market by offering a single, unified suite of developer tools. Under the hood, PostHog provides deep product analytics, session recording, feature flagging, A/B testing, and user data pipelines (CDP) all within one platform that can be seamlessly self-hosted or deployed in the cloud. Their primary target audience spans developers, product managers, and agile engineering teams at hyper-growth startups and enterprises who want full control over their user data without relying on a fragmented stack of third-party SaaS tools. What sets PostHog apart in the crowded data landscape is its transparent, open-source architecture, its incredibly vibrant developer community, and its ability to consolidate the entire modern product stack into one workflow—giving teams the ultimate power to understand user behavior and ship winning features with absolute confidence.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.


