Lead Site Reliability Engineer
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
We are looking for a hands-on Lead Site Reliability Engineer for a project based assignment to establish and continuously improve the performance, reliability, and scalability of our platform. This role will define what the platform can reliably sustain today, identify where constraints will emerge, and ensure the organization is prepared to scale before demand arrives.
This is not a traditional DevOps role or an advisory architecture position. You will work directly across application services, infrastructure, databases, networking, caching, queues, external dependencies, and operational processes to identify bottlenecks, validate system limits, and lead remediation.
You will partner closely with engineering, product, and leadership to provide clear, evidence-based answers around capacity, performance, reliability, risk, and the cost of scaling.
Key Responsibilities:
- Establish performance, throughput, latency, and capacity baselines for critical customer and platform workflows
- Define and maintain SLOs, error budgets, performance budgets, dashboards, alerts, and reliability thresholds
- Instrument and analyze the full request path across application services, compute, storage, networking, databases, caches, queues, DNS, registry dependencies, and third-party services
- Identify system bottlenecks and lead cross-functional remediation efforts with engineering teams
- Build capacity models that show what the platform can sustain, where constraints will emerge, and what additional scale will cost
- Lead load, stress, soak, spike, failure, and recovery testing in representative environments
- Develop realistic demand scenarios for major customers, partnerships, pilots, and high-volume events
- Drive architecture hardening, graceful degradation, dependency-failure planning, and resilience improvements
- Partner with Test Automation and Scalability Engineering to establish automated performance testing, regression coverage, and production release gates
- Own technical readiness assessments for major pilots, partnerships, and production launches
- Create operational runbooks for scale-up events, incidents, rollback, recovery, and dependency failures
- Lead performance and reliability investigations during incidents and ensure lessons are incorporated into future engineering work
- Make infrastructure cost, performance, and reliability tradeoffs visible to engineering and executive leadership
- Recommend capacity and reliability investments before they become production constraints
Required Skills:
- Significant experience in Site Reliability Engineering, performance engineering, platform engineering, distributed systems, or a closely related engineering discipline
- Experience supporting production systems with meaningful scale, traffic, latency, or availability requirements
- Deep understanding of observability, performance analysis, capacity planning, and reliability engineering
- Strong hands-on experience with cloud infrastructure and production distributed systems
- Deep knowledge of databases, networking, caching, queueing, compute, storage, and common distributed-system failure modes
- Experience defining and operating against SLOs, SLIs, error budgets, and production reliability metrics
- Hands-on experience performing load, stress, soak, scalability, and resilience testing
- Ability to profile systems, diagnose bottlenecks, tune architecture, and work directly with engineering teams to implement improvements
- Experience designing for graceful degradation, dependency failures, recovery, and high-demand scenarios
- Strong incident management and root-cause analysis experience
- Ability to translate technical performance and reliability risks into clear business implications for senior leadership
- Strong judgment around when systems genuinely require optimization versus when additional complexity is premature
Nice to have:
- Experience operating high-scale SaaS, identity, DNS, registry, infrastructure, or other highly distributed platforms
- Experience creating capacity-cost models and forecasting infrastructure requirements
- Experience building performance and reliability gates into CI/CD pipelines
- Experience preparing platforms for significant increases in traffic associated with enterprise customers or strategic partnerships
- Experience leading reliability or performance initiatives that span multiple engineering teams
What Success Looks Like
Within your first several months, you will have:
- Established measurable throughput, latency, and capacity baselines for critical platform journeys
- Defined initial SLOs, error budgets, dashboards, alerts, and performance thresholds
- Identified the platform's most significant scalability and reliability constraints and created an actionable remediation roadmap
- Validated representative high-scale scenarios through load, soak, stress, failure, and recovery testing
- Developed a capacity and cost model showing how the platform can support significant increases in demand
- Established production-readiness criteria and clear go/no-go evidence for major launches and partnerships
- Created repeatable scale-up, incident, rollback, and dependency-failure runbooks
- Given leadership a clear, evidence-based understanding of the platform's current capacity envelope and future scaling requirements
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Tech Holding
Explore Top Companies in this Space
Tech Holding
View Company ProfileTech Holding is an elite, global digital transformation, product engineering, and enterprise consulting powerhouse engineered to operate as the definitive, full-stack technology infrastructure partner for high-growth startups, mid-market leaders, and Fortune 500 enterprises. Operating as a mission-critical technology acceleration partner, the firm eliminates the severe systemic friction of legacy IT models and unoptimized engineering pipelines—which frequently suffer from talent scarcity, prolonged time-to-market loops, fragmented data architectures, and decoupled cloud systems—by deploying cross-functional, agile delivery teams. Moving far beyond traditional, rigid outsourcing practices or passive software consultancies, Tech Holding empowers global business units to dynamically synchronize their cloud infrastructure modernizations (AWS, Azure, GCP), custom mobile and web applications, AI/ML integrations, and data engineering pipelines with elite, scalable, and production-hardened precision. Under the hood, their sophisticated global operational core—bolstered by distributed engineering talent across North America, Latin America, Europe, and Asia, alongside deep compliance certifications for healthcare (HIPAA) and financial tech—natively orchestrates complex enterprise architecture design, rapid prototyping cycles, and end-to-end DevOps automation. What sets Tech Holding apart is its uncompromising dedication to blending visionary product craftsmanship with high-velocity software engineering; by bridging the gap between performance-intensive data infrastructures and intuitive, user-centric product designs, the firm enables modern organizations to radically accelerate their digital innovation, eliminate technical debt bottlenecks, and build an unassailable foundation for continuous commercial growth.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.

