Back to Jobs
Delinea
Engineering & Architecture 3h ago

Manager of Site Reliability Engineering

Delinea
🌍U
Full-time
Not Disclosed
Senior-Level

Job Description

Key Skills Required

Master these to land this role

DevOps1h 38mFree Trial ✨
Start 10-Day Free Trial
FedRAMPAWSAzure Kubernetes ServiceDatadogObservabilityIncident Command

Want to know if you're a match for this job?

Calculate My Match Score

Delinea is seeking a hands-on Manager of Site Reliability Engineering to lead the SRE and DevOps engineers supporting Delinea’s products. This is a working manager role focused on leading people and work, including writing and reviewing automation, managing AKS and Azure telemetry, handling Sev1 and Sev2 incidents, and improving observability and deployment practices.

Core Responsibilities:

  • Lead hands-on: Spend a meaningful portion of your week in the environment—reviewing pull requests, validating pipeline changes, tuning monitors and dashboards, running queries in Datadog, and troubleshooting production issues alongside engineers. This role is not purely managerial.

  • Own availability and performance: Ensure robust operational health of Delinea Platform production environments across Azure and AWS, including AKS workloads, ingress, networking, data services, messaging, and CDN/WAF layers.

  • Manage a blended team: Hire, onboard, coach, and develop full-time SRE engineers. Direct and manage contractor resources, including scoping work, setting quality expectations, and reviewing deliverables.

  • Lead a distributed team: Run one-on-ones, standups, and planning sessions at times that accommodate engineers across multiple time zones.

  • Participate in on-call: Carry the pager as part of the rotation, act as incident commander for Sev1 and Sev2 events, and ensure effective communication with support, engineering, and leadership until resolution.

  • Command incident response: Own detection, triage, mitigation, customer-facing status communication, and post-incident reviews. Ensure Root Cause Analyses (RCAs) are written to a customer-ready standard, with preventative actions having owners and target dates.

  • Raise observability standards: Improve detection coverage so issues are found by monitoring rather than customer tickets. Define SLIs and SLOs, refine alert quality, expand synthetic coverage, enhance APM instrumentation, ensure log hygiene, and set dashboard standards.

  • Support FedRAMP and regulated operations: Grow into supporting Delinea’s FedRAMP High environment, including change control, evidence collection, and operational differences between government and commercial environments.

  • Reduce toil through automation: Set the expectation that manual work becomes code. Prioritize automation backlog alongside project and reliability work.

  • Report on operational health: Produce and present incident metrics, trends, and reliability commitments to leadership, translating them into concrete improvement plans.

Requirements:

  • 6+ years in Site Reliability Engineering, DevOps, or Cloud Operations, with demonstrated ownership of production SaaS systems.

  • 2+ years of direct people leadership, including performance management, hiring, and coaching. Experience managing contractors or outsourced teams is desired.

  • Hands-on experience with the Delinea tech stack: Azure Kubernetes Service (AKS), core Azure services (SQL, Redis, Service Bus, Blob Storage), AWS services (SES, EC2, RDS), WAF, Azure DevOps pipelines, Datadog, and Atlassian Jira Service Management.

  • Hands-on experience across both Azure and AWS is required, including administration, troubleshooting, cost optimization, and security posture management.

  • Deep observability expertise: Ownership of an observability framework at scale, including metrics, logs, traces, synthetics, SLOs, and alerting strategy. Proficiency with Datadog, including APM trace analysis and log-based troubleshooting.

  • Proven incident command: Experience running major incidents as incident commander, coordinating responders under pressure, communicating with customers and executives, and authoring RCAs.

  • Strong cloud networking and security fundamentals: load balancing, DNS, TLS/certificate lifecycle, firewalls, VPN, routing, and identity and access management.

  • Automation and scripting ability in PowerShell, Python, Bash, plus practical Infrastructure-as-Code (Terraform, ARM, or Bicep) experience.

  • Experience with multi-region, multi-tenant SaaS architectures, including backup, redundancy, and disaster recovery approaches.

  • Excellent written communication for customer-facing status updates and incident summaries.

  • Willingness to work across time zones and participate in on-call rotations.

Nice-to-Haves:

  • Direct experience operating in a FedRAMP or regulated environment (Azure Government, IL4/IL5, SOC 2, ISO 27001).

  • Experience standing up or maturing an incident management program, including sev definitions, escalation paths, on-call structure, and post-incident review processes.

  • Experience with public status page operations and customer notification practices.

  • Experience with Atlassian Jira Service Management, Confluence, and Azure DevOps as the operational toolchain.

  • Track record of reducing customer-detected incidents through improved monitoring coverage.

  • Cost optimization experience across Azure and AWS on a meaningful scale.

How would you rate this job post?

See what other professionals think about this role.

banner

Delinea is a global leader in Identity Security and Privileged Access Management (PAM) solutions, fundamentally redefining how modern enterprises secure digital access. Formed through the strategic merger of industry pioneers Thycotic and Centrify, the company provides an AI-driven, cloud-native platform designed to secure human, machine, and emerging AI identities. Under the hood, Delinea replaces legacy, standing access models with a robust "Zero Standing Privilege" and Just-in-Time (JIT) runtime authorization architecture—significantly strengthened by their integration with StrongDM. Powered by Delinea Iris AI, the platform centralizes authorization, continuously monitors identity posture, and intelligently vaults credentials to eliminate the massive risks associated with credential sprawl. Their primary target audience spans Fortune 500 enterprises, hyper-growth tech companies, and highly regulated government sectors that require bulletproof compliance and dynamic threat protection without slowing down developer velocity. What sets Delinea apart in the fiercely competitive cybersecurity ecosystem is its ability to make deeply complex enterprise access controls almost invisible to the end-user, enabling organizations to grant granular, task-based access seamlessly and securely in real time.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More
Manager of Site Reliability Engineering at Delinea | HireSkys