Senior Site Reliability Engineer
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
About UJET: UJET leads the way in AI-powered contact center innovation, delivering a future-proof, cloud platform that redefines the customer experience with cutting-edge AI, true multimodality, and a mobile-first approach. We infuse AI across every aspect of your customer journey and contact center operations to drive automation and efficiency.
Opportunity: We’re looking for a Senior Site Reliability Engineer to help build and scale a high-impact SRE function. You’ll be a technical leader on a team responsible for improving system reliability, reducing operational toil, and establishing best practices across engineering.
In this position, you’ll design how reliability works in UJET, influence engineering decisions, and build the tooling and processes that make production safer and more predictable.
Responsibilities:
- Lead efforts to improve system reliability, scalability, and performance across critical services
- Define and implement SLIs/SLOs and error budgets, and use them to guide engineering priorities
- Design and develop observability systems (metrics, logging, tracing, alerting) that produce actionable alerts and data
- Lead complex incident response, acting as incident commander when needed
- Conduct postmortems focused on systemic causes rather than individual fault, and ensure corrective actions from those reviews are completed
- Identify and eliminate toil through automation, tooling, and improved workflows
- Partner with product and platform teams on architecture decisions, production readiness, and designing systems that recover from failure
- Build reusable systems and “paved roads” that make it easier for teams to operate their services reliably
- Mentor other engineers and raise the overall operational maturity of the organization
Requirements:
- 6-10+ years of experience in SRE, infrastructure, or backend systems engineering
- Demonstrated experience of owning reliability outcomes for complex, distributed systems
- Strong experience with cloud infrastructure (AWS, GCP, or Azure) and production-scale systems
- Deep understanding of observability, incident management, and system performance
- Proficiency in at least one programming language (e.g., Go, Python, Java) with a focus on automation and tooling
- Able to change how other teams work without having managerial authority over them
- Strong competency in making clear decisions during incidents by following a defined process without reacting emotionally
Stand Out Qualifications:
- Experience building or scaling SRE practices (SLOs, incident frameworks, on-call models)
- Kubernetes/container orchestration experience
- Infrastructure as Code (Terraform, etc.)
- Experience with high-growth or scaling systems
- Background in performance engineering or capacity planning
Success Criteria:
- Critical services have clear, meaningful SLOs that drive engineering decisions
- Alerts are actionable; irrelevant alerts are reduced; on-call workload is manageable
- Incidents are handled efficiently, and repeat issues decline over time
- Engineering teams adopt reliability best practices with minimal friction
- Toil is actively reduced through automation and better system design
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
System Administrator (DISA)
Horizon Industries
United StatesMicrosoft Dynamics 365 Finance & Operations Implementation Lead
GrĂĽn
United StatesSenior Implementation Specialist (Email/SMS) at Hightouch
Hightouch
United StatesDirector of Engineering - Enterprise Platform and Integrations
Zenbusiness
United StatesMore Openings at UJET
Explore Top Companies in this Space
Eltropy
Fintech & Digital Communications / Omnichannel Contact Center SaaS / Conversational AI & Customer Experience / B2B Enterprise Software
Verint
Enterprise Software / Customer Experience / Analytics / AI
Emplifi
Enterprise Software / AI / Customer Experience / Digital Marketing
Reliant Health Partners
Healthcare / Medical Technology / Business Services / Enterprise Software
UJET (operating at ujet.cx) is a cloud-native contact center platform engineered for modern customer experience (CX) transformation. Founded in San Francisco, UJET replaces fragmented, outdated contact center stacks with a unified, AI-powered solution designed for the smartphone era. The platform integrates voice, chat, email, and social media channels into a single interface, leveraging AI to automate repetitive tasks while empowering agents with real-time insights. This allows enterprises to deliver hyper-personalized interactions at scale, reducing operational costs by up to 30% while boosting agent productivity and customer satisfaction. Backed by a $177 million funding round across six rounds—including a $76 million Series D in 2024—UJET serves global brands seeking to elevate their CX without legacy constraints.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.