Back to Jobs
Nice
Engineering & Architecture 1h ago

SRE – NOC (Site Reliability Engineer)

Nice
United KingdomUnited Kingdom
Full-time
Not Disclosed
Mid-Level

Job Description

Key Skills Required

Master these to land this role

Python2h 41mFree Trial ✨
Start 10-Day Free Trial
DevOps1h 38mFree Trial ✨
Start 10-Day Free Trial
KubernetesAWSAutomation Engineer

Want to know if you're a match for this job?

Calculate My Match Score

The SRE – NOC role sits at the intersection of traditional Network Operations Center (NOC) responsibilities and engineering-driven reliability practices. This role focuses on 24/7 service reliability, incident response, operational automation, and observability, while actively reducing operational toil through software and automation.

Unlike a traditional NOC analyst, an SRE-NOC is expected to engineer problems away, not just respond to alerts.

How will you make an impact?

Incident Response & Operations

  • Act as a primary or escalation responder in a 24x7 on-call rotation
  • Lead or support Major Incident (MI) response, including triage, mitigation, and resolution
  • Coordinate across Engineering, Infrastructure, Security, and Product teams
  • Execute and improve runbooks, playbooks, and escalation paths
  • Drive blameless post-incident reviews (PIRs) and track corrective actions

Monitoring, Alerting & Observability

  • Own service health monitoring across infrastructure, applications, and dependencies
  • Design and maintain alerting strategies that align with SLIs/SLOs
  • Reduce alert fatigue through signal-to-noise improvements
  • Build dashboards using tools such as:
    • Grafana
    • Prometheus
    • Datadog / Splunk / CloudWatch

Reliability Engineering & Automation

  • Automate repetitive operational tasks to reduce manual toil
  • Improve mean time to detect (MTTD) and mean time to resolve (MTTR)
  • Develop scripts and tools (Python, Bash, Go, etc.) to support NOC/SRE workflows
  • Implement self-healing and auto-remediation where possible
  • Partner with engineering teams to improve system design for reliability

Platform & Infrastructure Support

  • Support and troubleshoot:
    • Linux-based systems
    • Cloud platforms (AWS, Azure, GCP)
    • Kubernetes / containerized environments
  • Assist with capacity planning and availability reviews
  • Ensure operational readiness for production releases

Have you got what it takes?

Technical

  • Strong Linux systems administration
  • Experience with incident management and production support
  • Familiarity with:
    • Cloud infrastructure (AWS preferred)
    • Containers & orchestration (Docker, Kubernetes)
    • Monitoring/alerting platforms
  • Scripting or programming experience in Python, Bash, Go, or similar
  • Understanding of networking fundamentals (DNS, TCP/IP, load balancing)

Operational

  • Experience working in 24x7 NOC or production operations environments
  • Ability to handle high-pressure incidents calmly and effectively
  • Strong written and verbal communication for incident coordination
  • Comfort working from runbooks—but improving them when they fall short

Preferred / Differentiators

  • Experience defining or operating to SLOs / SLIs
  • Prior migration from traditional NOC → SRE model
  • Infrastructure as Code experience (Terraform, Ansible, etc.)
  • Exposure to security, compliance, or regulated environments

How would you rate this job post?

See what other professionals think about this role.

banner

Nice is a company that embodies the principles of innovation and customer satisfaction. With a name that reflects a commitment to excellence, Nice operates with the goal of providing top-notch products and services that meet the evolving needs of its clients. The company's mission is centered around creating a positive experience for all stakeholders, fostering a culture of collaboration, and driving growth through sustainable practices. As a forward-thinking organization, Nice is well-positioned to capitalize on emerging trends and technologies, leveraging its expertise to stay ahead of the curve. With a strong focus on research and development, Nice continually seeks to improve its offerings, ensuring that they remain relevant and effective in an ever-changing landscape. Through its dedication to quality and customer-centric approach, Nice aims to establish long-lasting relationships with its customers, built on trust, reliability, and mutual benefit. By prioritizing the well-being of both its clients and the environment, Nice strives to make a meaningful impact, contributing to the betterment of society while pursuing its business objectives. The company's leadership team comprises experienced professionals who bring a wealth of knowledge and insight to the table, guiding Nice towards achieving its vision of becoming a leading player in its industry. With its solid foundation, Nice is poised for success, ready to adapt and evolve in response to new challenges and opportunities that arise. Through its relentless pursuit of innovation and commitment to excellence, Nice is an organization that is sure to leave a lasting impression, making it an exciting company to watch in the years to come.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More
SRE – NOC (Site Reliability Engineer) at Nice | HireSkys