Back to Jobs
Development 1h ago

Site Reliability Operations (SRO) Engineer

United StatesUnited States
Full-time
Not Disclosed
Mid-Level

Job Description

Key Skills Required

Master these to land this role

DevOpsBestseller 🔥
Learn in 63 Hours
PythonBestseller 🔥
Learn in 56 Hours
AWSLinuxCI/CD

Want to know if you're a match for this job?

Calculate My Match Score

Job Summary

The Site Reliability Operations (SRO) team ensures 24/7 stability of the internal IT infrastructure and mission-critical backend systems. This role is not DevOps-focused, but is crucial in monitoring, coordinating, and restoring operations during incidents, particularly in a high-stakes, regulated environment.

The role balances incident command, technical troubleshooting, project leadership, and communication with multiple internal and external stakeholders.

Responsibilities

  • Oversee multi-platform IT infrastructure health using AWS CloudWatch, New Relic, Nagios, and SumoLogic. Continuously refine alert thresholds to minimize noise and enable proactive remediation.
  • Serve as an escalation point for complex technical issues. Perform deep-dive troubleshooting across Linux/UNIX, Windows, virtual servers, and virtual desktop environments.
  • Coordinate, automate, and execute code deployments using Jenkins, GitLab, or similar CI/CD tools, driving toward non-disruptive releases and zero-downtime updates.
  • Collaborate closely with Application Developers, 3rd-party vendors, and internal Incident Management teams during critical service disruptions. Oversee ticket queues in ServiceNow and Jira.
  • Drive medium- to large-scale infrastructure projects, including migrations, cloud upgrades, and performance tuning.
  • Conduct risk assessments for production updates, maintain SOPs in team knowledge bases, and oversee enterprise backup operations (CommVault, Veeam, AWS Backup) to ensure strict regulatory compliance.

Qualifications & Requirements

Experience: 3–5+ years in an Operations Center (SRO/NOC) or cloud infrastructure environment with hands-on experience in full-stack application deployments.

Systems Expertise: Solid proficiency in both Windows and UNIX/Linux administration, deep-dive troubleshooting (scripting, grepping logs, analyzing performance metrics), and virtual server/desktop management.

Cloud & Monitoring: Direct experience with AWS cloud services (Storage, VMs, Networking) and enterprise monitoring tools (AWS CloudWatch, New Relic, Nagios, SumoLogic).

Scripting & Automation: Strong scripting or programming capability in PowerShell, Python, or bash to automate repetitive tasks and optimize run-time operations.

Tools & Backup: Practical experience with CI/CD platforms (Jenkins, GitLab), ITSM ticket platforms (ServiceNow, Jira), and backup solutions (CommVault, Veeam, AWS Backup).

Communication: Excellent verbal and written communication skills with experience serving as a bridge between technical teams, executive stakeholders, and external vendors.

Pluses

  • Advanced AWS Certifications (e.g., AWS Certified Solutions Architect Professional or DevOps Engineer Professional).
  • Background in or exposure to AI/ML tools for infrastructure monitoring and predictive analytics.
  • Previous experience in ITIL-aligned environments or enterprise Change/Incident Management frameworks.
  • Bachelor’s Degree in Computer Science, Information Technology, or a related field.

How would you rate this job post?

See what other professionals think about this role.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More