Site Reliability Operations (SRO) Engineer
United StatesJob Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Job Summary
The Site Reliability Operations (SRO) team ensures 24/7 stability of the internal IT infrastructure and mission-critical backend systems. This role is not DevOps-focused, but is crucial in monitoring, coordinating, and restoring operations during incidents, particularly in a high-stakes, regulated environment.
The role balances incident command, technical troubleshooting, project leadership, and communication with multiple internal and external stakeholders.
Responsibilities
- Oversee multi-platform IT infrastructure health using AWS CloudWatch, New Relic, Nagios, and SumoLogic. Continuously refine alert thresholds to minimize noise and enable proactive remediation.
- Serve as an escalation point for complex technical issues. Perform deep-dive troubleshooting across Linux/UNIX, Windows, virtual servers, and virtual desktop environments.
- Coordinate, automate, and execute code deployments using Jenkins, GitLab, or similar CI/CD tools, driving toward non-disruptive releases and zero-downtime updates.
- Collaborate closely with Application Developers, 3rd-party vendors, and internal Incident Management teams during critical service disruptions. Oversee ticket queues in ServiceNow and Jira.
- Drive medium- to large-scale infrastructure projects, including migrations, cloud upgrades, and performance tuning.
- Conduct risk assessments for production updates, maintain SOPs in team knowledge bases, and oversee enterprise backup operations (CommVault, Veeam, AWS Backup) to ensure strict regulatory compliance.
Qualifications & Requirements
Experience: 3–5+ years in an Operations Center (SRO/NOC) or cloud infrastructure environment with hands-on experience in full-stack application deployments.
Systems Expertise: Solid proficiency in both Windows and UNIX/Linux administration, deep-dive troubleshooting (scripting, grepping logs, analyzing performance metrics), and virtual server/desktop management.
Cloud & Monitoring: Direct experience with AWS cloud services (Storage, VMs, Networking) and enterprise monitoring tools (AWS CloudWatch, New Relic, Nagios, SumoLogic).
Scripting & Automation: Strong scripting or programming capability in PowerShell, Python, or bash to automate repetitive tasks and optimize run-time operations.
Tools & Backup: Practical experience with CI/CD platforms (Jenkins, GitLab), ITSM ticket platforms (ServiceNow, Jira), and backup solutions (CommVault, Veeam, AWS Backup).
Communication: Excellent verbal and written communication skills with experience serving as a bridge between technical teams, executive stakeholders, and external vendors.
Pluses
- Advanced AWS Certifications (e.g., AWS Certified Solutions Architect Professional or DevOps Engineer Professional).
- Background in or exposure to AI/ML tools for infrastructure monitoring and predictive analytics.
- Previous experience in ITIL-aligned environments or enterprise Change/Incident Management frameworks.
- Bachelor’s Degree in Computer Science, Information Technology, or a related field.
How would you rate this job post?
See what other professionals think about this role.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.