Incident Commander
CanadaJob Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
About the Role & Team
As part of the team, you will be working with a team of smart, friendly, and dedicated Engineers, Product Managers and Designers determined to deliver some of the best apps the market has to offer. We are looking for an Incident Commander to join our site reliability team, to work cross-functionally across engineering, and be the front line for incidents and working with Release Engineering to help prevent new events.
This is a position responsible for all incidents across various organizations within the company, both online and physical, which includes P1, P2, P3 and P4. Classifying and documenting all incidents and carrying out support, assisting and driving all incidents, regarding investigation, hierarchical and technical escalation, diagnosis and recovery and root cause analysis. Additionally driving improvements to our service delivery and release processes based on disruption reports.
About the Work
- Drive and enhance collaboration with other Incident Commanders, Customer Support, Application and Engineering teams - cross-functional teams to lead real-time incident management.
- Provides Leadership for developing Practices, Frameworks, Process Flows, Templates and Process Guides
- Continuously improve and enhance the internal framework, methodology, processes, and tools
- Developing and maintaining key practical capabilities
- Collaborating with SRE Teams and Infrastructure teams to identify requirements and gaps resulting in downtime or blindspots.
- Recommends innovative solutions that enable the organization to deliver on its objectives and goals.
- Promote opportunities for Continuous Service Improvements
- Manage and update Root Cause Analysis documentation.
- Lead SRE communications to stakeholders via E-mail, Slack, & Teams in timely manner
- Lead initiatives to promote JIRA Release Ticket management, quality and alignment with Incident management communication supporting SLAs
- Other duties as required.
About You
- Experience in a similar role or incident management role.
- Experience and understanding of Containerization (Docker & Kubernetes preferred)
- Automation: Understanding of configuration management and infrastructure as code tools. Terraform, Ansible, Helm, etc.
- Experience with a programming language.
- Comfortable within Linux environments and needs.
- Experience working with AWS, GCP, and on-premises environments.
- Ability to work independently and learn quickly with little supervision.
- Ability to handle multiple projects simultaneously.
- Willingness to drop everything and take on an ad-hoc task.
- You’re the type of individual who is tech-savvy and passionate about learning new technologies and tools.
- Out-going, and able to keep a conversation going naturally to extract needed information
- A degree in computer science, engineering, and/or similar experience.
- Nice to have: Postgres, MySQL, Elastic Search, Kafka, Redis, Terragrunt, Prometheus, Python, Talos Linux
What We Offer
- Competitive compensation package.
- Fun, relaxed work environment.
- Education and conference reimbursements.
- Parental leave top up.
- Opportunities for career progression and mentoring others.
How would you rate this job post?
See what other professionals think about this role.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.