Incident Commander (24x7 On-Call)
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Role
Serve as the on-duty commander for Alpaca’s most critical incidents, directing cross-functional response to restore service quickly, keeping the right people engaged and informed, and ensuring every incident leaves behind actionable organizational learnings.
You do not fix the outage. You make the response reliable: correct severity classification, the right engineers engaged, mitigation that does not stall, leaders informed in time, and follow-up work that survives the call.
Things You Get To Do
- Command incidents end to end. Take command from declaration to mitigation, keeping responders focused on stopping customer and partner impact as fast as possible. Run the bridge, keep observers out of the responders’ way, and name a stall out loud when you see one.
- Classify and hold the line on severity. Set severity at declaration and re-check it as facts arrive. Risk advises on financial and regulatory materiality; the call is yours.
- Engage the right people, fast. Identify the owning team by service, symptom, and blast radius, page them, and expand the responder set the moment the first team is wrong or not enough. When a page goes unanswered, escalate—and escalate the escalation. Bring in the leaders who must make business calls: feature flags, traffic shedding, failover, freeze-or-ship.
- Hold the bridge, and protect the people fixing it. Keep engineering and technical support uninterrupted—questions from stakeholders, partners, and executives come to you. Be the single source of truth to the partner communications team on impact, severity, and timing: you decide when a status page update or partner contact is needed, they write and send it, and chasing a late or stale update is yours.
- Run follow-the-sun handoffs. Deliver warm, high-fidelity handoffs across regions: current impact and severity, mitigation path and next actions, who is in the room, outstanding decisions, and what must not be dropped. The incoming commander confirms ownership before you step away—command never goes dark at a region boundary.
- Close the loop, on the clock. Maintain the timeline of facts as the incident runs rather than reconstructing it afterward—in a regulated business, that record has to hold up long after the call ends. Once mitigated, make sure a blameless retrospective is scheduled with a named owner and a timebox, and record where the cause sits—that choice sets which follow-up items are mandatory. Every action item needs a real ticket, one named accountable, a priority, and a category, delivered inside the agreed service level. If a postmortem produces nothing but low-priority items, treat that as a signal the analysis stopped early and escalate to SRE rather than passing it on.
- Automate the coordination away. Coordination is the part of this job that should eventually belong to a machine. Every manual prompt you send—the update that is due, the question nobody answered, the partner nobody contacted—is a candidate for automation, and the direction we are heading is AI handling the routine so commanders can spend their attention on judgment. You get us there by working to the decision trees, saying where they are wrong, and being honest about which of your instincts are actually rules.
Who You Are (Must-Haves)
- 4+ years commanding or co-commanding high-severity incidents in a production engineering, SRE, or technical operations environment.
- You direct technical responders under pressure without being the person writing the fix.
- You make and defend crisp severity and escalation decisions, and you take charge without waiting to be asked. Command means waking senior people at 03:00, interrupting an executive, and telling an experienced engineer to stop what they are doing—with an audience watching. It is a visible, directive role and it needs to be instinctive.
- You can read a dashboard and judge for yourself whether impact has actually stopped.
- You communicate clearly with engineers, executives, and partner-facing stakeholders—and you know the difference between briefing the comms function and speaking for the company.
- You are comfortable holding other teams to account in the moment, across a reporting line that is not yours, without turning it into friction.
- You thrive in a follow-the-sun model with clean cross-region handoffs.
- You understand FinTech concepts and the trust stakes of API-driven financial platforms.
- You use AI tools and agentic automation to reduce manual toil and speed up response.
- You will work a regional coverage window as part of a global 24x7 Incident Commander roster.
Who You Might Be (Nice-to-Haves)
- Formal incident command training—ITIL, Major Incident Management, or crisis management.
- Experience with modern incident management and on-call platforms.
- You have written severity rubrics, decision trees, escalation matrices, runbooks, or incident playbooks.
- You have commanded in game days, tabletop exercises, or incident simulations, not only in production.
- You have partnered with problem management or reliability program functions to roadmap incident follow-ups.
- Online securities trading or capital markets experience, or another regulated, market-hours-sensitive domain.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Alpaca
Explore Top Companies in this Space
Interactive Brokers
Financial Services / Online Brokerage / FinTech / Investment Technology
FP Markets
Financial Services / Online Trading / Forex / CFD Brokerage
Prosper
Banking / Finance / Lending & Brokerage
Freedom24
Financial Services / Brokerage / Investment Platforms
Alpaca
View Company ProfileAlpaca is a developer-first API brokerage platform and a fully registered, self-clearing US broker-dealer. Their mission is to open financial services to everyone on the planet. Instead of building a flashy consumer trading app like Robinhood, Alpaca built the underlying infrastructure. They offer powerful, modern APIs that allow algorithmic traders, hedge funds, and global FinTech startups to instantly build customizable investing applications. Whether a developer wants to write a Python script to automate their own stock trades, or an international startup wants to offer US stock investing to its users in Latin America or the Middle East, Alpaca provides the complete backend infrastructure to make it happen.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.
