Head of Cloud Operations
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Havoc is a leader in all-domain collaborative autonomy, specializing in military and commercial-grade autonomous systems for sea, air, and land. Their software-defined hardware approach enables real-time information sharing, adaptability, and continuous operation even in contested environments.
About the Role
HavocAI is seeking a Head of Cloud Operations to oversee change, release, incident, and reliability practices, ensuring systems remain dependable, auditable, and compliant as the company scales.
Reporting to the Director of Cloud and collaborating closely with the ISSO and engineering teams, this role involves defining approval and deployment processes for production changes, coordinating releases, managing incidents, and maintaining a trustworthy operational record.
The ideal candidate combines strong operational leadership with deep technical expertise to challenge assumptions, make high-pressure decisions, and translate security/compliance requirements into practical engineering processes.
What You’ll Do
SRE & DevOps Leadership
- Lead and manage SRE and DevOps teams, setting technical and operational direction across reliability, automation, and safe delivery.
- Hire, coach, and manage performance for engineers across both functions.
- Own reliability and delivery practices, including SLIs, SLOs, error budgets, on-call health, CI/CD, and infrastructure automation.
- Establish clear ownership and expectations for cloud reliability and delivery.
- Partner with engineering leaders to identify and prioritize systemic reliability risks.
- Build an engineering culture balancing speed, reliability, security, and operational discipline.
Change & Release Management
- Define and own change classes (standard, normal, emergency) with clear approval paths and requirements.
- Establish change approval processes, including impact assessments, rollback plans, and approval records.
- Integrate change management with GitOps workflows, using merged, signed, peer-reviewed pull requests as the foundation.
- Own maintenance windows, freeze periods, and emergency-change processes.
- Ensure production changes are traceable to approved requests, approvers, and rollback decisions.
- Manage release calendars, versioning strategies, and environment promotions.
- Establish pre-deployment verification requirements covering CI status, security scans, migrations, feature flags, and rollback triggers.
- Coordinate releases across Cloud Platform, Backend, Autonomy, and Frontend teams.
- Maintain complete, audit-ready deployment and release records.
Incident & Problem Management
- Own HavocAI’s incident management framework, including severity levels, escalation paths, and incident command.
- Establish clear authority and expectations for incident declaration and management.
- Run incident command during significant events, coordinating roles, communications, and stakeholder notifications.
- Own on-call health, alert quality, escalation practices, and operational readiness.
- Lead blameless post-incident reviews and ensure remediation actions are tracked and completed.
- Manage government-sponsor notification obligations for incidents affecting authorized systems.
- Distinguish between incidents and problems, driving analysis of recurring issues.
- Translate recurring operational issues into technical debt and remediation priorities.
Configuration, Baselines & Compliance
- Maintain authoritative records of deployed systems, including versions, digests, and dependencies.
- Synchronize deployment records with the ISSO’s system inventory.
- Establish system baselines and processes for detecting configuration drift.
- Produce audit-ready evidence for the ISSO, including change records, deployment logs, and incident reports.
- Track execution against POA&M commitments, ensuring remediation dates and commitments are met.
- Translate security/compliance controls into practical engineering processes.
- Build automation to reduce manual compliance work and improve reliability.
What We’re Looking For
- 8+ years of experience in change management, release management, incident management, technical program management, service management, SRE, DevOps, or platform engineering.
- Experience leading or managing SRE, DevOps, or Platform Engineering teams, including hiring, coaching, and performance management.
- Strong understanding of modern cloud operations, software delivery, infrastructure automation, and production reliability.
- Experience establishing and operating change, release, and incident management processes in complex environments.
- Proven ability to coordinate complex initiatives across engineering teams and stakeholders.
- Technical fluency to evaluate impact assessments, deployment strategies, rollback plans, and root-cause analyses.
- Ability to remain calm, decisive, and directive during active incidents.
- Strong written communication skills for incident reports, documentation, and executive updates.
- Ability to create scalable processes that provide control without slowing engineering teams.
- Strong ownership, judgment, and comfort in fast-moving, ambiguous environments.
- Must be a U.S. Citizen and obtain/maintain a U.S. Government security clearance.
Nice to Have
- Incident command experience in regulated, defense, government, or safety-relevant environments.
- Experience operating cloud systems subject to U.S. Government authorization or compliance requirements.
- Familiarity with POA&M, security authorization processes, configuration baselines, and audit evidence management.
- Experience with GitOps-based change and release processes.
- Knowledge of ITIL practices or equivalent hands-on experience in service management.
- Experience with tools like Jira, PagerDuty, status pages, runbook platforms, and incident management systems.
- Experience with Kubernetes, infrastructure as code, CI/CD platforms, observability systems, and modern cloud infrastructure.
What Success Looks Like
Within your first 12 months, you will have:
- Established an enforced, practical change and release management process.
- Ensured every production change is traceable to the appropriate approval and deployment record.
- Created a consistent incident management framework with clear severity levels, ownership, and communication standards.
- Established effective post-incident practices with remediation actions tracked through completion.
- Improved the health and effectiveness of SRE, DevOps, on-call, and reliability practices.
- Created an automated and trustworthy record of deployed systems across authorized environments.
- Kept deployment records synchronized with the ISSO’s inventory and maintained audit-ready evidence.
- Established effective tracking and execution against POA&M remediation commitments.
- Built operational processes that strengthen reliability and compliance without unnecessary friction.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Havoc AI
Explore Top Companies in this Space
K2 Space
Aerospace and Defense / Satellite Manufacturing / Space Technology / Commercial Spaceflight
ICEYE
Satellite Technology / Earth Observation / Defense & Aerospace / Space Industry
Ursa Major
Aerospace & Defense / Space Technology / Advanced Manufacturing
Swarm Aero
Aerospace & Defense / Autonomous Systems / Artificial Intelligence / Military Technology
Havoc AI
View Company ProfileHavoc AI (operating under havocai.com, legally HavocAI, Inc.) is the premier, enterprise-grade all-domain collaborative autonomy platform, defense technology pioneer, and automated uncrewed systems orchestration powerhouse engineered to act as the definitive, high-velocity command-and-control (C2), edge intelligence, and distributed fleet coordination layer for modern military operations, contested logistics networks, and maritime security ecosystems globally. Founded by former military veterans and defense tech visionaries including Paul Lwin (a former U.S. Naval Flight Officer and aerospace engineer) alongside Timothy Rhatigan and Andrew Gregg, the company completely eliminates the severe systemic friction of modern autonomous operations—where defense systems rely on isolated, single-asset control paradigms, suffer from low-signal communications over degraded networks, and require extensive manpower to supervise minimal uncrewed arrays—by deploying a sophisticated, multi-domain software operating matrix. Moving far beyond traditional, passive remote-control frameworks or isolated hardware drones, Havoc natively unifies a centralized "one-to-many" control layer (Havoc C2), an advanced edge intelligence optimization system (Havoc Insights), an interactive peer-to-peer data synchronization network (Havoc Connect) that holds operational stability across denied and communications-degraded (DDIL) environments, and a production-hardened edge operating system (Havoc OS) into a single high-availability all-domain intelligence workspace. Validated through more than 25,000 hours of autonomous real-world deployments and commanding over 100 fielded autonomous surface vessels (USVs) supporting critical U.S. Department of Defense (DoD) missions, the platform empowers a single warfighter to supervise thousands of heterogeneous autonomous assets across land, sea, and air simultaneously. Rapidly consolidating its all-domain vision, the high-growth enterprise has expanded its operational footprint through the strategic technical acquisitions of Mavrik and Teleo to cleanly bridge the gap between low-level edge hardware automation and high-level mission intent. Valued as an elite rising star in the defense technology landscape with a post-money valuation scaling past $900 million, the corporation has raised over $200 million in total institutional financing—anchored by a monumental $100 million Series A funding matrix in May 2026 led by prominent asset managers including Boardman Bay Capital Management and Cobalt Capital, alongside significant heavy-tier backing from In-Q-Tel, B Capital, Scout Ventures, Outlander VC, SAIC, and defense titan Lockheed Martin. Under the hood, its technology core utilizes sophisticated sensor-fusion and tracking frameworks, distributed peer-to-peer tactical mesh protocols, and strict "human-on-the-loop" gating boundaries designed to maximize wide-area situational awareness and execution velocity without sacrificing operational safety or mission control. What sets Havoc AI apart is its uncompromising dedication to replacing fragile, siloed uncrewed vehicles with absolute real-world collaborative coordination predictability, hardware-agnostic software flexibility, and hardened battlefield resilience; by combining tactical edge computing with enterprise-tier data integration, the company remains a definitive cornerstone of modern algorithmic defense architecture and global military systems transformation.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.

