---
title: "Head of Cloud Operations"
company: "HavocAI"
company_url: "https://www.remjobs.works/companies/havocai"
url: "https://www.remjobs.works/job/havocai-head-of-cloud-operations-9440c8f4-7765-406d-bc6b-876a395bb6ee"
apply_url: "https://jobs.ashbyhq.com/havocai/71676337-967d-4be1-943b-0999cdc173bc"
workplace: remote
location: "Remote "
employment_type: full-time
seniority: vp
role: operations
region: other
salary: "$200,000–$220,000 per year"
date_posted: 2026-09-23T03:13:26.012Z
first_seen_by_remjobs: 2026-09-23T03:38:10.318Z
---

# Head of Cloud Operations

**HavocAI** · Remote 

Salary: $200,000–$220,000 per year

Apply: https://jobs.ashbyhq.com/havocai/71676337-967d-4be1-943b-0999cdc173bc

## About the role

### About Us:

Havoc is a leader in all-domain collaborative autonomy. Its software-defined hardware approach powers military and commercial-grade autonomous systems across sea, air, and land to sense, decide, and act together in complex and contested environments. Havoc connects assets, enabling them to share information, adapt in real time, and continue operating even when communications are disrupted or denied. Havoc optimizes mission performance and minimizes human risk.

Havoc was founded in 2024 and headquartered in Providence, Rhode Island. Learn more at [Havoc: All-Domain Collaborative Autonomy](http://havocai.com/) .

#### About the Role

HavocAI is seeking a Head of Cloud Operations to own the change, release, incident, and reliability practices that keep our systems dependable, auditable, and compliant as we scale.

Reporting to the Director of Cloud and partnering closely with the ISSO and engineering teams, you will define how production changes are approved and deployed, how releases are coordinated, how incidents are managed, and how we maintain a trustworthy record of what is running across our environments.

You will also lead our SRE and DevOps teams, setting direction across reliability, infrastructure automation, CI/CD, observability, and safe delivery. This role requires someone who can build disciplined processes without creating unnecessary bureaucracy—using automation and engineering practices wherever possible to make the right way of working the easiest way of working.

The ideal candidate combines strong operational leadership with enough technical depth to challenge assumptions, make decisions under pressure, and translate security and compliance requirements into practical engineering processes.

#### What You’ll Do

##### SRE & DevOps Leadership

- Lead and manage the SRE and DevOps teams, setting technical and operational direction across reliability, automation, and safe delivery.

- Hire, coach, develop, and manage performance for engineers across both functions.

- Own reliability and delivery practices including SLIs, SLOs, error budgets, on-call health, CI/CD, and infrastructure automation.

- Establish clear ownership and operating expectations across cloud reliability and delivery.

- Partner with engineering leaders to identify systemic reliability risks and prioritize improvements.

- Build an engineering culture that balances speed, reliability, security, and operational discipline.

##### Change & Release Management

- Define and own change classes—including standard, normal, and emergency changes—with clear approval paths and requirements.

- Establish and operate an appropriate change approval process, including impact assessments, rollback plans, and approval records.

- Integrate change management with GitOps workflows, using merged, signed, peer-reviewed pull requests as the foundation of the change record.

- Own maintenance windows, freeze periods, and emergency-change processes, including retroactive approvals where appropriate.

- Ensure production changes are traceable to an approved request, approver, and rollback decision.

- Own the release calendar, versioning strategy, and promotion across environments and tenants.

- Establish pre-deployment verification requirements covering CI status, security scans, migrations, feature flags, and predefined rollback triggers.

- Coordinate releases across Cloud Platform, Backend, Autonomy, and Frontend teams to prevent conflicts and manage dependencies.

- Maintain complete, audit-ready deployment and release records.

##### Incident & Problem Management

- Own HavocAI’s incident management framework, including incident declaration, severity levels, escalation paths, and incident command.

- Establish clear authority and expectations for declaring and managing incidents.

- Run incident command during significant events, coordinating roles, communications, escalation, and stakeholder or customer notifications.

- Own on-call health, alert quality, escalation practices, and operational readiness.

- Lead blameless post-incident reviews and ensure remediation actions are assigned, tracked, and completed.

- Manage government-sponsor notification obligations and timelines for incidents affecting authorized systems, with company-wide scope beyond the IATT boundary.

- Establish clear distinctions between incidents and problems and drive analysis of recurring issues.

- Translate recurring operational issues into technical debt, reliability, and remediation priorities.

##### Configuration, Baselines & Compliance

- Maintain the authoritative record of deployed systems, including versions, digests, and dependencies.

- Keep deployment records synchronized with the ISSO’s system inventory.

- Establish and maintain system baselines and processes for detecting configuration drift.

- Produce audit-ready evidence for the ISSO, including change records, deployment logs, incident reports, and post-incident remediation actions.

- Own execution tracking against POA&M commitments, partnering with the ISSO to ensure remediation dates and engineering commitments are met.

- Translate security and compliance control language into practical engineering processes and clearly communicate engineering implementation back to security stakeholders.

- Build automation wherever possible to reduce manual compliance work and improve the reliability of operational evidence.

#### What We’re Looking For

- 8+ years of relevant experience across change management, release management, incident management, technical program management, service management, SRE, DevOps, platform engineering, or related disciplines.

- Demonstrated experience leading or managing SRE, DevOps, or Platform Engineering teams, including hiring, coaching, and performance management.

- Strong understanding of modern cloud operations, software delivery, infrastructure automation, and production reliability.

- Experience establishing and operating change, release, and incident management processes in complex technical environments.

- Proven ability to coordinate complex initiatives across engineering teams and stakeholders you do not directly manage.

- Technical fluency sufficient to evaluate and challenge engineering impact assessments, deployment strategies, rollback plans, and root-cause analyses.

- Ability to remain calm, decisive, and directive during active incidents and make sound decisions under pressure.

- Strong written communication skills, particularly for incident communications, post-incident reports, operational documentation, and executive updates.

- Ability to create scalable processes that provide appropriate control without unnecessarily slowing engineering teams.

- Strong ownership, judgment, and comfort operating in a fast-moving and ambiguous environment.

- Must be a U.S. Citizen and able to obtain and maintain a U.S. Government security clearance.

#### Nice to Have

- Prior incident command experience in a regulated, defense, government, or safety-relevant environment.

- Experience operating cloud systems subject to U.S. Government authorization or compliance requirements.

- Familiarity with POA&Ms, security authorization processes, configuration baselines, and audit evidence management.

- Experience implementing GitOps-based change and release processes.

- Knowledge of ITIL practices or equivalent hands-on experience developing effective service management processes.

- Experience with tools such as Jira, PagerDuty, status pages, runbook platforms, and incident management systems.

- Experience with Kubernetes, infrastructure as code, CI/CD platforms, observability systems, and modern cloud infrastructure.

#### What Success Looks Like

Within your first 12 months, you will have:

- Established an enforced, practical change and release management process that engineering teams consistently follow.

- Ensured every production change is traceable to the appropriate approval, deployment record, and rollback decision.

- Created a consistent incident management framework with clear severity levels, ownership, command structures, escalation paths, and communication standards.

- Established effective post-incident practices with remediation actions tracked through completion.

- Improved the health and effectiveness of SRE, DevOps, on-call, and reliability practices.

- Created an automated and trustworthy record of what is deployed across authorized environments.

- Kept deployment records synchronized with the ISSO’s inventory and maintained audit-ready operational evidence.

- Established effective tracking and execution against POA&M remediation commitments.

- Built operational processes that strengthen reliability and compliance without introducing unnecessary friction for engineering teams.

### Benefits:

- 100% Employer paid Health, Dental and Vision Insurance for you and your families

- Life Insurance (Employer Paid)

- Ability to participate in the companies 401k program (Matching)

- Unlimited PTO policy with an enforced 2 week minimum

- Equity Package

- Work / Home Office Stipend

- Global Entry

- 16 Week Paid Parental Leave

- Monthly Health and Wellness Stipend

### Our Values:

- Innovation: We are driven to break new ground. Every day presents an opportunity to challenge the status quo, think boldly, and deliver advanced solutions that transform the future of defense technology.

- Integrity: We hold ourselves to the highest ethical standards, ensuring transparency, accountability, and trust in all our actions and partnerships.

- Mission-Driven: We are focused on achieving impactful outcomes that align with our core mission—protecting lives through innovation.

- Forward-Leaning: We continuously seek out new opportunities and remain at the forefront of technological advancements. We embrace change and anticipate the challenges of tomorrow with confidence and creativity.

- Ownership of All Tasks: At HavocAI, no problem is too complex or too trivial. We believe that greatness comes from tackling the hardest challenges, but also in handling the smallest, sometimes thankless, tasks with the same level of commitment and care.

- Servant Leadership: We lead by serving others, whether it’s supporting our employees, partners, or the broader community. Empowering those around us is key to achieving long-term success and making a lasting impact.HavocAI is an Equal Opportunity Employer and is committed to creating an inclusive and diverse workplace. We welcome applicants from all backgrounds and do not discriminate based on race, color, religion, gender, sexual orientation, age, national origin, disability, veteran status, or any other legally protected status.

---

Source: HavocAI's own career page, read by RemJobs. Canonical HTML version: https://www.remjobs.works/job/havocai-head-of-cloud-operations-9440c8f4-7765-406d-bc6b-876a395bb6ee
