Back to Jobs
Havoc AI
AI & Machine Learning 2d ago

Machine Learning Cloud Infrastructure Engineer

Havoc
United StatesUnited States
Full-time
Not Disclosed
Mid-Level

Job Description

Key Skills Required

Master these to land this role

Python2h 41mFree Trial ✨
Start 10-Day Free Trial
KubernetesInfrastructure as CodeAWSMLOps

Want to know if you're a match for this job?

Calculate My Match Score

As a Machine Learning Cloud Infrastructure Engineer, you will build and operate the infrastructure HavocAI teams use to train, evaluate, deploy, and monitor machine learning models safely and reliably.

You will develop the pipelines, services, platforms, and integrations connecting data lakes, telemetry stores, simulation environments, training workloads, cloud compute, and deployed models. Your work will enable autonomy, data, and software engineers to move efficiently from raw field and simulation data to reproducible datasets, scalable training, rigorous evaluation, and production deployment.

This role is ideal for a strong software, infrastructure, or data engineer with hands-on ML experience who enjoys building production systems at the intersection of data, models, compute, and autonomy. You should be comfortable operating in a fast-paced environment, solving ambiguous infrastructure problems, and building systems that must remain scalable, observable, secure, and reliable.

ML Infrastructure & Pipelines

  • Build pipelines that transform raw multi-modal data—including telemetry, imagery, video, sensor, and simulation data—into curated, versioned training datasets.

  • Develop reproducible training and evaluation workflows that scale across cloud compute and GPU resources.

  • Build and maintain model deployment infrastructure for packaging, serving, inference, versioning, and rollback.

  • Implement experiment tracking, dataset lineage, model versioning, and other capabilities required for reproducible ML development.

  • Own data schema versioning and migration across pipelines, data lakes, and services as datasets and models evolve.

Cloud Platform & Infrastructure

  • Design, build, and operate scalable AWS infrastructure using Infrastructure as Code.

  • Build and maintain Kubernetes/EKS workloads and containerized environments for training, batch processing, evaluation, and model serving.

  • Develop self-service tooling and paved paths for compute scheduling, storage, data access, training, and deployment.

  • Improve utilization, scalability, and cost efficiency across cloud and accelerator infrastructure.

  • Build infrastructure that enables engineering teams to launch workloads safely without unnecessary operational overhead.

Reliability, Evaluation & Observability

  • Build evaluation frameworks and regression testing for model quality, dataset integrity, and pipeline correctness.

  • Establish quality and reliability signals that help determine whether models are ready for production use.

  • Develop monitoring, logging, tracing, and observability across training jobs, data pipelines, and deployed models.

  • Diagnose and resolve performance, scaling, reliability, and infrastructure bottlenecks.

  • Maintain high standards for automation, testing, documentation, and operational readiness.

Cross-Functional Engineering

  • Partner with Autonomy, Software, Data, Simulation, and Security teams to build ML infrastructure spanning edge data capture through cloud training and model deployment.

  • Contribute to CI/CD and release processes for models, datasets, and ML pipelines.

  • Translate engineering requirements into scalable platform capabilities that can support multiple teams and use cases.

  • Incorporate feedback from engineers and users to continuously improve ML development workflows.

Security & Data Management

  • Implement secure infrastructure practices, including IAM least privilege, secrets management, access controls, and secure handling of sensitive and defense-related data.

  • Build data and ML workflows with reproducibility, traceability, and appropriate controls from the start.

  • Partner with security and infrastructure teams to ensure ML systems meet applicable operational and compliance requirements.

How would you rate this job post?

See what other professionals think about this role.

banner

Havoc AI (operating under havocai.com, legally HavocAI, Inc.) is the premier, enterprise-grade all-domain collaborative autonomy platform, defense technology pioneer, and automated uncrewed systems orchestration powerhouse engineered to act as the definitive, high-velocity command-and-control (C2), edge intelligence, and distributed fleet coordination layer for modern military operations, contested logistics networks, and maritime security ecosystems globally. Founded by former military veterans and defense tech visionaries including Paul Lwin (a former U.S. Naval Flight Officer and aerospace engineer) alongside Timothy Rhatigan and Andrew Gregg, the company completely eliminates the severe systemic friction of modern autonomous operations—where defense systems rely on isolated, single-asset control paradigms, suffer from low-signal communications over degraded networks, and require extensive manpower to supervise minimal uncrewed arrays—by deploying a sophisticated, multi-domain software operating matrix. Moving far beyond traditional, passive remote-control frameworks or isolated hardware drones, Havoc natively unifies a centralized "one-to-many" control layer (Havoc C2), an advanced edge intelligence optimization system (Havoc Insights), an interactive peer-to-peer data synchronization network (Havoc Connect) that holds operational stability across denied and communications-degraded (DDIL) environments, and a production-hardened edge operating system (Havoc OS) into a single high-availability all-domain intelligence workspace. Validated through more than 25,000 hours of autonomous real-world deployments and commanding over 100 fielded autonomous surface vessels (USVs) supporting critical U.S. Department of Defense (DoD) missions, the platform empowers a single warfighter to supervise thousands of heterogeneous autonomous assets across land, sea, and air simultaneously. Rapidly consolidating its all-domain vision, the high-growth enterprise has expanded its operational footprint through the strategic technical acquisitions of Mavrik and Teleo to cleanly bridge the gap between low-level edge hardware automation and high-level mission intent. Valued as an elite rising star in the defense technology landscape with a post-money valuation scaling past $900 million, the corporation has raised over $200 million in total institutional financing—anchored by a monumental $100 million Series A funding matrix in May 2026 led by prominent asset managers including Boardman Bay Capital Management and Cobalt Capital, alongside significant heavy-tier backing from In-Q-Tel, B Capital, Scout Ventures, Outlander VC, SAIC, and defense titan Lockheed Martin. Under the hood, its technology core utilizes sophisticated sensor-fusion and tracking frameworks, distributed peer-to-peer tactical mesh protocols, and strict "human-on-the-loop" gating boundaries designed to maximize wide-area situational awareness and execution velocity without sacrificing operational safety or mission control. What sets Havoc AI apart is its uncompromising dedication to replacing fragile, siloed uncrewed vehicles with absolute real-world collaborative coordination predictability, hardware-agnostic software flexibility, and hardened battlefield resilience; by combining tactical edge computing with enterprise-tier data integration, the company remains a definitive cornerstone of modern algorithmic defense architecture and global military systems transformation.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More
Machine Learning Cloud Infrastructure Engineer at Havoc | HireSkys