Back to Jobs
Havoc AI
AI & Machine Learning 3h ago

AI Infrastructure Engineer – Agents & ML Systems

Havoc AI
United StatesUnited States
Full-time
Not Disclosed
Mid-Level

Job Description

Key Skills Required

Master these to land this role

Machine LearningBestseller πŸ”₯
Learn in 42 Hours
MLOpsData Science & AnalyticsPython ScriptingAI Engineer

Want to know if you're a match for this job?

Calculate My Match Score

About the Role

As an AI Infrastructure Engineer – Agents & ML Systems, you will help build the internal AI infrastructure that allows HavocAI teams to use modern AI systems safely, reliably, and effectively. You will develop tools, services, pipelines, and integrations that connect large language models, agentic workflows, internal data sources, engineering systems, and ML workflows.

This role is ideal for a strong software or infrastructure engineer who is excited about the practical application of AI. You need not have worked on every part of the AI stack, but you should be curious, hands-on, and comfortable building production systems that connect models, tools, data, and users.

You will work on systems that help internal teams search and reason over company data, automate engineering workflows, support simulation and autonomy development, curate data for future model training, and evaluate AI systems before they are trusted in critical workflows. This is a high-impact role at the intersection of software engineering, AI infrastructure, developer tooling, data systems, and applied ML.

Job Responsibilities

  • Build internal AI infrastructure that connects LLMs and AI agents with internal tools, APIs, data sources, data lakes, telemetry stores, simulation tools, code repositories, documentation systems, logs, and engineering workflows.
  • Develop and maintain agentic AI systems for task automation, data analysis, engineering support, simulation workflows, and internal productivity.
  • Build tool integration and connector infrastructure for AI agents, including MCP and other emerging tool-use standards, spanning servers, tools, resources, prompts, connectors, and secure tool-use patterns.
  • Create pipelines for retrieval, RAG, context management, document processing, embeddings, and internal knowledge search.
  • Support ML infrastructure workflows such as data preparation, dataset curation, experiment tracking, model evaluation, fine-tuning support, and model deployment.
  • Build evaluation frameworks for agent performance, tool-use reliability, task success, model quality, regression testing, and failure analysis.
  • Develop observability, logging, tracing, auditability, monitoring, and debugging tools for AI agents, model calls, MCP tools, and ML pipelines.
  • Partner with Autonomy, Software, Data, Simulation, Product, and Operations teams to identify high-value AI use cases and turn them into reliable internal tools.
  • Secure agentic AI systems end-to-end with least-privilege tool access, sandboxed tool execution, prompt-injection and misuse mitigation, secrets management, human-in-the-loop approvals, and safe handling of sensitive and defense data.
  • Maintain documentation, reusable examples, templates, and best practices that help internal teams adopt AI tools safely and effectively.

Qualifications

  • Bachelor's degree in Computer Science, Engineering, Machine Learning, Data Science, Applied Mathematics, or a related technical field.
  • 3+ years of experience in software engineering, infrastructure engineering, ML infrastructure, backend systems, data engineering, developer tools, or related technical roles.
  • Strong programming experience in Python, TypeScript, Go, C++, or similar languages.
  • Experience building production software systems, APIs, services, data pipelines, or internal platforms.
  • Experience with, or strong interest in, LLM applications, AI agents, tool-using systems, RAG pipelines, or AI developer tools.
  • Familiarity with modern AI infrastructure concepts such as embeddings, vector search, prompt management, evaluation, model serving, fine-tuning, or MLOps.
  • Ability to integrate systems across APIs, databases, object stores, documents, logs, internal tools, and structured or unstructured data sources.
  • Strong understanding of production engineering fundamentals, including reliability, observability, testing, and maintainability.
  • Strong grounding in securing AI and agentic systems, including least-privilege tool access, prompt-injection and misuse mitigation, secrets management, and safe handling of sensitive data.
  • Ability to work across Software, Data, ML, Infrastructure, and Product teams.
  • Strong debugging skills and comfort working with complex distributed systems.
  • U.S. citizenship and the ability to obtain and maintain a security clearance.

Preferred Skills

  • Experience with MCP, including MCP servers, clients, tools, resources, prompts, or connector patterns.
  • Experience with LangGraph, LangChain, LlamaIndex, Semantic Kernel, OpenAI APIs, Anthropic APIs, local/open-weight models, or similar AI tooling.
  • Experience with vector databases, embeddings, retrieval systems, knowledge graphs, document processing, search, or RAG systems.
  • Experience with LLM fine-tuning, supervised fine-tuning, preference tuning, synthetic data generation, evaluation datasets, or model benchmarking.
  • Experience with Kubernetes, Docker, Ray, Airflow, Dagster, MLflow, Weights & Biases, Kafka, Postgres, S3-compatible storage, or similar infrastructure.
  • Experience building internal platforms, developer tools, workflow automation systems, or data/ML infrastructure.
  • Experience with human-in-the-loop workflows, approval systems, audit logs, policy enforcement, or safe tool invocation patterns.
  • Experience integrating AI tools with engineering workflows such as GitHub, CI/CD, issue trackers, documentation systems, simulation platforms, or data lakes.
  • Familiarity with autonomy, robotics, simulation, telemetry, perception, or defense technology workflows.
  • Experience building secure AI systems for sensitive, regulated, government, defense, or enterprise environments.
  • Active or prior security clearance.

How would you rate this job post?

See what other professionals think about this role.

banner

Havoc AI (operating under havocai.com, legally HavocAI, Inc.) is the premier, enterprise-grade all-domain collaborative autonomy platform, defense technology pioneer, and automated uncrewed systems orchestration powerhouse engineered to act as the definitive, high-velocity command-and-control (C2), edge intelligence, and distributed fleet coordination layer for modern military operations, contested logistics networks, and maritime security ecosystems globally. Founded by former military veterans and defense tech visionaries including Paul Lwin (a former U.S. Naval Flight Officer and aerospace engineer) alongside Timothy Rhatigan and Andrew Gregg, the company completely eliminates the severe systemic friction of modern autonomous operationsβ€”where defense systems rely on isolated, single-asset control paradigms, suffer from low-signal communications over degraded networks, and require extensive manpower to supervise minimal uncrewed arraysβ€”by deploying a sophisticated, multi-domain software operating matrix. Moving far beyond traditional, passive remote-control frameworks or isolated hardware drones, Havoc natively unifies a centralized "one-to-many" control layer (Havoc C2), an advanced edge intelligence optimization system (Havoc Insights), an interactive peer-to-peer data synchronization network (Havoc Connect) that holds operational stability across denied and communications-degraded (DDIL) environments, and a production-hardened edge operating system (Havoc OS) into a single high-availability all-domain intelligence workspace. Validated through more than 25,000 hours of autonomous real-world deployments and commanding over 100 fielded autonomous surface vessels (USVs) supporting critical U.S. Department of Defense (DoD) missions, the platform empowers a single warfighter to supervise thousands of heterogeneous autonomous assets across land, sea, and air simultaneously. Rapidly consolidating its all-domain vision, the high-growth enterprise has expanded its operational footprint through the strategic technical acquisitions of Mavrik and Teleo to cleanly bridge the gap between low-level edge hardware automation and high-level mission intent. Valued as an elite rising star in the defense technology landscape with a post-money valuation scaling past $900 million, the corporation has raised over $200 million in total institutional financingβ€”anchored by a monumental $100 million Series A funding matrix in May 2026 led by prominent asset managers including Boardman Bay Capital Management and Cobalt Capital, alongside significant heavy-tier backing from In-Q-Tel, B Capital, Scout Ventures, Outlander VC, SAIC, and defense titan Lockheed Martin. Under the hood, its technology core utilizes sophisticated sensor-fusion and tracking frameworks, distributed peer-to-peer tactical mesh protocols, and strict "human-on-the-loop" gating boundaries designed to maximize wide-area situational awareness and execution velocity without sacrificing operational safety or mission control. What sets Havoc AI apart is its uncompromising dedication to replacing fragile, siloed uncrewed vehicles with absolute real-world collaborative coordination predictability, hardware-agnostic software flexibility, and hardened battlefield resilience; by combining tactical edge computing with enterprise-tier data integration, the company remains a definitive cornerstone of modern algorithmic defense architecture and global military systems transformation.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More