Infrastructure Engineer (Platform & AI Agent Orchestration)
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
About CrewAI
CrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production. This role owns the infrastructure foundation that keeps it reliable, secure, and fast.
The Role
You'll build and operate the platform infrastructure behind CrewAI's cloud and enterprise deployments. You'll work across multiple hyperscalers—AWS, Azure, and GCP. You’ll work on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. Your job is to make the product and runtime teams faster while making customers’ production environments safer.
This is not a pure DevOps support role. You'll write code, improve systems, design deployment paths, harden production, and build the internal platform that lets CrewAI scale and scale our customer deployments.
What You'll Do
- Own and improve the infrastructure that runs CrewAI's platform: AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related services.
- Build and maintain CI/CD pipelines for build, test, image publishing, migrations, environment promotion, rollbacks, and deploy safety.
- Improve reliability across cloud and enterprise deployments: health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks—and own the front-line on-call rotation and its SLAs.
- Partner with runtime engineers on Celery/FastAPI/Redis workloads and with product engineers on Rails/Solid Queue/Postgres production behavior.
- Manage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems.
- Harden security and compliance posture across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least-privilege access.
- Build the tooling and automation that lets field engineers and customers run self-hosted installs themselves—Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks—so engineering does fewer hands-on installs over time.
- Reduce operational toil by automating recurring workflows and making deployments boring.
Requirements
What We're Looking For
- Strong infrastructure/platform engineering experience in production SaaS environments.
- Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services.
- Experience with ECS and/or Kubernetes; Helm experience is a strong plus.
- Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production.
- Strong debugging instincts across app, infra, network, deploy, and dependency layers.
- Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access.
- Ability to write reliable automation in Python, Ruby, Go, Bash, or similar.
- Calm, rigorous approach to incidents, rollbacks, migrations, and production change management.
Bonus
- Experience with AI/agent platforms, workflow runtimes, or high-volume async execution systems.
- Experience supporting enterprise/self-hosted deployments.
- Terraform or other IaC experience.
- SRE background: SLOs, incident review, capacity planning, load testing.
- Familiarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
Explore Top Companies in this Space
Metaview
AI Recruiting Automation / Interview Intelligence SaaS / Talent Acquisition Technology / B2B Enterprise Software
AppZen
FinTech & Autonomous Finance / Agentic AI & AP Automation / Corporate Expense Auditing / Enterprise Software SaaS
Virtuous
Nonprofit CRM & Responsive Fundraising Software / Enterprise Marketing Automation SaaS / Philanthropic AI & Wealth Intelligence Platforms / Digital Giving Infrastructure
HappyCo
PropTech & Real Estate Technology / AI Maintenance Automation / Property Operations SaaS / B2B Enterprise Software
CrewAI
View Company ProfileCrewAI (operating at crewai.com) is an AI-driven automation platform engineered for streamlining workflows and enhancing productivity through collaborative AI agents. Founded to address the inefficiencies of traditional task management systems, CrewAI leverages advanced AI to orchestrate multi-agent workflows, enabling seamless coordination between tools and processes. Under the hood, the platform integrates AI agents with APIs and human inputs, automating repetitive tasks while maintaining flexibility for complex decision-making. This allows businesses and developers to optimize operations, reduce manual effort, and accelerate project timelines. While specific funding details remain undisclosed, CrewAI positions itself as a cutting-edge solution for enterprises seeking to harness AI for scalable automation.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.
