Back to Jobs
Roboflow
AI & Machine Learning 3h ago

Infrastructure Engineer

Roboflow
San FranciscoSan Francisco
Full-time
$165,000 - $200,000 base
Mid-Level

Job Description

Key Skills Required

Master these to land this role

Python2h 41mFree Trial ✨
Start 10-Day Free Trial
DevOps1h 38mFree Trial ✨
Start 10-Day Free Trial
Machine Learning41mFree Trial ✨
Start 10-Day Free Trial
TensorFlowKubernetes

Want to know if you're a match for this job?

Calculate My Match Score

As a member of our infrastructure team, you'll be at the heart of a fast-paced startup environment. Your primary focus will be on striking the right balance between rapid delivery, high reliability, and robust security. This isn't a traditional, siloed role; you'll need to wear many hats—acting as an infrastructure engineer one moment, and a developer, or even a security analyst.

You will be securing, scaling, and maintaining the core infrastructure that powers our product. This includes our cloud architecture, databases, file storage, search clusters, microservices, and machine learning pipelines. You'll work closely with our product team and collaborate across the company on product, operations, and customer-facing projects, constantly context-switching to solve the next critical challenge.

Skillset

We're looking for a versatile engineer excited by high-impact challenges. At Roboflow, we are AI-native: we expect our team to use AI to accelerate everything from writing code and fixing bugs to analyzing security, cost, and performance. Experience in some or all of the following areas will be crucial:

  • Production experience with Kubernetes: Building and managing containerized applications at scale.

  • Infrastructure-as-Code (IaC): Using Terraform, Helm charts, bash scripting, and Python to automate everything.

  • Scale & Site Reliability: Operating, monitoring, and scaling large-scale applications (especially in ML/AI) in AWS and/or GCP.

  • Development Skills: Proficiency in Node.js and Python, with the ability to collaborate with full-stack developers on designing and operating SaaS applications.

  • ML/Big Data Ops: Hands-on experience with the infrastructure required for machine learning at scale (GPUs, Docker, Kubernetes) and familiarity with libraries like PyTorch or Tensorflow.

  • CI/CD Automation: Experience with tools like GitHub Actions or Spacelift to build and deploy code efficiently.

  • Pragmatic Security: Awareness of security best practices for cloud operations and how they can be applied to startup environments.

  • AI-Native Engineering: Leveraging LLMs and AI tools to accelerate the development lifecycle—from writing and refactoring code to identifying security vulnerabilities and optimizing infrastructure costs.

A Glimpse of Your Work

No two days will be the same. Your tasks will be a blend of strategic projects and hands-on implementation. Examples include:

  • Running and optimizing a high-availability machine learning inference service.

  • Collaborating with customer security teams to ensure secure integration.

  • Developing creative IaC solutions to scale our platform cost-effectively.

  • Working with the engineering team to define SLOs/SLAs and participating in incident response.

  • Improving the Observability and Alerting stack and the processes built around it.

  • Diving deep into our stack to identify and act on cost-optimization opportunities.

  • Contributing code (Python, JavaScript, etc.) as part of a team designing and deploying new product features.

  • Fixing security vulnerabilities and bugs

  • Hardening our systems and processes to meet SOC 2, HIPAA, and GDPR requirements, making us audit-ready.

  • Participating in an on-call rotation to ensure platform reliability.

📅 Within one week, you will…

  • Learn all about computer vision, our product, company, customers, and vision.

  • Ship something substantial to an end user

  • Start learning our infrastructure and security practices.

📅 Within one month, you will…

  • Onboard in person with your manager

  • Build your first computer vision project with Roboflow (if you haven't already)

  • Start contributing to infra-as-code

  • Start working with customers to help with their security questions and onboarding

  • Understand the architecture of Roboflow

📅 Within six months, you will…

  • Attend your first company onsite

  • Be ramped up on other relevant parts of the Roboflow product.

How would you rate this job post?

See what other professionals think about this role.

banner

Roboflow is a pioneering artificial intelligence company that provides a comprehensive, end-to-end platform for building, training, and deploying computer vision models. Founded in 2019 by Joseph Nelson and Brad Dwyer, the company is on a mission to democratize visual intelligence, allowing any developer to give their software the sense of sight. Under the hood, Roboflow offers a massive suite of developer-first tools—including AI-assisted image annotation, dataset management, automated preprocessing, one-click model training, and highly scalable cloud or edge deployment via API. Their primary target audience spans over a million individual software engineers, machine learning researchers, and massive Fortune 100 enterprises across industries like manufacturing, logistics, and healthcare who need to automate complex visual tasks such as defect detection and inventory tracking. What sets Roboflow apart in the rapidly evolving AI landscape is its incredibly vibrant open-source ecosystem, its low-code interface that drastically reduces the friction of model training, and its ability to seamlessly bridge the gap between raw image data and production-ready visual AI applications.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More