Back to Jobs
Skydio
Development 6h ago

Staff Site Reliability Engineer

Skydio
SanSan
CaliforniaCalifornia
Full-time
$240,000-$300,000
Senior-Level

Job Description

Key Skills Required

Master these to land this role

DevOpsBestseller 🔥
Learn in 63 Hours
KubernetesCI/CDAWSTerraform

Want to know if you're a match for this job?

Calculate My Match Score

Skydio is the leading US drone company and the world leader in autonomous flight, the key technology for the future of drones and aerial mobility. The Skydio team combines deep expertise in artificial intelligence, best-in-class hardware and software product development, operational excellence, and customer obsession to empower a broader, more diverse audience of drone users.

About the team:

Skydio’s Cloud infrastructure team is here to ensure that the Skydio Cloud platform is always available to our customers when they need it: whether that’s performing a routine bridge inspection or it’s aiding in rescue operations during a natural disaster. With tens of thousands of drones operating all over the world, we are constantly improving the continuous delivery and deployment of our infrastructure. Our technology helps save lives. You’ll play a critical role in keeping the infrastructure behind it reliable, scalable, and available when it matters most.

About the role:

We are looking for a hands-on Staff Site Reliability Engineer to build, operate, and scale the cloud infrastructure that powers our products. This role is focused on owning production infrastructure, including Kubernetes, AWS, infrastructure as code, CI/CD, observability, networking, and reliability.

You don't need to be an expert in every area, but you should have strong Kubernetes and cloud fundamentals with meaningful depth in at least one infrastructure domain.

How you'll make an impact:

  • Build, operate, and troubleshoot production Kubernetes/EKS clusters.
  • Perform Kubernetes upgrades, node rollouts, and cluster maintenance.
  • Build and manage AWS infrastructure including VPCs, networking, subnets, load balancers, IAM, EKS, databases, and storage.
  • Define and maintain infrastructure using Terraform.
  • Build and operate CI/CD and deployment infrastructure.
  • Troubleshoot production issues across Kubernetes, AWS, Linux, networking, and databases.
  • Build monitoring, alerting, and observability for critical infrastructure.
  • Participate in on-call rotations and respond to production incidents.
  • Identify and solve infrastructure scaling and reliability problems.
  • Automate operational work using Python, Go, or similar languages.
  • Help expand infrastructure across new regions and deployment environments.

What makes you a good fit:

  • 8+ years of experience as a Site Reliability Engineer, Platform Engineer, DevOps, Production Engineer or equivalent infrastructure role.
  • Strong hands-on experience operating Kubernetes, not simply deploying applications to existing clusters.
  • Experience managing Kubernetes/EKS upgrades and production clusters.
  • Strong AWS fundamentals, including VPCs, public/private subnets, networking, load balancers, EKS, IAM, and databases.
  • Production experience with Terraform or similar infrastructure-as-code tooling.
  • Experience owning or maintaining CI/CD and deployment systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins.
  • Experience diagnosing production infrastructure and networking problems.
  • Experience solving meaningful scaling or reliability challenges.
  • This position requires access to export-controlled technical data, restricted government information, and/or information systems subject to U.S. government security and access-control requirements. Employment in this role is contingent upon verification of U.S. person status and the ability to access controlled or restricted information as required for the position.

Bonus points:

  • Helm and GitOps experience.
  • Datadog or similar observability tooling.
  • PostgreSQL/database operations experience.
  • Multi-region infrastructure experience.
  • On-premises or disconnected deployment experience.
  • Streaming or high-throughput distributed systems experience.

How would you rate this job post?

See what other professionals think about this role.

banner

Skydio (operating at skydio.com) is a drone manufacturer and autonomous flight technology platform engineered for safer, smarter, and faster work. Founded in 2014 by Adam Bry, Abe Bachrach, and Matt Donahoe and headquartered in San Mateo, California, Skydio leverages breakthrough artificial intelligence to create flying drones that are used by consumers, enterprises, and government customers. Under the hood, Skydio's drones use AI to navigate complex environments and capture high-quality video and images. This allows customers to conduct inspections, capture data, and complete tasks more efficiently and safely. Backed by top investors and strategic partners, including Andreessen Horowitz, Linse Capital, N47, IVP, Playground, and NVIDIA, Skydio has raised $562 million in funding.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More