Senior Storage Engineer at Runpod
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Runpod is the AI Developer Cloud, serving over one million developers—from indie researchers to teams running frontier models in production. The platform has processed more than 20 billion inference requests and recently closed a $100M Series A in June 2026. Runpod is at a pivotal moment in AI infrastructure, building the platform the next generation of developers will rely on.
The company operates as a small, remote-first team, emphasizing ownership, speed, and delivering work that impacts millions of developers daily.
You will join the Infrastructure organization, specifically the team managing Runpod's multi-region storage ecosystem, including network volumes, local NVMe, and S3-compatible object storage. Storage is critical at Runpod, influencing cold start velocity, training job data streaming efficiency, and the reliable persistence of model weights and checkpoints.
This senior, hands-on role operates in close coordination with SRE, networking, and supply chain teams, as well as global hardware partners. The team eliminates the divide between architectural design and operational execution, with engineers defining systems, writing code, optimizing infrastructure, managing on-call rotations, and leading capacity planning.
Runpod is hiring a Senior Storage Engineer to contribute to the design, scaling, and reliability of Runpod's storage platform. This is a hands-on engineering role, not just an administrative one. You will implement and maintain distributed storage deployments, writing code and automation to operate them.
Your responsibilities will include shaping Runpod's storage strategy for the next several years, deciding on distributed storage systems, data tiering and placement, network tuning, and hardware procurement at petabyte scale. You will have the freedom to innovate, automate manual operations, design metrics and SLOs, and lead capacity expansions and migrations end-to-end. Improvements you make will directly impact faster cold starts, faster jobs, and fewer incidents for over a million developers.
Responsibilities:
- Own capacity, durability, availability, and performance characteristics of network volumes, local NVMe, and S3-compatible object storage.
- Tune the full I/O path: device and filesystem configuration, caching and read-ahead strategies, replication and erasure coding trade-offs, and client-side mount behavior.
- Diagnose hard performance problems end-to-end.
- Lead capacity expansions, hardware refreshes, migrations, and rebalances without customer-visible disruption.
- Work with Runpod and partner networking teams to design and tune network paths for storage, including high-throughput east-west fabric, MTU and jumbo frames, congestion and flow control, multipath, and NIC/offload configuration.
- Understand and optimize RDMA/RoCE and high-speed IB/Ethernet fabrics for storage traffic.
- Collaborate with network engineering on topology decisions, oversubscription ratios, and cross-region data movement.
- Write production code in Go, Python, or similar for storage control-plane services, provisioning workflows, data movement pipelines, and monitoring.
- Build against and extend APIs: Runpod’s control plane, S3-compatible interfaces, CSI drivers, Kubernetes APIs, and vendor/cloud provider APIs.
- Automate operations to reduce manual runbooks.
- Treat infrastructure as code and participate in code review, testing, and CI.
- Instrument the storage fleet to ensure legible behavior, including IOPS, throughput, latency, error and retry rates, capacity utilization, and per-tenant consumption.
- Build dashboards, SLOs, and alerts to catch degradation before customers notice.
- Participate in on-call rotations for storage systems and drive blameless post-incident follow-through.
Requirements:
- 8+ years in infrastructure, storage, or systems engineering, with substantial ownership of production storage at scale.
- Deep, practical experience with at least one distributed storage system, such as Ceph, MinIO, Lustre, GPFS/Spectrum Scale, MooseFS, WekaFS, VAST, or ZFS-based systems.
- Strong Linux internals and storage-stack knowledge, including block layer, filesystems, NVMe, page cache, I/O schedulers, NFS/SMB, iSCSI/NVMe-oF.
- Building and/or operating S3-compatible object storage services.
- Solid networking fundamentals with specific experience tuning networks for storage workloads.
- Proficiency in writing and shipping production code in Go, Python, Rust, or similar (not just scripting).
- Hands-on experience with observability tooling (Prometheus, Grafana, Datadog, or equivalent), including designing metrics.
- A track record of performance analysis and debugging under real production pressure.
- Self-starting with general direction; capable of diagnosing problems and proposing solutions independently.
- Continuous improvement mindset; eliminating recurring toil and manual steps.
- Ownership mentality; following problems across team boundaries to resolution.
- Collaborative and low-ego but high-confidence.
Preferred:
- Storage for AI/ML workloads, including checkpointing, dataset streaming, model weight distribution, GPU-adjacent data locality, and GPUDirect Storage.
- Kubernetes storage internals, such as CSI drivers, PV/PVC lifecycle, StatefulSets, and local persistent volumes.
- Bare-metal and colocation experience, including hardware selection, vendor management, firmware, and physical failure domains.
- Multi-tenant environments where isolation, fairness, and QoS are critical.
- Experience in a fast-growing cloud or infrastructure provider.
What You’ll Receive:
- The competitive base pay for this position ranges from $180,000 - $260,000.
- Meaningful equity in a fast-growing company; everyone on the team receives stock options.
- Generous medical, dental, and vision plans.
- Flexible PTO to recharge.
- Remote work first with an inclusive, collaborative team utilizing Slack for communication.
- A $1,200 Home Office & Equipment Stipend to set up your ideal workspace.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at RunPod
Explore Top Companies in this Space
Lambda
AI Infrastructure & GPU Cloud / Supercomputing & Data Centers / Machine Learning Platforms / High-Performance Computing (HPC)
Belkins
B2B Marketing Services / Sales Development / Lead Generation / Appointment-Setting
WaterAid
Nonprofit / Public Health / Development / Advocacy
PDI Technologies
Software Development / Enterprise Software / Petroleum & Energy / Retail Technology
RunPod
View Company ProfileRunPod is a premier, enterprise-grade GPU cloud platform engineered to orchestrate massive-scale AI/ML compute ecosystems and intelligent, frictionless infrastructure-delivery workflows. Operating as a developer-first, high-throughput cloud hub, the company eliminates the operational friction of traditional, legacy-cloud providers—which frequently lock users into rigid, overpriced, and manual-heavy compute models—by seamlessly deploying advanced serverless GPU telemetry, rigorous multi-region container-orchestration architectures, and cohesive cross-platform scaling frameworks. Moving beyond rigid legacy VM-based paradigms, RunPod empowers over 500,000 global developers, researchers, and Fortune 500 enterprises to dynamically synchronize their training, fine-tuning, and inference pipelines with elite, autonomous, and cost-effective execution. Under the hood, their sophisticated proprietary data infrastructure natively manages complex multi-node cluster ingestion (A100/H100/H200 architectures), instantaneous autoscaling endpoint routing, and automated ephemeral-pod management, providing the necessary operational foundation to support everything from individual experimental models to large-scale, trillion-parameter distributed training. What sets RunPod apart is its uncompromising dedication to frictionless compute orchestration; by bridging the gap between highly technical, performance-intensive GPU-infrastructure demands and accessible, low-latency deployment interfaces, the platform empowers modern AI-native organizations to radically accelerate their production-AI velocity, eliminate prohibitive infrastructure-management bottlenecks, and build an unassailable foundation for continuous commercial and institutional dominance in the modern, AI-transformed digital landscape.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.
