Back to Jobs
RunPod
Development 9h ago

Technical Program Manager (TPM)

RunPod
United StatesUnited States
Full-time
$140,000 - $165,000
Senior-Level

Job Description

Key Skills Required

Master these to land this role

DevOps1h 38mFree Trial ✨
Start 10-Day Free Trial
Project Management3h 35mFree Trial ✨
Start 10-Day Free Trial
Backend42mFree Trial ✨
Start 10-Day Free Trial
AI EngineerData Science & Analytics

Want to know if you're a match for this job?

Calculate My Match Score

Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.

We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.

We're seeking a Technical Program Manager (TPM) who combines strong technical depth with exceptional organizational and communication skills. The ideal candidate is a self-starter who thrives in a fast-paced, high-growth environment and drives execution through ambiguity with confidence and clarity.

As a TPM at RunPod, you'll play a central role in leading complex, cross-functional programs across product and platform teams. This role sits inside our Supply organization — you'll be the connective tissue between Engineering, Product, Data Center Management, and our host partner teams, working to make sure GPU capacity keeps pace with demand. Through strong communication, alignment, and proactive risk management, you'll drive the successful delivery of high-impact programs that keep supply, cost, and reliability in balance for our customers.

Responsibilities:

  • Partner with Data Center Management and Host/Supply teams to run the capacity and utilization review cadence that feeds Supply's input into the product roadmap — surfacing constraints early (host churn, new capacity coming online, regional gaps) and turning them into actionable planning input.
  • Own the relationship and escalation path with host and infrastructure partners — resolving capacity, performance, or contractual issues quickly, and keeping Engineering and Product informed of anything that could affect delivery.
  • Translate capacity and utilization data into concrete tradeoffs: where to add supply, where to consolidate, and how each option nets out on cost and reliability.
  • Build and drive project plans across Engineering, Product, and Supply stakeholders — scoping work, sequencing dependencies, and setting timelines that hold up against real-world constraints like partner lead times and hardware delivery.
  • Act as the primary point of contact across these teams: spot risks and roadblocks early, and keep stakeholders aligned with clear, regular status updates.
  • Bring technical judgment to evaluations of infrastructure tools, platforms, and vendor products, recommending what best fits RunPod's scale and constraints.
  • Run post-program reviews on completed initiatives, capturing what worked, what didn't, and what should change for the next cycle.

Requirements:

  • Previous experience in infrastructure-focused organizations, with exposure to capacity planning, data center operations, or vendor/partner management.
  • 4+ years of experience or equivalent expertise in technical program management, leading complex technology programs from planning through delivery.
  • Proven track record of partnering with 3+ cross-functional teams to deliver high-impact programs from planning through launch.
  • Proven track record of working with external infrastructure or hardware vendors globally.
  • Strong technical background in engineering and scalable SaaS/IaaS products or services.
  • Strong quantitative and analytical skills, with the ability to turn capacity, cost, and delivery data into clear planning recommendations.
  • Deep understanding of the software development lifecycle (SDLC) and how engineering teams design, build, test, and ship products.
  • Highly organized and detail-oriented, with a track record of catching issues before they become blockers.
  • Demonstrated ability to independently build, track, and execute programs in dynamic environments with multiple dependencies, competing priorities, and tight deadlines.
  • A collaborative team player who values ownership, clarity, accountability, and follow-through over rigid process.
  • Exceptional verbal and written communication skills, with the ability to align stakeholders across teams and functions.
  • Solution-oriented and adaptable, comfortable navigating ambiguity and resolving challenges creatively under pressure.

Preferred:

  • Previous experience in AI/ML developer platforms, GPU/cloud infrastructure, or other hardware-capacity-constrained environments.
  • Experience at an early-stage or high-growth startup.
  • Technical degree in Computer Science, Engineering, or a related field.

What You’ll Receive:

  • The competitive base pay for this position ranges from ($140,000 - $165,000). This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location.
  • Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside.
  • Generous medical, dental & vision plans.
  • Flexible PTO- take the time you need to recharge.
  • Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication.
  • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.
  • $1,200 Home Office & Equipment Stipend- We set you up for success from day one with gear and support to create your ideal workspace.

How would you rate this job post?

See what other professionals think about this role.

banner

RunPod is a premier, enterprise-grade GPU cloud platform engineered to orchestrate massive-scale AI/ML compute ecosystems and intelligent, frictionless infrastructure-delivery workflows. Operating as a developer-first, high-throughput cloud hub, the company eliminates the operational friction of traditional, legacy-cloud providers—which frequently lock users into rigid, overpriced, and manual-heavy compute models—by seamlessly deploying advanced serverless GPU telemetry, rigorous multi-region container-orchestration architectures, and cohesive cross-platform scaling frameworks. Moving beyond rigid legacy VM-based paradigms, RunPod empowers over 500,000 global developers, researchers, and Fortune 500 enterprises to dynamically synchronize their training, fine-tuning, and inference pipelines with elite, autonomous, and cost-effective execution. Under the hood, their sophisticated proprietary data infrastructure natively manages complex multi-node cluster ingestion (A100/H100/H200 architectures), instantaneous autoscaling endpoint routing, and automated ephemeral-pod management, providing the necessary operational foundation to support everything from individual experimental models to large-scale, trillion-parameter distributed training. What sets RunPod apart is its uncompromising dedication to frictionless compute orchestration; by bridging the gap between highly technical, performance-intensive GPU-infrastructure demands and accessible, low-latency deployment interfaces, the platform empowers modern AI-native organizations to radically accelerate their production-AI velocity, eliminate prohibitive infrastructure-management bottlenecks, and build an unassailable foundation for continuous commercial and institutional dominance in the modern, AI-transformed digital landscape.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More
Technical Program Manager (TPM) at RunPod