Back to Jobs
RunPod
AI & Machine Learning 1h ago

Forward Deployed Engineer (AI Infrastructure) at Runpod

RunPod
United StatesUnited States
Full-time
$100,000 - $160,000 USD
Mid-Level

Job Description

Key Skills Required

Master these to land this role

Machine Learning41mFree Trial ✨
Start 10-Day Free Trial
Python2h 41mFree Trial ✨
Start 10-Day Free Trial
AI EngineerFull-stack DevelopmentCloud Platform

Want to know if you're a match for this job?

Calculate My Match Score

Runpod is the AI Developer Cloud, serving over one million developers—from indie researchers to teams running frontier models in production. The platform has processed more than 20 billion inference requests and recently closed a $100M Series A in June 2026. Runpod is at a pivotal moment in AI infrastructure, building the platform the next generation of developers will rely on.

The team is small, remote-first, and values ownership, speed, and impact. They ship work that over a million developers depend on daily. They seek individuals who care deeply, build with urgency, and want to make a difference at scale.

As part of the Revenue Team, this role involves close collaboration with customers, sales, product, engineering, and support teams to ensure a seamless and delightful Runpod experience. The focus is on providing technical insights, resolving challenges, and building trust with customers. The team aims to deliver exceptional solutions and foster cross-functional collaboration to optimize platform performance.

This Forward Deployed Engineer will support the growth of Runpod’s cloud platform. The role ensures customers experience seamless operations by addressing technical challenges, guiding onboarding, and collaborating with engineering and sales teams. Ideal candidates combine technical expertise with strong communication skills to enhance customer satisfaction and improve the platform.

Key responsibilities include solving technical challenges, shaping the customer experience, and driving product improvements. By building strong customer relationships and collaborating across teams, contributions will directly enhance operational efficiency, customer satisfaction, and Runpod’s success.

Responsibilities:

  • Participate in sales meetings with customers, explain Runpod’s specific technologies, provide architectural recommendations, and build proof-of-concept solutions to support onboarding for potential high-spending customers.

  • Troubleshoot and resolve critical or complex technical issues escalated by customers, including those related to configuration, performance, functionality, compatibility, or code errors in Runpod’s products or services.

  • Utilize various tools and methods, such as code analysis, scripting for testing, log analysis, and remote access, to identify root causes and deliver solutions or workarounds.

  • Communicate effectively with customers and internal teams, including engineering, sales, supply, and product management teams, to ensure customer satisfaction and integrate valuable feedback.

  • Assist the support team in troubleshooting and resolving escalated technical tickets, collaborate with the engineering team to provide workarounds or bug fixes, and work with the infrastructure team to address GPU server-related issues.

  • Contribute to product development and testing efforts by relaying feedback from customers and the support team to the engineering team, helping to shape product improvements.

  • Create and maintain technical documentation, such as knowledge base articles, FAQs, guides, and manuals, while also developing and delivering training sessions, webinars, and demos for customers, partners, and internal teams.

Requirements:

  • A Bachelor's degree in a relevant field (e.g., Computer Science, Computer Engineering, Software Engineering, Information Technology, or a related field) or equivalent professional experience.

  • 3+ years of professional experience in software development.

  • Strong problem-solving skills and ability to work in a collaborative environment.

  • Familiarity with applied AI use cases such as inference, fine-tuning, LLM-based applications, or agentic systems.

  • Excellent communication skills and attention to detail.

  • Located in the APAC region (Ideally in Malaysia, Singapore, or South Korea).

  • Successful completion of a background check.

Preferred:

  • Strong technical skills, including proficiency in Python (Django, Flask, PyTorch), JavaScript (React, Node.js), and Go, with hands-on experience in full-stack development.

  • Knowledge of various operating systems (Linux/Ubuntu), containerization (Docker), SQL databases (MySQL, PostgreSQL), NoSQL databases (MongoDB), and an understanding of data center components such as server hardware and network operations, as well as familiarity with common network protocols like TCP/IP, HTTP/HTTPS, SSH, and DNS.

  • Proficiency in deploying and managing AI/ML models as services or APIs, along with a deep understanding of machine learning models, including their training, inference, and deployment processes on Runpod platforms.

  • Ability to write and execute scripts, queries, or commands, and the ability to read and understand code, logs, errors, or traces.

How would you rate this job post?

See what other professionals think about this role.

banner

RunPod is a premier, enterprise-grade GPU cloud platform engineered to orchestrate massive-scale AI/ML compute ecosystems and intelligent, frictionless infrastructure-delivery workflows. Operating as a developer-first, high-throughput cloud hub, the company eliminates the operational friction of traditional, legacy-cloud providers—which frequently lock users into rigid, overpriced, and manual-heavy compute models—by seamlessly deploying advanced serverless GPU telemetry, rigorous multi-region container-orchestration architectures, and cohesive cross-platform scaling frameworks. Moving beyond rigid legacy VM-based paradigms, RunPod empowers over 500,000 global developers, researchers, and Fortune 500 enterprises to dynamically synchronize their training, fine-tuning, and inference pipelines with elite, autonomous, and cost-effective execution. Under the hood, their sophisticated proprietary data infrastructure natively manages complex multi-node cluster ingestion (A100/H100/H200 architectures), instantaneous autoscaling endpoint routing, and automated ephemeral-pod management, providing the necessary operational foundation to support everything from individual experimental models to large-scale, trillion-parameter distributed training. What sets RunPod apart is its uncompromising dedication to frictionless compute orchestration; by bridging the gap between highly technical, performance-intensive GPU-infrastructure demands and accessible, low-latency deployment interfaces, the platform empowers modern AI-native organizations to radically accelerate their production-AI velocity, eliminate prohibitive infrastructure-management bottlenecks, and build an unassailable foundation for continuous commercial and institutional dominance in the modern, AI-transformed digital landscape.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More
Forward Deployed Engineer (AI Infrastructure) at Runpod at RunPod | HireSkys