Senior HPC Support Engineer
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.
If you'd like to build the world's best AI cloud, join us.
This position is expected to participate in an on-call rotation.
What You’ll Do
- Serve as a senior technical escalation point, troubleshooting the hardest infrastructure and platform issues down to the hardware, driver, or kernel level when needed
- Quickly and accurately distinguish between hardware failures, driver issues, kernel-level problems, and customer workload misconfiguration, so issues get resolved correctly the first time
- Proactively identify process, tooling, and documentation gaps, and go fix them, not just wait for them to be assigned
- Use AI tools effectively to build scripts, automations, or small internal tools that close real operational gaps (no professional development background required)
- Perform root-cause analysis across distributed systems, clusters, and GPU infrastructure
- Craft clear documentation of solutions and contribute to evolving support procedures
- Collaborate closely with engineering teams to turn recurring customer pain points into permanent fixes
- Take escalations from peers while training and mentoring them in the process
- Participate in a rotating on-call schedule, owning major incidents and major customer issues
- Be ready to roll up your sleeves and pitch in wherever needed, especially during fast, high-volume deployments
You
- 3+ years of hands-on HPC experience in an administration, support, or engineering role.
- Very strong understanding and experience supporting Linux in a system administration role.
- Proven experience in HPC environments, showcasing your expertise in Linux cluster administration, with strong preference for Kubernetes and/or Slurm for cluster orchestration.
- Strong coding ability and CI/CD experience, with a track record of using AI-assisted tools to move fast.
- Proficiency with monitoring/logging tools (Prometheus, Grafana, Datadog).
- Strong skills in log analysis, debugging kernel-level issues, and performance profiling.
- Experience with CUDA, NCCL, NVLink, GPUDirect RDMA.
- Experience with high throughput networking technologies(IB/RoCE).
- Knowledge of distributed AI/ML or HPC workloads.
- Knowledge of TCP/IP, VPN, and firewalls in cloud environments.
- Ability to work independently and mentor junior support engineers.
Nice to Have
- Experience with virtualization and container (Docker, Kubernetes) technologies.
- Experience with neoclouds/GPU cloud providers.
- Flexible availability for potential shifts outside of normal working hours/weekends.
- Experience with high performance storage systems.
- Familiarity with infrastructure-as-code tools (Terraform, Ansible, etc.)
- Experience with Nvidia GPUs and Infiniband.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Lambda
Explore Top Companies in this Space

Twilio
Cloud Communications
Pila8
Data Annotation & Labeling
Rubrik
Cloud Data Management

Datadog
Cloud Observability
Lambda
View Company ProfileLambda (operating under lambda.ai, formerly Lambda Labs) is the premier, enterprise-grade AI computing platform, modular AI factory pioneer, and GPU infrastructure powerhouse engineered to operate as the definitive, high-velocity compute backbone for the world’s most advanced AI research labs, engineering teams, and superintelligence initiatives. Founded in 2012 by ML engineers, the company completely eliminates the severe systemic friction of modern AI training—where high-performance model development is often throttled by limited GPU accessibility, complex cluster orchestration, and inefficient interconnect bottlenecks—by deploying an advanced, purpose-built AI factory matrix. Moving far beyond traditional, general-purpose cloud providers, Lambda natively unifies high-density, liquid-cooled, and InfiniBand-interconnected GPU compute infrastructure with a curated, AI-optimized software stack (Lambda Stack). The ecosystem offers a tiered deployment model: on-demand GPU instances for rapid prototyping, \"1-Click Clusters\" for production-ready model training (16 to 2,000+ NVIDIA B200 or H100 GPUs), and bespoke \"Superclusters\" (4,000 to 165,000+ GPUs) designed for foundation model training at gigawatt scale. Trusted by 97% of top U.S. research universities and a massive global footprint of enterprise AI developers, Lambda’s infrastructure is explicitly tuned for PyTorch, TensorFlow, and JAX workloads. Backed by $480M+ in Series D funding and recognized as a key NVIDIA partner for healthcare and supercomputing, the organization maintains a strict \"AI-only\" engineering focus, providing 24/7 technical support specifically centered on resolving complex ML cluster performance, networking, and throughput issues. What sets Lambda apart is its uncompromising dedication to replacing complex, fragmented infrastructure management with absolute compute accessibility and performance predictability; by bridging the gap between mission-critical hardware engineering and the rapid-fire requirements of AI reasoning and generation, the enterprise remains the definitive, category-defining cornerstone of modern algorithmic supercomputing and worldwide AI transformation.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.

