Back to Jobs
TensorWave
Engineering & Architecture 14h ago

Senior Manager, Cluster Engineering & Deployment

TensorWave
United StatesUnited States
Full-time
Not Disclosed
Senior-Level

Job Description

Key Skills Required

Master these to land this role

Python2h 41mFree Trial ✨
Start 10-Day Free Trial
Network EngineeringAutomation EngineerHPCCluster Deployment

Want to know if you're a match for this job?

Calculate My Match Score

TensorWave’s mission is to deliver seamless, secure, reliable, and resilient AI compute at scale. Their cloud platform eliminates infrastructure barriers, allowing builders to focus on innovation rather than technical challenges.

About the Role

The Senior Manager, Cluster Engineering & Deployment oversees the entire process of transforming delivered racks into fully operational clusters. Key responsibilities include:

  • Network bring-up, fabric cabling verification against port maps, GPU node integration with the fabric, cluster-level validation, and burn-in (including RCCL/collective performance testing).

  • Acts as the critical gatekeeper for cluster acceptance into production, directly impacting cluster revenue timelines.

What You’ll Do

  • Own and evolve the cluster deployment playbook, including staged bring-up, automated configuration push, link/optics validation, cabling verification, and fault triage during deployment windows.

  • Lead deployment engineering across concurrent cluster builds, coordinating with team leads, on-site engineers, Data Center Integration teams, and cabling vendors.

  • Drive deployment velocity by optimizing bring-up time per cluster through tooling, pre-staging, and defect-source elimination, setting measurable targets for the team.

  • Establish and enforce defect feedback loops with Network Engineering, Layer One (cabling quality), and vendors (hardware/optics RMA patterns), ensuring accountability for resolution.

  • Define and standardize spares, test equipment, and deployment tooling requirements across all sites.

Who You Are

Required Qualifications

  • At least 10+ years of experience in network deployment, cluster/HPC bring-up, or large-scale infrastructure delivery, including managing engineers in a field/deployment setting.

  • Hands-on experience with fabric bring-up at scale (hundreds of switches / thousands of links per deployment).

  • Strong operational rigor, including building and enforcing playbooks, gates, metrics, and blameless defect loops.

  • Proven team leadership with schedule accountability across multiple concurrent builds or sites.

Preferred Qualifications

  • GPU cluster validation experience, including NCCL/RCCL benchmarking.

  • Automation skills using Python and Ansible for deployment tasks.

  • Depth in optics/link-layer debugging.

  • Experience with acceptance testing as a commercial gate (revenue-linked).

  • End-to-end ownership of cluster validation, including bandwidth/latency baselines, collective (RCCL) performance tests, burn-in criteria, and go/no-go acceptance gates.

How would you rate this job post?

See what other professionals think about this role.

banner

TensorWave is a premier, enterprise-grade cloud infrastructure provider engineered to orchestrate massive-scale artificial intelligence (AI) and high-performance computing (HPC) workloads. Operating as a high-velocity digital ecosystem, the company eliminates the operational friction and supply chain bottlenecks of legacy hyperscalers by exclusively deploying advanced AMD Instinct™ accelerators—including the MI300X and MI325X—across highly optimized bare-metal environments. Moving beyond the rigid memory constraints of traditional GPU cloud models, TensorWave provides industry-leading capacity with up to 288GB of HBM3e per accelerator, directly tackling the immense requirements of next-generation Large Language Models (LLMs) and complex machine learning training clusters. Under the hood, their UEC-ready networking architecture and direct liquid cooling systems seamlessly integrate into existing AI pipelines, ensuring ultra-low latency inference and blistering training speeds. What sets TensorWave apart is its uncompromising dedication to open-source democratization and performance accessibility; by bridging the gap between top-tier compute resources and massive cost-efficiency, the platform empowers scaling enterprises to radically accelerate AI innovation without the burden of building internal hardware infrastructure.

Safety First

  • Never pay for a job application.
  • Do not share sensitive bank info.
  • Verify the client before starting work.
Learn More