Senior Manager, Cluster Engineering & Deployment
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
TensorWave’s mission is to deliver seamless, secure, reliable, and resilient AI compute at scale. Their cloud platform eliminates infrastructure barriers, allowing builders to focus on innovation rather than technical challenges.
About the Role
The Senior Manager, Cluster Engineering & Deployment oversees the entire process of transforming delivered racks into fully operational clusters. Key responsibilities include:
Network bring-up, fabric cabling verification against port maps, GPU node integration with the fabric, cluster-level validation, and burn-in (including RCCL/collective performance testing).
Acts as the critical gatekeeper for cluster acceptance into production, directly impacting cluster revenue timelines.
What You’ll Do
Own and evolve the cluster deployment playbook, including staged bring-up, automated configuration push, link/optics validation, cabling verification, and fault triage during deployment windows.
Lead deployment engineering across concurrent cluster builds, coordinating with team leads, on-site engineers, Data Center Integration teams, and cabling vendors.
Drive deployment velocity by optimizing bring-up time per cluster through tooling, pre-staging, and defect-source elimination, setting measurable targets for the team.
Establish and enforce defect feedback loops with Network Engineering, Layer One (cabling quality), and vendors (hardware/optics RMA patterns), ensuring accountability for resolution.
Define and standardize spares, test equipment, and deployment tooling requirements across all sites.
Who You Are
Required Qualifications
At least 10+ years of experience in network deployment, cluster/HPC bring-up, or large-scale infrastructure delivery, including managing engineers in a field/deployment setting.
Hands-on experience with fabric bring-up at scale (hundreds of switches / thousands of links per deployment).
Strong operational rigor, including building and enforcing playbooks, gates, metrics, and blameless defect loops.
Proven team leadership with schedule accountability across multiple concurrent builds or sites.
Preferred Qualifications
GPU cluster validation experience, including NCCL/RCCL benchmarking.
Automation skills using Python and Ansible for deployment tasks.
Depth in optics/link-layer debugging.
Experience with acceptance testing as a commercial gate (revenue-linked).
End-to-end ownership of cluster validation, including bandwidth/latency baselines, collective (RCCL) performance tests, burn-in criteria, and go/no-go acceptance gates.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at TensorWave
Senior Technical Program Manager, Data Center Build and Design
TensorWave
United StatesSenior Manager, Technical Program Management, Data Center Build and Design
TensorWave
United StatesSenior Network Engineer - Operations
TensorWave
United StatesData Center Mechanical Engineer
TensorWave
United StatesExplore Top Companies in this Space
Lambda
AI Infrastructure & GPU Cloud / Supercomputing & Data Centers / Machine Learning Platforms / High-Performance Computing (HPC)
TSMG
AI Infrastructure / Cloud Computing / Data Science / HPC
Construction.com
Construction / Digital Marketplaces / B2B Services / Enterprise Software
ReversingLabs
Cybersecurity / Artificial Intelligence / Enterprise Software / Threat Intelligence
TensorWave
View Company ProfileTensorWave is a premier, enterprise-grade cloud infrastructure provider engineered to orchestrate massive-scale artificial intelligence (AI) and high-performance computing (HPC) workloads. Operating as a high-velocity digital ecosystem, the company eliminates the operational friction and supply chain bottlenecks of legacy hyperscalers by exclusively deploying advanced AMD Instinct™ accelerators—including the MI300X and MI325X—across highly optimized bare-metal environments. Moving beyond the rigid memory constraints of traditional GPU cloud models, TensorWave provides industry-leading capacity with up to 288GB of HBM3e per accelerator, directly tackling the immense requirements of next-generation Large Language Models (LLMs) and complex machine learning training clusters. Under the hood, their UEC-ready networking architecture and direct liquid cooling systems seamlessly integrate into existing AI pipelines, ensuring ultra-low latency inference and blistering training speeds. What sets TensorWave apart is its uncompromising dedication to open-source democratization and performance accessibility; by bridging the gap between top-tier compute resources and massive cost-efficiency, the platform empowers scaling enterprises to radically accelerate AI innovation without the burden of building internal hardware infrastructure.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.