AI Cluster Architect
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
Vultr is looking for an AI Cluster Architect who will be responsible for creating and refining large-scale GPU cluster architectures within strict power and infrastructure limits. This role focuses heavily on power-aware design: starting from a fixed power envelope, the architect determines the optimal number of GPUs while accounting for the full stack of services needing to be deployed—compute nodes, storage systems, networking fabric, cooling, and facility constraints. This role requires deep experience navigating heterogeneous environments, multiple generations of hardware, and end user requirements.
The architect must understand how different GPU SKUs, NICs, switches, and fabrics interact at scale, including their individual and aggregate power and thermal characteristics. They will evaluate multi-plane, rail-optimized, and tiered fabric designs across technologies like InfiniBand, RoCE, and SpectrumX to ensure the networking architecture supports the intended GPU count without overrunning facility limits or switch radix and/or topology constraints. This role balances customer-specific requirements for compute, storage, and service density, ensuring that the final cluster design maintains acceptable levels of GPU and fabric performance, while maximizing the number of usable GPUs within the total power budget.
Key Responsibilities
- Architect large-scale GPU clusters within fixed site power budgets that optimizes for maximum GPU density while reserving necessary headroom for compute services, storage, and networking.
- Model and validate power consumption across the full cluster bill of materials (GPUs, CPUs, NICs, switches, fabric components, storage, and facility limits).
- Evaluate tradeoffs across multiple fabric networking architectures (InfiniBand, RoCE, SpectrumX) as well as multi-plane, 2-tier/3-tier, and rail-optimized topologies.
- Determine network scale limits based on switch radix, link speed, topology, and blocking requirements.
- Gather, interpret, and maintain detailed SKU-level power and thermal specifications for GPUs, NICs, switches, DPUs, storage, and server platforms.
- Develop power-aware cluster configuration templates and capacity-planning models that can scale across sites with varying constraints and allow for quick iteration and ideation.
- Document architecture, design choices, tradeoff analyses, and operational considerations for deployment and lifecycle management.
- Provide guidance on future-proofing, including the ability to incorporate next-gen GPUs, NICs, or fabrics.
- Collaborate with vendors on novel fabric architectures that enable large-scale cluster deployments (100k+ GPUs).
Qualifications
- 7+ years designing or building large-scale HPC, AI, or hyperscale GPU clusters.
- Expert understanding of GPU and accelerator system design, including node topology, PCIe/NVLink/NVSwitch/ROCm, and NIC-to-GPU affinity considerations.
- Strong familiarity with InfiniBand, RoCE, and SpectrumX networking, including multi-tier, multi-plane, Clos/dragonfly variants, and large-radix switch design.
- Demonstrated experience modeling power draw and thermal characteristics of servers, GPUs, NICs, switches, optics, and storage systems.
- Ability to design networks that maintain full non-blocking performance or intentionally introduce over/under-subscription while understanding impacts on workload performance.
- Proven ability to gather and analyze vendor SKU-level specifications and incorporate them into scalable cluster architectures.
- Experience balancing customer-driven requirements for compute, storage, and service density in combination with overall GPU count.
- Strong documentation, communication, and cross-functional collaboration skills.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at Vultr
Explore Top Companies in this Space

Datadog
Cloud Observability
Rubrik
Cloud Data Management
Render
Cloud Computing

Twilio
Cloud Communications
Vultr
View Company ProfileVultr is a cloud computing platform that provides a range of infrastructure services, including virtual machines, storage, and networking. With a strong focus on scalability, reliability, and security, Vultr empowers businesses and individuals to build, deploy, and manage their applications and services with ease. The company's cloud platform is designed to support a wide range of use cases, from simple web applications to complex enterprise environments. By leveraging the latest advancements in cloud technology, Vultr delivers high-performance, low-latency, and cost-effective solutions that enable its customers to achieve their goals. With a user-friendly interface and a robust set of features, Vultr makes it easy for users to provision, manage, and optimize their cloud resources. Whether you're a developer, a startup, or an established enterprise, Vultr provides the flexibility, control, and support you need to succeed in today's fast-paced digital landscape. With its cutting-edge technology, exceptional customer support, and commitment to innovation, Vultr is an ideal choice for anyone looking to harness the power of the cloud.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.


