Director of Data Center Facilities - AI Infrastructure
Job Description
Key Skills Required
Master these to land this role
Want to know if you're a match for this job?
We’re looking for a Director of Data Center Facilities - AI Infrastructure to join our team during an exciting phase of growth. In this role, you will own the physical infrastructure strategy, operations, reliability, and expansion of a high-density AI data center built to support GPU computing at scale.
You'll ensure the facility can deliver the power, cooling, connectivity, and operational capacity that fast-growing AI compute platforms demand, while leading multidisciplinary facilities teams and partnering closely with Operations, Engineering, Network, Construction, Security, Finance, Procurement, utilities, and technology vendors.
What You’ll Do
AI Data Center Operations
- Provide overall leadership for facilities operations supporting high-density AI/GPU compute environments.
- Ensure the facility consistently delivers the required power, cooling, environmental conditions, and infrastructure availability for AI compute clusters.
- Establish operational readiness standards for new AI compute deployments, including infrastructure validation, capacity verification, and turnover.
- Lead facility response to electrical, mechanical, thermal, controls, and cooling-related incidents.
High-Density Power Infrastructure
- Oversee utility service, substations, medium-voltage distribution, transformers, switchgear, UPS systems, generators, busways, PDUs, and rack-level power distribution.
- Partner with utilities and engineering teams on utility capacity, grid interconnection, power availability, resiliency, and future expansion.
- Evaluate power architectures and redundancy strategies appropriate for AI infrastructure.
- Develop long-range power capacity plans aligned with GPU/accelerator deployment schedules.
Advanced Cooling & Thermal Management
- Lead the design, operation, and optimization of cooling systems supporting high-density AI compute.
- Establish standards for liquid-cooling reliability, leak detection, water quality, filtration, redundancy, and maintenance.
- Partner with AI hardware and compute teams to understand evolving thermal requirements for GPUs, accelerators, CPUs, networking equipment, and future-generation platforms.
- Evaluate cooling capacity at the rack, row, room, and facility level.
- Develop strategies for increasing rack density without compromising thermal performance or reliability.
AI Capacity Planning & Infrastructure Strategy
- Develop multi-year facilities capacity plans based on AI compute roadmaps and anticipated GPU/accelerator deployments.
- Translate compute requirements into MW, cooling tonnage, rack density, water, space, and infrastructure requirements.
- Partner with AI infrastructure and data center engineering teams to establish deployment timelines and facility readiness requirements.
- Develop scalable architectures capable of supporting rapid expansion from individual clusters to multi-megawatt AI campuses.
- Evaluate emerging infrastructure technologies and determine their suitability for large-scale AI environments.
- Participate in site selection, utility strategy, infrastructure due diligence, and campus master planning.
New Construction, Expansion & Commissioning
- Lead facilities participation in the design, construction, commissioning, and turnover of new AI data center capacity.
- Establish commissioning requirements for electrical, mechanical, controls, liquid-cooling, and life-safety systems.
- Review engineering designs, equipment selections, redundancy models, sequence-of-operations documents, and commissioning plans.
- Develop standardized infrastructure designs that can be replicated across multiple AI data center sites.
Reliability & Incident Management
- Establish a world-class reliability program for AI data center infrastructure.
- Lead root-cause analysis for critical facility incidents and implement sustainable corrective actions.
- Develop predictive and condition-based maintenance strategies for critical assets.
- Establish reliability metrics for power, cooling, controls, and liquid-cooling systems.
- Conduct failure-mode analysis and identify potential infrastructure risks before they become operational events.
- Lead emergency response and recovery for facility incidents affecting AI compute availability.
- Conduct regular scenario-based exercises involving major power failures, cooling failures, loss of utility service, controls failures, and liquid-cooling events.
Automation, Controls & Data
- Drive increased use of BMS, EPMS, DCIM, telemetry, analytics, and predictive monitoring across the facility.
- Establish real-time visibility into electrical and thermal capacity.
- Identify opportunities to automate facility operations while maintaining appropriate safeguards and human oversight.
Team Leadership
- Build and lead a high-performing organization of facilities engineers, managers, technicians, and specialized contractors.
- Establish organizational structures capable of supporting 24x7 AI data center operations at scale.
- Develop technical training programs focused on high-density power, liquid cooling, controls, and AI infrastructure.
- Establish succession planning and technical career-development programs.
- Foster a culture of operational excellence, safety, urgency, accountability, and continuous improvement.
Financial & Vendor Management
- Manage strategic relationships with utilities, OEMs, engineering firms, contractors, cooling providers, and equipment manufacturers.
- Establish performance requirements and service-level agreements for critical vendors.
- Evaluate lifecycle costs, reliability, availability, maintainability, and scalability when selecting infrastructure technologies.
- Identify opportunities to reduce operating costs while maintaining or improving infrastructure reliability.
Safety, Compliance & Risk
- Establish a safety-first culture appropriate for high-voltage electrical systems, heavy mechanical equipment, industrial cooling systems, and construction environments.
- Develop comprehensive emergency response and business-continuity programs.
- Maintain accurate facility documentation, drawings, asset records, operating procedures, and maintenance histories.
- Conduct regular infrastructure risk assessments and resilience reviews.
- Ensure facilities are prepared for internal, customer, regulatory, insurance, and third-party audits.
Energy & Sustainability
- Develop strategies to manage the significant energy demands associated with AI compute.
- Optimize PUE, WUE, power utilization, cooling efficiency, and carbon impact.
- Partner with utilities and energy teams on renewable energy, energy procurement, demand management, and grid-related initiatives.
- Evaluate alternative cooling technologies, heat-reuse opportunities, water-conservation strategies, and other sustainability initiatives.
- Balance sustainability objectives with AI compute availability, reliability, and growth requirements.
How would you rate this job post?
See what other professionals think about this role.
Similar Opportunities
More Openings at TensorWave
Senior Manager, Cluster Engineering & Deployment
TensorWave
United StatesSenior Technical Program Manager, Data Center Build and Design
TensorWave
United StatesSenior Manager, Technical Program Management, Data Center Build and Design
TensorWave
United StatesSenior Network Engineer - Operations
TensorWave
United StatesExplore Top Companies in this Space
Lambda
AI Infrastructure & GPU Cloud / Supercomputing & Data Centers / Machine Learning Platforms / High-Performance Computing (HPC)
TSMG
AI Infrastructure / Cloud Computing / Data Science / HPC
MySigrid
Remote Staffing / Virtual Assistance / AI-Powered Solutions / Business Services
Label Your Data
Information Technology & Services / Artificial Intelligence / Data Annotation / Enterprise Software
TensorWave
View Company ProfileTensorWave is a premier, enterprise-grade cloud infrastructure provider engineered to orchestrate massive-scale artificial intelligence (AI) and high-performance computing (HPC) workloads. Operating as a high-velocity digital ecosystem, the company eliminates the operational friction and supply chain bottlenecks of legacy hyperscalers by exclusively deploying advanced AMD Instinct™ accelerators—including the MI300X and MI325X—across highly optimized bare-metal environments. Moving beyond the rigid memory constraints of traditional GPU cloud models, TensorWave provides industry-leading capacity with up to 288GB of HBM3e per accelerator, directly tackling the immense requirements of next-generation Large Language Models (LLMs) and complex machine learning training clusters. Under the hood, their UEC-ready networking architecture and direct liquid cooling systems seamlessly integrate into existing AI pipelines, ensuring ultra-low latency inference and blistering training speeds. What sets TensorWave apart is its uncompromising dedication to open-source democratization and performance accessibility; by bridging the gap between top-tier compute resources and massive cost-efficiency, the platform empowers scaling enterprises to radically accelerate AI innovation without the burden of building internal hardware infrastructure.
Safety First
- Never pay for a job application.
- Do not share sensitive bank info.
- Verify the client before starting work.