Vice President of Infrastructure & Deployment
Location & Travel
Remote or hybrid, with regular travel to data center markets, customer deployment sites, vendor locations, and company operating sessions expected.
About Acasia
Acasia builds, deploys, and operates high-performance GPU infrastructure for enterprise AI workloads. Our customers rely on Acasia to deliver production-grade GPU environments that are performant, reliable, scalable, and supportable in real-world data center conditions.
As Acasia expands across multiple data center markets, we are hiring a Vice President of Infrastructure & Deployment to serve as our senior infrastructure subject matter expert, technical authority, and escalation leader for GPU infrastructure design, deployment readiness, complex troubleshooting, ongoing support, and production reliability.
Role Summary
The Vice President of Infrastructure & Deployment owns the successful delivery, implementation, commissioning, and operational readiness of Acasia's AI infrastructure across customer and data center environments.
This executive owns the complete delivery lifecycle—from customer handoff after contract execution through deployment planning, installation, networking, cluster bring-up, validation, production acceptance, and ongoing optimization—ensuring every GPU cluster is delivered safely, on schedule, on budget, and to Acasia's standards before entering production.
The VP will build and lead Acasia's Infrastructure Delivery organization, managing field engineering teams that implement large-scale GPU clusters across multiple data centers and customer locations. They will establish deployment methodologies, technical standards, playbooks, commissioning procedures, and quality controls while driving continuous improvement in deployment speed, consistency, and customer experience. They will work hand-in-hand with customers, Sales, Customer Success, Engineering, Product, Operations, Procurement, OEM partners, networking vendors, and data center operators.
Mission: Build the industry's fastest, most reliable GPU infrastructure deployment organization—taking customer environments from signed contract to production in weeks instead of months.
Key Responsibilities
Infrastructure Architecture & Technical Leadership
- Own infrastructure architecture standards for GPU clusters across all Acasia deployments.
- Define reference architectures for rack layouts, power distribution, cooling, networking, storage, monitoring, telemetry, and remote management.
- Review customer infrastructure designs for scalability, resilience, serviceability, deployment efficiency, and lifecycle management.
- Identify technical and deployment risks before implementation begins.
- Establish repeatable infrastructure standards that simplify deployment while improving reliability and operational excellence.
- Partner with Engineering, Product, Security, Sales, Customer Success, Procurement, OEMs, and customers to ensure infrastructure designs meet both technical and commercial objectives.
Networking & Cluster Integration
- Own deployment and validation of high-performance AI networking environments.
- Lead implementation of InfiniBand, Ethernet, RoCE, RDMA, NVLink, NVSwitch, GPUDirect, and NVIDIA networking technologies.
- Oversee spine-leaf architecture deployment, east-west networking, switch configuration, storage networking, and cluster connectivity.
- Establish standards for network validation, benchmarking, congestion analysis, latency optimization, and fabric performance tuning.
- Validate cluster communication performance using NCCL, MPI, RDMA, and other AI infrastructure benchmarking tools.
- Partner with NVIDIA and networking OEMs to optimize production environments.
Production Readiness & Commissioning
- Define production acceptance standards for every customer deployment.
- Own Factory Acceptance Testing (FAT), Site Acceptance Testing (SAT), cluster commissioning, and customer handoff.
- Establish standards for hardware acceptance, BIOS and firmware consistency, GPU validation, driver installation, CUDA readiness, storage validation, telemetry, monitoring, and observability.
- Lead burn-in procedures, stress testing, soak testing, redundancy validation, performance baselining, and workload validation before production acceptance.
- Ensure every customer environment meets Acasia's technical standards before entering production.
Program Management & Deployment Execution
- Build and manage deployment schedules for large-scale infrastructure implementations.
- Coordinate execution across internal teams, customers, OEMs, contractors, and deployment partners.
- Track milestones, critical path activities, dependencies, and implementation risks.
- Lead executive deployment reviews and customer implementation meetings.
- Develop standardized governance, reporting, and deployment management processes.
- Drive continuous improvement in deployment efficiency, predictability, and execution quality.
Technical Escalation & Problem Resolution
- Serve as Acasia's highest-level technical escalation point for complex infrastructure issues.
- Lead cross-functional war-room responses during critical deployments or production-impacting incidents.
- Troubleshoot issues spanning GPU hardware, Linux, networking, firmware, drivers, storage, environmental systems, and customer workloads.
- Partner with Engineering, OEMs, vendors, data center providers, and customers on root cause analysis and permanent corrective actions.
- Convert recurring issues into improved standards, tooling, automation, documentation, and deployment methodologies.
Leadership & Organizational Development
- Build and lead Acasia's Infrastructure Delivery and Field Engineering organization.
- Recruit, mentor, and develop world-class infrastructure engineers and deployment leaders.
- Establish technical certification and training programs across the organization.
- Build repeatable knowledge transfer mechanisms, technical playbooks, deployment standards, and troubleshooting frameworks.
- Foster a culture of accountability, precision, technical excellence, customer obsession, and continuous improvement.
Vendor & Strategic Partner Management
- Build executive technical relationships with NVIDIA, OEM partners, networking providers, storage vendors, and data center operators.
- Evaluate new infrastructure technologies, deployment methodologies, and vendor capabilities.
- Participate in technical due diligence for new data center markets and deployment strategies.
- Hold strategic partners accountable during implementation and complex technical escalations.
- Support commercial and procurement teams with technical guidance on infrastructure quality, lifecycle planning, supportability, and deployment risk.
Required Qualifications
- Bachelor's degree in Computer Science, Electrical Engineering, Information Technology, or a related technical discipline (Master's preferred).
- 10+ years leading infrastructure engineering, field engineering, or large-scale infrastructure deployment organizations.
- Demonstrated success deploying production AI or HPC infrastructure environments exceeding 500 GPUs; experience with deployments of 1,000+ GPUs preferred.
- Extensive hands-on experience implementing GPU clusters, Linux environments, storage systems, and high-performance networking.
- Deep expertise with rack-scale infrastructure, structured cabling, power distribution, cooling systems, and data center operations.
- Strong Linux systems administration and troubleshooting expertise.
- Strong networking expertise including InfiniBand, RoCE, Ethernet, switching, routing, VLANs, RDMA, DNS, DHCP, firewalls, and production network troubleshooting.
- Experience leading Factory Acceptance Testing (FAT), Site Acceptance Testing (SAT), commissioning, burn-in, and production validation.
- Experience serving as the senior technical escalation point for complex production infrastructure issues.
- Strong project leadership, vendor management, and cross-functional execution skills.
- Excellent communication and executive presentation skills.
- Willingness to travel extensively to customer sites, deployment locations, OEM facilities, and data center markets.
Preferred Qualifications
- NVIDIA Certified Professional (NCP) strongly preferred, ideally NVIDIA Certified Professional – AI Infrastructure (NCP-AII) or equivalent NVIDIA AI Infrastructure certification.
- Experience deploying NVIDIA HGX, DGX, GB300, GB200, B300, B200, H200, H100, or equivalent AI infrastructure platforms.
- Expert knowledge of NVIDIA Networking, NVLink, NVSwitch, CUDA, NCCL, GPUDirect Storage, BlueField DPUs, DCGM, NVML, Redfish, BMC/IPMI, iDRAC, and iLO.
- Experience working directly with NVIDIA and leading OEM partners including Supermicro, Dell, HPE, Lenovo, ASUS, and GIGABYTE.
- Experience with Kubernetes, Slurm, Docker, bare-metal provisioning, AI workload orchestration, and GPU resource management.
- Experience in cloud infrastructure, hyperscale data centers, AI infrastructure companies, managed infrastructure providers, or HPC environments.
- Professional certifications such as CCNP, CCIE, RHCE, RHCSA, or equivalent enterprise infrastructure credentials.
- Own technical architecture standards for Acasia’s GPU infrastructure environments across data center markets.
- Design and review infrastructure patterns for GPU servers, racks, power, cooling, networking, cabling, storage, monitoring, telemetry, and remote management.
- Establish reference architectures for repeatable GPU cluster deployments.
- Review proposed customer deployments for performance, reliability, maintainability, scalability, and supportability.
- Identify technical risks in infrastructure designs before they become production issues.
- Partner with engineering, product, security, customer success, sales, vendors, and data center providers to ensure infrastructure designs meet customer and business requirements.
- Serve as Acasia’s internal authority on GPU infrastructure architecture, data center deployment patterns, and production readiness.