TECHNICAL MANAGER-- GPU CLOUD & AI INFRASTRUCTURE
Zy Future International · Singapore
TECHNICAL MANAGER - GPU CLOUD & AI INFRASTRUCTURELocation: SingaporeEmployment Type: Full-Time, PermanentMonthly Salary: S$15,000-S$20,000Reporting To: Head of GPU Cloud and AI InfrastructureTravel Requirement: Regional travel within Southeast AsiaROLE OVERVIEWWe are seeking an experienced Technical Manager to lead the architecture, development, deployment and technical operation of our GPU Cloud and AI infrastructure platform.Based in Singapore, the successful candidate will serve as the technical lead for GPU infrastructure architecture, AI computing clusters, cloud-platform development, customer solution design, project delivery and platform operations.The role will connect GPU servers, data center infrastructure, high-speed networking, parallel storage, cloud-native platforms and AI model workloads to create a stable, scalable, schedulable, measurable and commercially viable GPU Cloud service.This is a technical leadership position. Candidates must possess hands-on experience in GPU clusters, AI infrastructure, high-performance computing or large-scale parallel-computing platforms.KEY RESPONSIBILITIES1. GPU Infrastructure Architecture• Lead the overall technical architecture and solution design of GPU servers and AI computing clusters.• Evaluate high-performance GPU platforms, server configurations, CPUs, GPUs, memory, high-speed local storage, network interfaces and rack configurations.• Design or review InfiniBand, RoCE and high-speed Ethernet network architectures.• Evaluate NV Link, NV Switch, RDMA and multi-node, multi-GPU communication solutions.• Work with data center teams to confirm rack power density, liquid or air-cooling requirements, PDU allocation, network connectivity and cabling specifications.• Establish technical standards for GPU cluster deployment, performance testing, production acceptance and capacity expansion.2. GPU Cloud Platform Development• Lead the technical planning, implementation and operation of GPU resource pools, computing clusters and GPU Cloud platforms.• Build and maintain GPU resource scheduling and management platforms using Kubernetes and/or Slurm.• Manage GPU drivers, acceleration libraries, container runtimes, NVIDIA GPU Operator and related infrastructure software stacks.• Design multiple GPU service models, including bare metal, virtual machines, containers and GPU instances.• Implement full-GPU allocation, MIG partitioning, GPU sharing, quota management and multi-tenant resource isolation.• Support user management, access control, resource provisioning, order activation, resource recovery and API integration.• Work with product and commercial teams to develop standardized and scalable GPU Cloud services.3. AI Workload Support• Support customers with large language model training, fine-tuning, inference deployment and AI workload optimization.• Analyze customer requirements relating to model size, dataset size, concurrency, throughput and latency.• Recommend appropriate GPU models, node quantities, network topologies and storage configurations.• Support distributed training environments using PyTorch, TensorFlow an dother mainstream AI frameworks.• Deploy and optimise inference solutions using vLLM, NVIDIA Triton, TensorRT-LLM or similar AI inference frameworks.• Lead customer proof-of-concept projects, performance benchmarks, technical testing and acceptance processes.4. Platform Operations and SLA Management• Establish monitoring and alerting systems covering GPUs, computing nodes, networks and storage using technologies such as DCGM, Prometheus and Grafana.• Monitor GPU utilization, GPU memory usage, power consumption, temperature errors and idle capacity.• Develop accurate usage-metering logic based on GPU hours, instance hours or Token consumption.• Establish service-level agreements, incident-classification standards, response procedures, escalation mechanisms and disaster-recovery plans.• Continuously improve platform availability, GPU utilisation and commercial output per GPU.5. Customer Solutions and Pre-Sales Support• Work with sales and commercial teams to understand customers’ AI and infrastructure requirements.• Translate customer requirements into solution architectures, technical proposals, RFP or RFQ responses, bills of materials and implementation plans.• Lead technical presentations, architecture discussions, proof-of-concept validation, customer onboarding and capacity-expansion activities.• Define service boundaries, delivery standards, SLA commitments and technical acceptance criteria.6. Project and Vendor Management• Manage server OEMs, data center providers, network and storage vendors, and system integrators.• Oversee equipment delivery, rack installation, cabling, environment initialization, cluster commissioning and final acceptance.• Manage project schedules, technical risks, issue lists and remediation activities.• Establish technical documentation, operational procedures, deployment standards and maintenance manuals.• Recruit and manage infrastructure and operations engineers as the business expands.REQUIREMENTS• Bachelor’s degree in Computer Science, Electronic Engineering, Telecommunications, Software Engineering or a related discipline.• At least 7 years of relevant experience in cloud computing, high-performance computing, AI infrastructure or data center environments.• At least 3 years of hands-on experience building or operating GPU clusters, AI Cloud platforms or large-scale parallel-computing environments.• Strong practical knowledge of Linux, Docker, Kubernetes, Helm and cloud-native technologies.• Strong understanding of GPU hardware architecture, GPU drivers, acceleration libraries and AI server software stacks.• Hands-on experience with Kubernetes GPU scheduling and/or Slur cluster-resource management.• Good understanding of InfiniBand, RoCE, RDMA, NV Link, NV Switch and high-speed networking technologies.• Practical experience supporting large language model training, fine-tuning or inference workloads.• Experience with platform monitoring, incident management, capacity planning and SLA management.• Strong customer communication, cross-functional collaboration, technical documentation and vendor-management capabilities.• Excellent written and spoken English, as English will be used for regional operations, technical documentation, customer communication and vendor management.• Professional proficiency in Mandarin is required for regular technical coordination with China-based engineering teams, technology suppliers and Mandarin-speaking customers.• Willing and able to travel within Southeast Asia when required.PREFERRED QUALIFICATIONS• Experience building or managing NVIDIA HGX-based or other enterprise-scale AI computing clusters.• Experience with Slurm, HPC or large-scale distributed training platforms.• Experience implementing GPU virtualization, MIG partitioning, GPU sharing or usage-metering systems.• Hands-on deployment experience with vLLM, NVIDIA Triton, TensorRT-LLM or similar AI inference frameworks.• Experience building GPU-hour, instance-hour or Token-based metering and billing systems.• Previous experience with a hyperscale, AI Cloud provider, GPU computing service provider, data center operator or major server OEM.• Relevant professional certifications such as CKA, CKAD, Linux, cloud computing or networking certifications.KEY PERFORMANCE INDICATORS• On-time and successful launch of GPU clusters and GPU Cloud platforms.• Platform availability and customer SLA achievement.• Average GPU utilization and idle-capacity management.• Customer PoC conversion and technical acceptance success rate.• Mean Time to Recovery for critical incidents.• Improvement in billable GPU usage and commercial output per GPU.• Quality and completeness of technical documentation, operating procedures and delivery standards.