GPU as a Service in Malaysia
AI Training, Fine-Tuning and Inference
Run demanding AI workloads on enterprise-grade NVIDIA GPUs without purchasing and maintaining your own hardware.
Deploy your workloads in Malaysia when you require local data residency, lower regional latency and greater control over where your datasets, applications and model weights are stored and processed.
From cost-efficient inference on NVIDIA L20 GPUs to large-scale AI reasoning and training on NVIDIA H200, B300 and GB300 NVL72 infrastructure, our GPU as a Service offering supports organisations at every stage of their AI journey.
Why GPU as a Service
Lower Capital Expenditure and Predictable Operating Costs
Avoid large upfront GPU purchases, hardware depreciation and the risk of investing in infrastructure that may become outdated.
Access Newer GPU Generations
Upgrade to newer GPU models or higher service tiers as your workload requirements grow, without replacing physical infrastructure.
Scale According to Your Workload
Increase GPU capacity for model training, fine-tuning or high-volume inference, then scale down to smaller instances for development, testing and steady-state production workloads.
Faster Time to Deployment
Access enterprise GPU infrastructure without managing procurement, installation, rack space, power, cooling, networking or hardware maintenance.
Malaysia-Based Deployment
Keep AI workloads closer to Malaysia-based applications, users and data sources to support data-residency requirements and reduce regional network latency.
Support for Both Training and Inference
Use the same GPU infrastructure for model development, fine-tuning, evaluation, batch processing and production inference.
Supported AI Workloads
AI Inference and Model Serving
- Large language model chatbots and AI assistants
- Retrieval-augmented generation, or RAG
- Vector embeddings and semantic search
- Agentic AI and multi-agent workflows
- AI reasoning and test-time compute
- Document understanding, extraction and summarisation
- Content classification and sentiment analysis
- Speech recognition, voice AI and text-to-speech
- Image and video generation
- Computer vision inference
- e-KYC, OCR, facial verification and liveness detection
- Fraud detection and anomaly identification
- Recommendation and personalisation engines
AI Training and Fine-Tuning
- LoRA and QLoRA fine-tuning for large language models
- Full-parameter model fine-tuning
- Domain-specific language model adaptation
- Reinforcement learning and post-training workloads
- Multimodal model training and fine-tuning
- Training and tuning image-generation models
- Object detection, classification and segmentation
- e-KYC, OCR and document fraud model development
- Synthetic data generation
- Distributed and multi-GPU model training
- Model benchmarking, evaluation and optimisation
Common Use Cases
University AI Centre of Excellence
Provide a shared GPU resource pool for students, lecturers and researchers.
GPU resources can be allocated across classes, laboratories, research projects, hackathons and student experimentation, with user quotas, access controls and cost monitoring.
Enterprise AI Centre of Excellence
Create a centralised GPU platform that enables different departments to experiment with and deploy AI applications while maintaining governance, security and resource controls.
In-House Research and Development
Conduct rapid prototyping, model benchmarking, model evaluation, internal tool development and innovation pilots without committing to long-term hardware ownership.
Train and Fine-Tune Proprietary Models
Fine-tune large language models using organisational or industry-specific knowledge, including financial, legal, healthcare, manufacturing and government datasets.
Train or optimise computer vision models for e-KYC, OCR, facial verification, liveness detection, quality inspection and document fraud detection.
Production AI Applications
Deploy scalable inference endpoints for:
- Customer service
- Enterprise knowledge search
- AI assistants and copilots
- Personalised recommendations
- Business analytics
- Compliance workflows
- Document processing
- Fraud monitoring
- Voice AI
- Image and video generation
- Agentic AI applications
GPU Options in Malaysia
Keep Your AI Workloads in Malaysia
These options are suitable for organisations that prioritise Malaysia-based deployment, local data residency and low-latency connectivity to Malaysia-based applications and users.
Final availability, instance configuration and commercial terms are subject to capacity confirmation.
NVIDIA L20 — 48GB
A cost-efficient option for production inference, computer vision, smaller language models and moderate fine-tuning workloads.
1 × L20 GPU
Instance: ecs.gni3cl.5xlarge
Compute: 22 vCPU
System memory: 120GB RAM
GPU memory: 48GB per GPU
Pricing: Contact us for the best available price
2 × L20 GPUs
Instance: ecs.gni3cl.11xlarge
Compute: 44 vCPU
System memory: 240GB RAM
GPU memory: 48GB per GPU
Total GPU memory: 96GB
Pricing: Contact us for the best available price
4 × L20 GPUs
Instance: ecs.gni3cl.22xlarge
Compute: 90 vCPU
System memory: 480GB RAM
GPU memory: 48GB per GPU
Total GPU memory: 192GB
Pricing: Contact us for the best available price
8 × L20 GPUs
Instance: ecs.gni3cl.45xlarge
Compute: 180 vCPU
System memory: 960GB RAM
GPU memory: 48GB per GPU
Total GPU memory: 384GB
Pricing: Contact us for the best available price
NVIDIA H100 — 80GB
A high-performance Hopper-generation GPU for demanding AI training, large-scale fine-tuning and production inference.
8 × H100 GPUs
Instance: ecs.hpcpni3h.42xlarge
Compute: 168 vCPU
System memory: 1,960GB RAM
GPU memory: 80GB per GPU
Total GPU memory: 640GB
Local storage: 3.84TB × 8
Network: 400G × 8 InfiniBand
Pricing: Contact us for the best available price
The NVIDIA H100 SXM configuration provides 80GB of HBM3 memory per GPU and supports high-speed NVLink connectivity for multi-GPU workloads. (NVIDIA)
NVIDIA H200 — 141GB HBM3e
Designed for memory-intensive large language models, RAG, larger context windows, high-throughput inference and advanced model fine-tuning.
Available Configuration
GPU: NVIDIA H200 Tensor Core GPU
GPU memory: 141GB HBM3e per GPU
Memory bandwidth: 4.8TB/s per GPU
Recommended deployment: Dedicated multi-GPU node
CPU, RAM, storage and network: Configured according to workload requirements
Pricing: Contact us for the best available price
An eight-GPU H200 deployment provides approximately 1.13TB of aggregate GPU memory, making it suitable for larger models, larger batch sizes and workloads that are constrained by GPU memory.
NVIDIA specifies 141GB of HBM3e memory and 4.8TB/s of memory bandwidth for each H200 GPU. (NVIDIA)
NVIDIA B300 Blackwell Ultra — 288GB HBM3e
A premium GPU tier for large-scale AI reasoning, agentic AI, advanced post-training, multimodal models and high-throughput training and inference.
8 × B300 GPUs
GPU: 8 × NVIDIA B300 Blackwell Ultra GPUs
GPU memory: 288GB HBM3e per GPU
Total GPU memory: Approximately 2.3TB
FP8 training performance: Up to 72 PFLOPS per eight-GPU system
FP4 inference performance: Up to 144 PFLOPS per eight-GPU system
Interconnect: Fifth-generation NVIDIA NVLink and NVSwitch
Network: Up to 800Gb/s InfiniBand or Ethernet connectivity, depending on configuration
CPU, RAM and storage: Configured according to deployment requirements
Pricing: Contact us for the best available price
The NVIDIA DGX B300 reference system combines eight B300 GPUs, each with 288GB of HBM3e memory, for approximately 2.3TB of total GPU memory. (NVIDIA Docs)
NVIDIA GB300 NVL72 — Rack-Scale AI Infrastructure
The highest infrastructure tier for frontier-scale AI reasoning, large multimodal models, distributed training and extremely high-throughput inference.
Unlike a conventional GPU virtual machine, the GB300 NVL72 is a fully liquid-cooled, rack-scale system that operates as one large NVLink compute domain.
GB300 NVL72 Configuration
GPU: 72 × NVIDIA Blackwell Ultra GPUs
CPU: 36 × NVIDIA Grace CPUs
GPU memory: Approximately 20TB
GPU memory bandwidth: Up to 576TB/s aggregate
CPU memory: Approximately 17TB LPDDR5X
Total fast memory: Approximately 37TB
CPU cores: 2,592 Arm Neoverse V2 cores
NVLink bandwidth: Approximately 130TB/s
Cooling: Fully liquid-cooled rack-scale architecture
Deployment: Dedicated rack or reserved infrastructure capacity
Pricing: Contact us for availability and customised commercial terms
GB300 NVL72 is designed for AI reasoning, test-time scaling, agentic AI, trillion-parameter-class models and large-scale AI factory deployments.
NVIDIA specifies that the GB300 NVL72 integrates 72 Blackwell Ultra GPUs and 36 Grace CPUs, with approximately 20TB of GPU memory, 576TB/s aggregate GPU memory bandwidth and 130TB/s of NVLink bandwidth. (NVIDIA)
GPU Model Selection Guide
L20 — Best Value for Inference
Best for:
- Production LLM inference
- RAG and semantic search
- Embeddings
- Computer vision
- OCR and e-KYC
- Smaller and quantised models
- Moderate LoRA and QLoRA fine-tuning
- Stable, continuous production workloads
Why choose it:
L20 provides a practical balance between GPU memory, performance and cost for organisations deploying mainstream enterprise AI applications.
H100 — Proven High-Performance Training
Best for:
- Large-scale model training
- Full-parameter fine-tuning
- Multi-GPU experimentation
- High-performance inference
- Scientific computing
- Performance-sensitive AI workloads
Why choose it:
H100 remains a widely supported enterprise GPU platform with a mature software ecosystem and strong training and inference performance.
H200 — More Memory for Larger Models
Best for:
- Larger large language models
- Long-context inference
- Larger RAG pipelines
- Larger batch sizes
- Memory-intensive fine-tuning
- High-throughput model serving
- Scientific and high-performance computing
Why choose it:
The H200 increases GPU memory to 141GB and provides 4.8TB/s of memory bandwidth, reducing memory bottlenecks for large models and data-intensive workloads. (NVIDIA Newsroom)
B300 Blackwell Ultra — Advanced Reasoning and Agentic AI
Best for:
- Agentic AI systems
- AI reasoning models
- Reinforcement learning and post-training
- Large-scale multimodal models
- Video-generation models
- Large-context and high-concurrency inference
- Frontier-scale model fine-tuning
- Advanced distributed training
Why choose it:
Each B300 GPU provides 288GB of HBM3e memory. An eight-GPU system delivers approximately 2.3TB of total GPU memory, providing substantially more room for large models, context windows and batch sizes. (NVIDIA Docs)
GB300 NVL72 — Rack-Scale AI Factory
Best for:
- Frontier-scale AI reasoning
- Test-time scaling
- Trillion-parameter-class models
- Large multimodal and mixture-of-experts models
- National or enterprise AI infrastructure
- High-volume token generation
- Distributed training
- Ultra-high-throughput inference
- AI factory and sovereign AI deployments
Why choose it:
GB300 NVL72 connects 72 Blackwell Ultra GPUs through a rack-scale NVLink architecture, providing approximately 20TB of GPU memory and 130TB/s of NVLink bandwidth. (NVIDIA)
Which GPU Should You Choose?
Choose L20 when cost-efficient inference and moderate fine-tuning are your main priorities.
Choose H100 when you need proven high-performance training and enterprise-grade multi-GPU computing.
Choose H200 when your workload requires more GPU memory for larger models, longer context windows or larger batch sizes.
Choose B300 when you are developing advanced reasoning models, agentic AI systems, multimodal applications or high-throughput generative AI services.
Choose GB300 NVL72 when you require dedicated rack-scale AI infrastructure for frontier models, large-scale reasoning or AI factory deployments.
Commercial and Technical Notes
- Monthly subscription discounts may be available for reserved capacity.
- On-demand, monthly reserved and longer-term dedicated arrangements may be offered subject to GPU availability.
- Final pricing depends on the GPU model, number of GPUs, deployment term, storage, networking and support requirements.
- GPU availability may vary according to capacity and delivery schedule.
- CPU, RAM, storage and network specifications for H200, B300 and GB300 deployments will be designed according to the selected architecture.
- Dedicated environments, private networking and custom security configurations are available upon request.
- Malaysia deployment supports local data residency but does not, by itself, guarantee compliance with every regulatory requirement.
- Customers remain responsible for evaluating the regulatory, data-protection and industry-specific requirements applicable to their workloads.
- For H100, H200, B300 and GB300 NVL72, please contact us to confirm capacity and obtain customised commercial terms.
Contact Us
Speak with our team to identify the most suitable GPU configuration for your AI workload.
Email: [email protected]
