Civo's GPU infrastructure: Full range and use cases
Written by
Chief Technology Officer (CTO) at Civo
Written by
Chief Technology Officer (CTO) at Civo
The GPU landscape has changed faster in the past two years than in the previous decade. What was a cutting-edge H100 setup in 2024 is now entry-level for frontier inference. Blackwell has landed, Blackwell Ultra is shipping, and Vera Rubin is already reserved.
This blog covers the complete Civo NVIDIA GPU range, every GPU we offer, what it's built for, how to match it to your workload, and why where you run your AI infrastructure matters as much as which GPU you choose.
The full Civo GPU range
Civo offers seven NVIDIA GPUs spanning five architectural generations. Every GPU in this range is available on transparent, hourly pricing with no egress fees, pre-installed NVIDIA drivers, and full Kubernetes-native support.
Each GPU explained
NVIDIA A100
The A100 is built on NVIDIA's Ampere architecture and remains the most versatile GPU in the range for teams that don't need frontier-generation compute. With 40GB or 80GB of HBM2e memory, it handles a wide band of workloads — AI training, inference, scientific computing, and data analytics — without the cost profile of newer hardware.
Key specs: 6,912 CUDA Cores | 432 Tensor Cores (3rd gen) | 40GB or 80GB HBM2e | 2,039 GB/s bandwidth | Up to 400W TGP | MIG support
Ideal for: Teams running models up to 34B parameters in FP16, scientific computing, and workloads where cost-efficiency matters more than raw throughput.
NVIDIA A100 GPU for AI workloads
Run training, inference, and data-intensive workloads with proven, high-performance GPU kubernetes and compute.
NVIDIA L40S
The L40S is the only GPU in the Civo range built on NVIDIA's Ada Lovelace architecture, and it's the only one using GDDR6 rather than HBM memory. That combination makes it uniquely suited to workloads that mix AI computation with graphics rendering, video encoding, and visual AI — tasks where HBM's bandwidth advantage matters less than Ada's graphics pipeline improvements.
Key specs: 18,176 CUDA Cores | 568 Tensor Cores (4th gen) | 48GB GDDR6 | 864 GB/s bandwidth | Up to 300W TGP | No MIG support
Ideal for: Generative AI applications with visual outputs, media and entertainment workflows, inference combined with rendering, and streaming/encoding pipelines.
NVIDIA L40s GPU for inference and rendering
Run inference, rendering, and real-time workloads with fast, flexible GPU compute.
NVIDIA H100
The H100 is NVIDIA's flagship Hopper-architecture GPU and defined the standard for serious AI infrastructure from 2022 through 2025. Fourth-generation Tensor Cores, a dedicated Transformer Engine with FP8 support, and NVLink 4 interconnect make it well-suited for large-scale LLM training and distributed inference across multiple GPUs.
Available in SXM and PCIe configurations. The SXM variant delivers higher memory bandwidth and supports NVLink for multi-GPU communication — the right choice for distributed training. PCIe is better suited to single-GPU inference workloads where interconnect bandwidth is less critical.
Key specs: 16,896 CUDA Cores | 528 Tensor Cores (4th gen) | 80GB HBM2e | 3,350 GB/s bandwidth | Up to 700W TGP | MIG support
Ideal for: LLM training at scale, large-scale inference, distributed multi-GPU workloads, and teams that need a proven, well-supported production platform.
H100 GPU for AI workloads
Run large-scale AI training and inference on H100 GPUs with scalable compute and Kubernetes.
NVIDIA H200
The H200 shares its compute die with the H100 — same Hopper architecture, same Tensor Core count, same FP8 throughput. What changed is the memory: the H200 replaces HBM2e with HBM3e and nearly doubles capacity to 141GB, with a 43% bandwidth improvement to 4.8 TB/s.
In practice, this means the H200 and H100 perform similarly on compute-bound training runs, but the H200 pulls ahead significantly on inference workloads where large models previously had to be spread across multiple H100s. For teams serving 70B parameter models, the H200 often halves the GPU count required.
Key specs: 16,896 CUDA Cores | 528 Tensor Cores (4th gen) | 141GB HBM3e | 4,800 GB/s bandwidth | Up to 700W TGP | MIG support
Ideal for: Large LLM inference where memory is the binding constraint, serving 70B models on fewer GPUs, and teams upgrading from H100 who need more headroom before moving to Blackwell.
NVIDIA H200 GPU for advanced AI workloads
Run large models and memory-intensive training with high-performance GPU compute.
NVIDIA B200
The B200 is the first GPU in NVIDIA's Blackwell generation and represents the largest architectural leap in the range. Key changes: a dual-reticle die design with 208 billion transistors, fifth-generation Tensor Cores with NVFP4 support, 192GB of HBM3e at 8 TB/s, and NVLink 5 at 1.8 TB/s per GPU — double the Hopper generation's interconnect bandwidth.
The introduction of NVFP4 is the most significant capability shift. At 4-bit precision with hardware-accelerated quantization, the B200 delivers 10 petaFLOPS of dense compute at nearly FP8-equivalent accuracy — 5x more throughput than the H100 or H200 for inference workloads running NVFP4. The practical result: more concurrent model instances, lower cost per token, and models previously requiring multi-GPU H200 setups running on a single B200.
Key specs: 20,480 CUDA Cores | 640 Tensor Cores (5th gen) | 192GB HBM3e | 8,000 GB/s bandwidth | 10 PFLOPS dense NVFP4 | ~1,200W TGP | MIG support | NVLink 5
Ideal for: Frontier LLM inference, training models from 130B to 300B parameters, teams moving off Hopper who need a serious step up in throughput, and generative AI at production scale.
Get a NVIDIA Blackwell B200 GPU today
Deploy NVIDIA's Blackwell B200 architecture, purpose-built for AI, ML, and high-performance computing
NVIDIA B300 (Blackwell Ultra)
The B300 is the highest-performance GPU in the Blackwell lineup — the same dual-die architecture as the B200, pushed to its maximum. Memory capacity increases by 50% to 288GB, NVFP4 throughput increases by 50% to 15 petaFLOPS, and SFU throughput for attention-layer operations is doubled, delivering 2x faster softmax computation for reasoning models with long context windows.
These three changes were targeted at exactly the workloads that arrived faster than the B200 was designed for: long-chain-of-thought reasoning models, large MoE architectures like DeepSeek-R1 (671B), and agentic AI pipelines that generate far more tokens per request than standard LLM inference. In MLPerf Inference v6.0 (April 2026), Blackwell Ultra systems delivered 2.5 million tokens per second on DeepSeek-R1 and achieved up to 50x higher throughput per megawatt than Hopper for low-latency agentic workloads.
Key specs: 20,480 CUDA Cores | 640 Tensor Cores (5th gen) | 288GB HBM3e | 8,000 GB/s bandwidth | 15 PFLOPS dense NVFP4 | ~1,400W TGP | MIG support | NVLink 5
Ideal for: AI reasoning workloads, agentic AI, frontier LLMs with 200B+ parameters, DeepSeek-scale inference, and any workload where the B200 was hitting memory limits.
Get a Blackwell Ultra B300 GPU today with Civo
Accelerate your most demanding inference, agentic AI, and large-scale generative workloads with dedicated B300 compute.
NVIDIA Vera Rubin (coming Q1 2027)
Vera Rubin is NVIDIA's next-generation AI platform — the successor to Blackwell entirely, not an iteration of it. Built on TSMC's N3 process with 336 billion transistors, HBM4 memory delivering 22 TB/s of bandwidth, and NVLink 6 at 3.6 TB/s per GPU, the Rubin GPU delivers 50 petaFLOPS of dense FP4 inference — 3.3x the B300, 25x the H100.
The platform is designed for the workloads that define 2027 and beyond: trillion-parameter mixture-of-experts models, rack-scale agentic AI systems, and inference pipelines that need to serve tens of millions of tokens per second. NVIDIA has confirmed that Vera Rubin can train MoE models with one-quarter the number of GPUs compared to Blackwell, and in NVL72 configurations delivers up to 35x inference performance per watt for trillion-parameter models relative to Blackwell.
Civo has confirmed early-access allocation for Vera Rubin infrastructure, with delivery from Q1 2027, from $11.00/hr. Allocations are limited and first-come, first-served.
Key specs: 288GB HBM4 | 22,000 GB/s bandwidth | 50 PFLOPS dense FP4 | NVLink 6 at 3.6 TB/s | TSMC N3 | 336B transistors
Ideal for: Large-scale training and inference for the next generation of frontier AI models, teams building infrastructure for 2027+, and organisations running trillion-parameter MoE workloads.
Reserve your Vera Rubin capacity
2,016 Vera Rubin GPUs. Q1 2027 delivery confirmed. Pricing from $11.00/hr. Allocations are first-come, first-served. Once they are gone, they are gone.
GPU selection guide: workload to hardware
Different workloads stress different parts of the GPU. Matching your workload to the right hardware is the fastest way to reduce both latency and cost.
Inference (production serving)
Production inference demands high throughput and low latency simultaneously. The binding constraints are usually memory capacity (can the full model fit on the GPU?), memory bandwidth (how fast can the model weights be accessed?), and compute throughput (how many tokens can be generated per second per GPU?).
For standard LLM inference at FP8 or NVFP4 precision, the B200 and B300 are the recommended starting points. Their NVFP4 compute throughput — 10 and 15 petaFLOPS respectively — is 5-7.5x higher than H100 or H200 FP8 throughput, which translates directly to more concurrent users, lower cost per million tokens, and faster time-to-first-token.
For reasoning models specifically (models that generate long chains of thought before answering), the B300's 2x faster attention computation makes it the better choice — softmax and attention-layer latency becomes a meaningful bottleneck in long-context inference that the B300's doubled SFU throughput directly addresses.
Recommended: B300 for reasoning/agentic inference; B200 for standard LLM inference at scale; H200 for teams serving 70B models who haven't yet moved to Blackwell.
Fine-tuning
Fine-tuning memory requirements depend heavily on the method. Full fine-tuning requires storing the full model, gradients, and optimiser states in VRAM — roughly 16-20 bytes per parameter in FP16 with Adam. LoRA and QLoRA reduce this significantly by only training adapter weights, but the base model still needs to be loaded.
For full fine-tuning of 7B-13B models, the A100 80GB or H100 provide sufficient headroom. For 34B-70B models with LoRA, the H200's 141GB is often the minimum; full fine-tuning at that scale typically requires a B200 or B300. For frontier models (130B+), the B300's 288GB is the practical minimum for single-GPU fine-tuning at NVFP4 precision.
Recommended: A100 80GB or H100 for 7B-34B fine-tuning; H200 or B200 for 34B-70B; B300 for 70B+ full fine-tuning or 130B+ LoRA.
Full training (pretraining from scratch)
Pre-training from scratch is where NVLink bandwidth and multi-GPU scaling matter most. Gradient synchronisation across GPUs is a core bottleneck in distributed training — the faster each GPU can communicate with its peers, the higher the effective utilisation across the cluster.
The H100 SXM (NVLink 4 at 900 GB/s per GPU) is a well-established baseline for distributed training. The B200 and B300 (NVLink 5 at 1.8 TB/s per GPU) double that interconnect bandwidth, which translates to meaningfully higher scaling efficiency in large clusters. For the next generation of frontier training runs — models at 200B+ parameters — Vera Rubin's NVLink 6 at 3.6 TB/s is the right infrastructure to plan around.
Recommended: H100 SXM for established distributed training workloads; B200 or B300 for new frontier training runs; Vera Rubin for 2027+ training infrastructure.
RAG (Retrieval-Augmented Generation)
RAG pipelines combine embedding generation, vector search, and LLM inference. The GPU workload is typically inference-only: encode a query, retrieve relevant context, and run a completion. Models are usually in the 7B-13B range, though 70B models are increasingly common for higher-quality retrieval and generation.
Because RAG is inference-dominated and model sizes are often moderate, the A100, L40S, and H100 all handle it effectively. For production-scale RAG with high concurrency or 70B+ generation models, the B200 or B300 provide the throughput needed to keep latency low under load.
Recommended: A100 or H100 for standard RAG pipelines; B200 or B300 for high-concurrency production RAG with large generation models.
VRAM planning: Model size to minimum GPU
VRAM requirements depend on model size, precision, and workload type. The table below gives practical minimum recommendations for inference (model weights only — add 20-30% headroom for KV cache and activations in production).
Note: These figures are for model weights only. Production inference typically needs an additional 20–30% for KV cache, activations, and batch overhead.
Get started
Every GPU in the Civo range is available via our product pages, with transparent pricing and no commitment required for on-demand access.
- A100: www.civo.com/ai/a100-gpu
- L40S: www.civo.com/ai/l40s-gpu
- H100: www.civo.com/ai/h100-gpu
- H200: www.civo.com/ai/h200-gpu
- B200: www.civo.com/ai/b200-blackwell-gpu
- B300: www.civo.com/ai/b300-blackwell-ultra-gpu
- Vera Rubin (early access reservation — Q1 2027 delivery): www.civo.com/ai/vera-rubin-gpu
Not sure which GPU fits your workload? Talk to our team and we'll help you work it out.

Chief Technology Officer (CTO) at Civo
Dinesh Majrekar is Chief Technology Officer at Civo, where he leads the company’s technology strategy and platform development. His work focuses on building scalable cloud infrastructure and advancing the technologies that power the Civo platform.
Before becoming CTO, Dinesh served as Director of Innovation at Civo and held senior leadership roles at ServerChoice. His experience spans infrastructure architecture, platform engineering, and large-scale operations across hosting, cloud, and cybersecurity environments.
Share this article