Civo's GPU infrastructure: Full range and use cases

10 minutes reading time

Written by

Dinesh Majrekar
Dinesh Majrekar

Chief Technology Officer (CTO) at Civo

The GPU landscape has changed faster in the past two years than in the previous decade. What was a cutting-edge H100 setup in 2024 is now entry-level for frontier inference. Blackwell has landed, Blackwell Ultra is shipping, and Vera Rubin is already reserved.

This blog covers the complete Civo NVIDIA GPU range, every GPU we offer, what it's built for, how to match it to your workload, and why where you run your AI infrastructure matters as much as which GPU you choose.

The full Civo GPU range

Civo offers seven NVIDIA GPUs spanning five architectural generations. Every GPU in this range is available on transparent, hourly pricing with no egress fees, pre-installed NVIDIA drivers, and full Kubernetes-native support.

GPUArchitectureMemoryBandwidthDense ComputeStatusBest for

Ampere

40GB / 80GB HBM2e

2.0 TB/s

312 TFLOPS (TF32)

Available

Versatile AI/ML, HPC, mixed workloads

Ada Lovelace

48GB GDDR6

864 GB/s

362 TFLOPS (TF32)

Available

Graphics, rendering, mixed AI/visual

Hopper

80GB HBM2e

3.35 TB/s

3.9 PFLOPS (FP8)

Available

LLM training, large-scale AI

Hopper

141GB HBM3e

4.8 TB/s

3.9 PFLOPS (FP8)

Available

Large LLM inference, memory-bound workloads

Blackwell

192GB HBM3e

8 TB/s

10 PFLOPS (NVFP4)

Available

Frontier inference, training 130B+ models

Blackwell Ultra

288GB HBM3e

8 TB/s

15 PFLOPS (NVFP4)

Available from $5.45/hr

Reasoning models, agentic AI, frontier LLMs

Rubin

288GB HBM4

22 TB/s

50 PFLOPS (FP4)

Early access — Q1 2027 delivery

Rack-scale training, trillion-parameter inference

Each GPU explained

NVIDIA A100

The A100 is built on NVIDIA's Ampere architecture and remains the most versatile GPU in the range for teams that don't need frontier-generation compute. With 40GB or 80GB of HBM2e memory, it handles a wide band of workloads — AI training, inference, scientific computing, and data analytics — without the cost profile of newer hardware.

Key specs: 6,912 CUDA Cores | 432 Tensor Cores (3rd gen) | 40GB or 80GB HBM2e | 2,039 GB/s bandwidth | Up to 400W TGP | MIG support

Ideal for: Teams running models up to 34B parameters in FP16, scientific computing, and workloads where cost-efficiency matters more than raw throughput.

NVIDIA A100 GPU for AI workloads

Run training, inference, and data-intensive workloads with proven, high-performance GPU kubernetes and compute.

Access the A100 on Civo →

NVIDIA L40S

The L40S is the only GPU in the Civo range built on NVIDIA's Ada Lovelace architecture, and it's the only one using GDDR6 rather than HBM memory. That combination makes it uniquely suited to workloads that mix AI computation with graphics rendering, video encoding, and visual AI — tasks where HBM's bandwidth advantage matters less than Ada's graphics pipeline improvements.

Key specs: 18,176 CUDA Cores | 568 Tensor Cores (4th gen) | 48GB GDDR6 | 864 GB/s bandwidth | Up to 300W TGP | No MIG support

Ideal for: Generative AI applications with visual outputs, media and entertainment workflows, inference combined with rendering, and streaming/encoding pipelines.

NVIDIA L40s GPU for inference and rendering

Run inference, rendering, and real-time workloads with fast, flexible GPU compute.

Access the L40S on Civo →

NVIDIA H100

The H100 is NVIDIA's flagship Hopper-architecture GPU and defined the standard for serious AI infrastructure from 2022 through 2025. Fourth-generation Tensor Cores, a dedicated Transformer Engine with FP8 support, and NVLink 4 interconnect make it well-suited for large-scale LLM training and distributed inference across multiple GPUs.

Available in SXM and PCIe configurations. The SXM variant delivers higher memory bandwidth and supports NVLink for multi-GPU communication — the right choice for distributed training. PCIe is better suited to single-GPU inference workloads where interconnect bandwidth is less critical.

Key specs: 16,896 CUDA Cores | 528 Tensor Cores (4th gen) | 80GB HBM2e | 3,350 GB/s bandwidth | Up to 700W TGP | MIG support

Ideal for: LLM training at scale, large-scale inference, distributed multi-GPU workloads, and teams that need a proven, well-supported production platform.

H100 GPU for AI workloads

Run large-scale AI training and inference on H100 GPUs with scalable compute and Kubernetes.

Access the H100 on Civo →

NVIDIA H200

The H200 shares its compute die with the H100 — same Hopper architecture, same Tensor Core count, same FP8 throughput. What changed is the memory: the H200 replaces HBM2e with HBM3e and nearly doubles capacity to 141GB, with a 43% bandwidth improvement to 4.8 TB/s.

In practice, this means the H200 and H100 perform similarly on compute-bound training runs, but the H200 pulls ahead significantly on inference workloads where large models previously had to be spread across multiple H100s. For teams serving 70B parameter models, the H200 often halves the GPU count required.

Key specs: 16,896 CUDA Cores | 528 Tensor Cores (4th gen) | 141GB HBM3e | 4,800 GB/s bandwidth | Up to 700W TGP | MIG support

Ideal for: Large LLM inference where memory is the binding constraint, serving 70B models on fewer GPUs, and teams upgrading from H100 who need more headroom before moving to Blackwell.

NVIDIA H200 GPU for advanced AI workloads

Run large models and memory-intensive training with high-performance GPU compute.

Access the H200 on Civo →

NVIDIA B200

The B200 is the first GPU in NVIDIA's Blackwell generation and represents the largest architectural leap in the range. Key changes: a dual-reticle die design with 208 billion transistors, fifth-generation Tensor Cores with NVFP4 support, 192GB of HBM3e at 8 TB/s, and NVLink 5 at 1.8 TB/s per GPU — double the Hopper generation's interconnect bandwidth.

The introduction of NVFP4 is the most significant capability shift. At 4-bit precision with hardware-accelerated quantization, the B200 delivers 10 petaFLOPS of dense compute at nearly FP8-equivalent accuracy — 5x more throughput than the H100 or H200 for inference workloads running NVFP4. The practical result: more concurrent model instances, lower cost per token, and models previously requiring multi-GPU H200 setups running on a single B200.

Key specs: 20,480 CUDA Cores | 640 Tensor Cores (5th gen) | 192GB HBM3e | 8,000 GB/s bandwidth | 10 PFLOPS dense NVFP4 | ~1,200W TGP | MIG support | NVLink 5

Ideal for: Frontier LLM inference, training models from 130B to 300B parameters, teams moving off Hopper who need a serious step up in throughput, and generative AI at production scale.

Get a NVIDIA Blackwell B200 GPU today

Deploy NVIDIA's Blackwell B200 architecture, purpose-built for AI, ML, and high-performance computing

Access the B200 on Civo →

NVIDIA B300 (Blackwell Ultra)

The B300 is the highest-performance GPU in the Blackwell lineup — the same dual-die architecture as the B200, pushed to its maximum. Memory capacity increases by 50% to 288GB, NVFP4 throughput increases by 50% to 15 petaFLOPS, and SFU throughput for attention-layer operations is doubled, delivering 2x faster softmax computation for reasoning models with long context windows.

These three changes were targeted at exactly the workloads that arrived faster than the B200 was designed for: long-chain-of-thought reasoning models, large MoE architectures like DeepSeek-R1 (671B), and agentic AI pipelines that generate far more tokens per request than standard LLM inference. In MLPerf Inference v6.0 (April 2026), Blackwell Ultra systems delivered 2.5 million tokens per second on DeepSeek-R1 and achieved up to 50x higher throughput per megawatt than Hopper for low-latency agentic workloads.

Key specs: 20,480 CUDA Cores | 640 Tensor Cores (5th gen) | 288GB HBM3e | 8,000 GB/s bandwidth | 15 PFLOPS dense NVFP4 | ~1,400W TGP | MIG support | NVLink 5

Ideal for: AI reasoning workloads, agentic AI, frontier LLMs with 200B+ parameters, DeepSeek-scale inference, and any workload where the B200 was hitting memory limits.

Get a Blackwell Ultra B300 GPU today with Civo

Accelerate your most demanding inference, agentic AI, and large-scale generative workloads with dedicated B300 compute.

Talk to our team →

NVIDIA Vera Rubin (coming Q1 2027)

Vera Rubin is NVIDIA's next-generation AI platform — the successor to Blackwell entirely, not an iteration of it. Built on TSMC's N3 process with 336 billion transistors, HBM4 memory delivering 22 TB/s of bandwidth, and NVLink 6 at 3.6 TB/s per GPU, the Rubin GPU delivers 50 petaFLOPS of dense FP4 inference — 3.3x the B300, 25x the H100.

The platform is designed for the workloads that define 2027 and beyond: trillion-parameter mixture-of-experts models, rack-scale agentic AI systems, and inference pipelines that need to serve tens of millions of tokens per second. NVIDIA has confirmed that Vera Rubin can train MoE models with one-quarter the number of GPUs compared to Blackwell, and in NVL72 configurations delivers up to 35x inference performance per watt for trillion-parameter models relative to Blackwell.

Civo has confirmed early-access allocation for Vera Rubin infrastructure, with delivery from Q1 2027, from $11.00/hr. Allocations are limited and first-come, first-served.

Key specs: 288GB HBM4 | 22,000 GB/s bandwidth | 50 PFLOPS dense FP4 | NVLink 6 at 3.6 TB/s | TSMC N3 | 336B transistors

Ideal for: Large-scale training and inference for the next generation of frontier AI models, teams building infrastructure for 2027+, and organisations running trillion-parameter MoE workloads.

Reserve your Vera Rubin capacity

2,016 Vera Rubin GPUs. Q1 2027 delivery confirmed. Pricing from $11.00/hr. Allocations are first-come, first-served. Once they are gone, they are gone.

Contact the Civo sales team to reserve today >

GPU selection guide: workload to hardware

Different workloads stress different parts of the GPU. Matching your workload to the right hardware is the fastest way to reduce both latency and cost.

Inference (production serving)

Production inference demands high throughput and low latency simultaneously. The binding constraints are usually memory capacity (can the full model fit on the GPU?), memory bandwidth (how fast can the model weights be accessed?), and compute throughput (how many tokens can be generated per second per GPU?).

For standard LLM inference at FP8 or NVFP4 precision, the B200 and B300 are the recommended starting points. Their NVFP4 compute throughput — 10 and 15 petaFLOPS respectively — is 5-7.5x higher than H100 or H200 FP8 throughput, which translates directly to more concurrent users, lower cost per million tokens, and faster time-to-first-token.

For reasoning models specifically (models that generate long chains of thought before answering), the B300's 2x faster attention computation makes it the better choice — softmax and attention-layer latency becomes a meaningful bottleneck in long-context inference that the B300's doubled SFU throughput directly addresses.

Recommended: B300 for reasoning/agentic inference; B200 for standard LLM inference at scale; H200 for teams serving 70B models who haven't yet moved to Blackwell.

Fine-tuning

Fine-tuning memory requirements depend heavily on the method. Full fine-tuning requires storing the full model, gradients, and optimiser states in VRAM — roughly 16-20 bytes per parameter in FP16 with Adam. LoRA and QLoRA reduce this significantly by only training adapter weights, but the base model still needs to be loaded.

For full fine-tuning of 7B-13B models, the A100 80GB or H100 provide sufficient headroom. For 34B-70B models with LoRA, the H200's 141GB is often the minimum; full fine-tuning at that scale typically requires a B200 or B300. For frontier models (130B+), the B300's 288GB is the practical minimum for single-GPU fine-tuning at NVFP4 precision.

Recommended: A100 80GB or H100 for 7B-34B fine-tuning; H200 or B200 for 34B-70B; B300 for 70B+ full fine-tuning or 130B+ LoRA.

Full training (pretraining from scratch)

Pre-training from scratch is where NVLink bandwidth and multi-GPU scaling matter most. Gradient synchronisation across GPUs is a core bottleneck in distributed training — the faster each GPU can communicate with its peers, the higher the effective utilisation across the cluster.

The H100 SXM (NVLink 4 at 900 GB/s per GPU) is a well-established baseline for distributed training. The B200 and B300 (NVLink 5 at 1.8 TB/s per GPU) double that interconnect bandwidth, which translates to meaningfully higher scaling efficiency in large clusters. For the next generation of frontier training runs — models at 200B+ parameters — Vera Rubin's NVLink 6 at 3.6 TB/s is the right infrastructure to plan around.

Recommended: H100 SXM for established distributed training workloads; B200 or B300 for new frontier training runs; Vera Rubin for 2027+ training infrastructure.

RAG (Retrieval-Augmented Generation)

RAG pipelines combine embedding generation, vector search, and LLM inference. The GPU workload is typically inference-only: encode a query, retrieve relevant context, and run a completion. Models are usually in the 7B-13B range, though 70B models are increasingly common for higher-quality retrieval and generation.

Because RAG is inference-dominated and model sizes are often moderate, the A100, L40S, and H100 all handle it effectively. For production-scale RAG with high concurrency or 70B+ generation models, the B200 or B300 provide the throughput needed to keep latency low under load.

Recommended: A100 or H100 for standard RAG pipelines; B200 or B300 for high-concurrency production RAG with large generation models.

VRAM planning: Model size to minimum GPU

VRAM requirements depend on model size, precision, and workload type. The table below gives practical minimum recommendations for inference (model weights only — add 20-30% headroom for KV cache and activations in production).

Model sizeFP16/BF16 weightsFP8 weightsNVFP4 weightsMinimum GPU for inference

Up to 3B

~6GB

~3GB

~1.5GB

L40S, A100 40GB

3B–7B

6–14GB

3–7GB

1.5–3.5GB

L40S, A100 40GB

7B–13B

14–26GB

7–13GB

3.5–6.5GB

A100 80GB, H100

13B–34B

26–68GB

13–34GB

6.5–17GB

34B–70B

68–140GB

34–70GB

17–35GB

H200 (FP8/FP4), B200 (FP16)

70B–130B

140–260GB

70–130GB

35–65GB

B200 or B300 (FP8/FP4)

130B–200B

260–400GB

130–200GB

65–100GB

B300 (NVFP4), multi-GPU B200

200B+ (dense)

400GB+

200GB+

100GB+

B300 cluster, Vera Rubin

671B MoE (e.g. DeepSeek-R1)

~1.3TB

~670GB

~335GB

B300 cluster (4+ GPUs), Vera Rubin

Note: These figures are for model weights only. Production inference typically needs an additional 20–30% for KV cache, activations, and batch overhead.

Get started

Every GPU in the Civo range is available via our product pages, with transparent pricing and no commitment required for on-demand access.

Not sure which GPU fits your workload? Talk to our team and we'll help you work it out.

Dinesh Majrekar
Dinesh Majrekar

Chief Technology Officer (CTO) at Civo

Dinesh Majrekar is Chief Technology Officer at Civo, where he leads the company’s technology strategy and platform development. His work focuses on building scalable cloud infrastructure and advancing the technologies that power the Civo platform.

Before becoming CTO, Dinesh served as Director of Innovation at Civo and held senior leadership roles at ServerChoice. His experience spans infrastructure architecture, platform engineering, and large-scale operations across hosting, cloud, and cybersecurity environments.

View author profile