An introduction to the NVIDIA B300: The Blackwell Ultra GPU

4 minutes reading time

Written by

Dinesh Majrekar
Dinesh Majrekar

Chief Technology Officer (CTO) at Civo

AI wasn't supposed to move this fast.

Twelve months ago, the H100 was still the benchmark everyone measured themselves against. Six months ago, the B200 changed the calculus for serious inference workloads. Now there's the B300, NVIDIA's Blackwell Ultra GPU, and it doesn't just move the goalposts. It takes them off the pitch entirely.

The B300 is the highest-performance GPU in the Blackwell family. It's purpose-built for the era of AI reasoning: multi-step thinking models, trillion-parameter inference, agentic AI pipelines that don't pause between steps. If AI is getting more demanding by the quarter, the B300 is the answer that meets it where it's going, not where it's been.

This blog breaks down what the B300 actually is, where it came from, what's changed inside the chip, and why it matters for the teams building real AI infrastructure right now.

Get a Blackwell Ultra B300 GPU today with Civo

Accelerate your most demanding inference, agentic AI, and large-scale generative workloads with dedicated B300 compute.

Talk to our team →

The history of the Blackwell architecture 

NVIDIA unveiled the Blackwell architecture at GTC in March 2024. It was positioned from day one as the platform for the next era of generative AI, not just incremental performance gains, but a rethinking of how GPUs are designed and deployed at scale.

💡 NVIDIA has a tradition of naming its GPU architectures after scientists who changed how we understand the world. Blackwell is named after David Harold Blackwell, one of the most important American statisticians and mathematicians of the twentieth century.

The original Blackwell GPU (B200) introduced dual-reticle die design, fifth-generation Tensor Cores, the NVFP4 precision format, and NVLink 5. It was, by every measure, a leap forward. But AI moved faster than even NVIDIA expected.

By the time Blackwell systems were shipping at scale, the workload landscape had already shifted. Reasoning models (models that think before they answer) were generating five, ten, sometimes fifty times more tokens per query than the models everyone had optimized for. Longer context windows. Heavier attention computation. Bigger KV caches. More memory pressure. The original Blackwell spec, impressive as it was, was being pushed to its limits almost immediately.

NVIDIA's answer was Blackwell Ultra: the B300. Announced in 2025 and shipping now, it takes the same dual-die architecture, the same TSMC 4NP manufacturing process, the same NVLink 5 fabric, and turns up every dimension that matters for inference at scale. More memory. More compute. Faster attention. The same chip, pushed to its full potential.

NVIDIA DGX B300 specifications

The DGX B300 is NVIDIA's turnkey deployment of Blackwell Ultra, eight NVIDIA Blackwell Ultra SXM in a single system, purpose-built for enterprise AI infrastructure.

It delivers 144 petaFLOPS of FP4 inference performance across 2.1 TB of total GPU memory. Two NVIDIA NVLink Switch Systems provide 14.4 TB/s of aggregate GPU-to-GPU bandwidth. Networking runs through eight OSFP ports serving ConnectX-8 VPI cards at up to 800 Gb/s, and two BlueField-3 DPUs add 400 Gb/s InfiniBand/Ethernet for data processing offload.

It's a 10U system, paired with Intel Xeon 6776P CPUs, and draws roughly 14 kW. Storage spans 2x 1.9 TB NVMe M.2 for the OS and 8x 3.84 TB NVMe E1.S for internal workload storage.

SpecDGX B300

GPU

8x NVIDIA Blackwell Ultra SXM

CPU

Intel® Xeon® 6776P Processors

Total GPU Memory

2.1 TB

FP4 Performance

144 PFLOPS (dense)

NVLink Bandwidth

14.4 TB/s aggregate

Networking

Up to 800 Gb/s InfiniBand/Ethernet

Storage

8x 3.84 TB NVMe E1.S internal

Power consumption

~14 kW

Rack units

10U

Software

NVIDIA AI Enterprise, NVIDIA Mission Control, DGX OS

What workloads is the B300 built for?

The B300 was not designed as an all-rounder. It was designed for the workloads where scale and latency both matter, where running a smaller GPU means either degrading the model or adding more machines.

UsageDescription

AI inference and reasoning at scale

The combination of 288GB memory and 15 petaFLOPS of NVFP4 compute means large reasoning models like DeepSeek-R1 (671B MoE) can run on fewer B300s with lower tensor parallelism overhead than any previous generation. Lower parallelism overhead means faster per-token latency and simpler infrastructure.

Frontier LLM training and fine-tuning

The memory capacity, bandwidth, and NVLink 5 interconnect make the B300 well-suited for training models with 200B+ parameters, where the primary challenge is keeping the GPUs fed with data and gradients across a distributed fabric.

Agentic AI pipelines

Agents don't run one query and wait. They chain together tool calls, retrieval, reasoning, and generation in sequences that can span hundreds of steps. The B300's faster attention and higher throughput mean individual steps complete faster, and the headroom to run multiple agents concurrently is substantially larger.

Multimodal AI

The B300 includes dedicated hardware for video and JPEG decoding (NVDEC and NVJPEG), enabling high-throughput image and video preprocessing directly on the GPU, critical for vision-language models and multimodal pipelines.

What's the B300 compared to previous generations?

FeatureH100 (Hopper)B200 (Blackwell)B300 (Blackwell Ultra)

HBM Memory

80 GB

192 GB

288 GB

Memory Bandwidth

3.35 TB/s

8 TB/s

8 TB/s

Dense NVFP4 Compute

2 PFLOPS (FP8)

10 PFLOPS

15 PFLOPS

NVLink Bandwidth (per GPU)

900 GB/s

1,800 GB/s

1,800 GB/s

Attention Acceleration

4.5 TeraExp/s

5 TeraExp/s

10.7 TeraExp/s

Max TGP

700W

1,200W

1,400W

Get a Blackwell Ultra B300 GPU today with Civo

The B300 is in stock and available for deployment on Civo now, starting from $5.45/hr. No surprise egress charges, no infrastructure headaches, no weeks-long provisioning cycles.

We include data ingress and egress as standard, so your compute budget goes to compute, not to moving data around. Pre-installed NVIDIA drivers and software mean you go from provisioned to productive in minutes.

The B300 joins our full range of NVIDIA GPU infrastructure, from L40S for graphics and rendering workloads through to H100, H200, B200, and now B300 for AI training and inference, all on the same platform, with the same transparent pricing, backed by the same team that actually picks up the phone.

Talk to our team about the B300 →

FAQs

Dinesh Majrekar
Dinesh Majrekar

Chief Technology Officer (CTO) at Civo

Dinesh Majrekar is Chief Technology Officer at Civo, where he leads the company’s technology strategy and platform development. His work focuses on building scalable cloud infrastructure and advancing the technologies that power the Civo platform.

Before becoming CTO, Dinesh served as Director of Innovation at Civo and held senior leadership roles at ServerChoice. His experience spans infrastructure architecture, platform engineering, and large-scale operations across hosting, cloud, and cybersecurity environments.

View author profile