NVIDIA B200 180GB, Confidential GPU Compute (Intel TDX, Blackwell)
What makes B200 different from H100 and H200
The B200 is NVIDIA's Blackwell-generation GPU and is the first data-center GPU with native FP4 tensor cores. On Blackwell, FP4 throughput reaches roughly 5× the FP8 throughput of an H100 (Hopper) and ~2.5× an H200, which means you train and serve frontier-scale models (400B–2T parameters) on fewer GPUs and lower wall-clock time. 180 GB of HBM3e at 8 TB/s(vs 141 GB / 4.8 TB/s on H200, 80 GB / 3.35 TB/s on H100) lets a single B200 hold a 70B model in FP16 with full KV cache headroom, or a 405B model sharded across just 4 GPUs instead of 8.
B200 is listed, never available to date, not attested. Our confidential tier seals each GPU inside an Intel TDX trust domain with trust-domain GPU isolation (GPU confidential-computing mode off on the container tier), AES-256 memory encryptionbacked by CPU-fused hardware keys, LUKS full-disk encryption, and hardware remote attestation. The hypervisor and the platform operator (VoltageGPU) sit outside the trust boundary, the same silicon-level confidential computing model used by Microsoft Azure Confidential VMs and Google Cloud Confidential Computing, but at $10.60/GPU/hour with per-second billinginstead of $14+/hr on Azure.
B200 vs H200 vs H100, quick spec comparison
| Metric | B200 | H200 | H100 |
|---|---|---|---|
| Architecture | Blackwell | Hopper refresh | Hopper |
| VRAM | 180 GB HBM3e | 141 GB HBM3e | 80 GB HBM3 |
| Memory bandwidth | 8 TB/s | 4.8 TB/s | 3.35 TB/s |
| Native FP4 tensor cores | Yes | No | No |
| NVLink generation | NVLink 5 (1.8 TB/s) | NVLink 4 (900 GB/s) | NVLink 4 (900 GB/s) |
| TDP | 1000 W | 700 W | 700 W |
| Confidential Computing | Intel TDX + trust-domain GPU isolation | Intel TDX + trust-domain GPU isolation | Intel TDX + trust-domain GPU isolation |
| VoltageGPU price | $10.60/hr | $6.58/hr | $5.00/hr |
See the dedicated pages for NVIDIA H200 141GB and NVIDIA H100 80GB, the live pricing comparison, and the full Intel TDX security architecture.
B200 memory: 192 GB or 180 GB? What NVIDIA actually documents
NVIDIA announced the B200 with 192 GB of HBM3e at launch, and that figure is still repeated across the web. The systems NVIDIA ships and documents carry less: the official DGX B200 specification lists 1,440 GB of GPU memory for eight GPUs, that is 180 GB HBM3e per B200, with 64 TB/s of aggregate bandwidth (8 TB/s per GPU); the HGX B200 page lists 1.4 TB for eight GPUs. We quote 180 GB everywhere on this site as of 12 September 2026 and treat 192 GB as the launch number, not the shipping one. If a provider quotes 192 GB, ask which system it is measured on.
When to actually choose B200 over H200
- Frontier-model pre-training, Training 70B–405B parameter models from scratch. FP8/FP4 mixed precision on Blackwell delivers roughly 2–3× wall-clock speedup vs H200 and reduces cluster size proportionally. NVLink 5 keeps gradient all-reduce from becoming the bottleneck.
- Long-context inference (128K+), 180 GB of HBM3e fits a 70B model plus a 256K-token KV cache without offloading. On H200 you'd start swapping past ~64K context.
- FP4 production inference, Quantizing DeepSeek-R1, Llama 3.1 405B, or GPT-OSS-120B to FP4 on B200 typically halves the per-token cost vs FP8 on H200, with negligible quality loss in evals. Native FP4 tensor cores mean no software emulation overhead.
- Multi-modal video / world models, Diffusion transformers and video-generation models are HBM-bandwidth bound; 8 TB/s is what unlocks them at production latency.
- Trillion-parameter MoE, DeepSeek-V3 (671B), Kimi-K2, GPT-OSS-120B route through experts that benefit from B200's combination of VRAM, bandwidth, and NVLink 5 expert-parallel sharding.
If you don't need any of the above, H200 at $6.58/hr is a better $/throughput pick. The B200 premium pays for itself only when you saturate FP4 or VRAM.
Pricing, single GPU, monthly equivalent, 8× cluster
- 1× B200 180GB Confidential: $10.60/hour (listed, never available to date, not attested; per-second billing, no commitment)
- Monthly equivalent (730h continuous): $7738/GPU/month
- 8× B200 NVLink 5 cluster: $84.8/hour (1.44 TB pooled HBM3e)
- Comparison: Azure ND B200 v6 lists at $28.50/hr per GPU (April 2026 reading) and is not a confidential VM; on 12 September 2026 no hyperscaler lists a confidential B200. Our B200 container runs inside an Intel TDX trust domain; GPU attestation is not verified by us on this SKU. Dated, sourced comparison: four clouds compared
- $5 referral credit on signup, no credit card required, Bitcoin and crypto accepted
Frequently asked, NVIDIA B200 confidential compute
Is the NVIDIA B200 192 GB or 180 GB?
180 GB HBM3e per GPU on the systems NVIDIA documents: the DGX B200 specification lists 1,440 GB for eight GPUs and the HGX B200 page lists 1.4 TB for eight. The 192 GB figure comes from the launch announcement. VoltageGPU quotes 180 GB, checked on 12 September 2026.
Should I rent B200 or H200?
Pick B200 if you need native FP4, train 405B+ models, serve 128K+ context, or run video diffusion workloads. For everything else, including most fine-tuning, RAG, and 7B–70B inference, H200 at$6.58/hr gives a better cost-per-throughput ratio. The B200's 8 TB/s HBM3e and FP4 tensor cores only earn their premium on workloads that actually saturate them.
Does Intel TDX add overhead on B200?
Independent benchmarks (Phoronix, Intel TDX 1.5 release notes) measure 3–7% throughput overhead from TDX on CPU-bound workloads. On B200 GPU workloads the overhead is below the noise floor because compute happens inside the GPU. trust-domain GPU isolation adds a one-time ~50ms warm-up at pod start; steady-state throughput matches non-confidential B200 within margin of error.
Can I train a 70B model on a single B200?
Yes for LoRA / QLoRA fine-tuning, a 70B model fits comfortably in 180 GB with optimizer states in FP16 plus 32K context KV cache. For full fine-tuning of 70B you'd use 2–4 B200s with NVLink 5 tensor parallelism. For 405B full fine-tuning, plan on a 4×–8× B200 cluster.
Is the B200 GPU GDPR / HIPAA compatible on VoltageGPU?
Our confidential tier runs GPUs inside Intel TDX trust domains with AES-256 memory encryption, and the tenant generates the attestation rather than taking our word for it. We have run that verification and published the evidence on H200, H100 including an eight-GPU node, and RTX PRO 6000 Blackwell. B200 is not in our inventory today, so we have not verified it and do not claim it. On GDPR, a DPA covering Article 28 processor obligations is available on request. On HIPAA, we implement the technical safeguards but do not offer a BAA yet; it is planned after SOC 2 Type I (see /trust). EU-based legal entity (VOLTAGE EI, France).
What you can deploy today, and where B200 stands
Sign up at voltagegpu.com/register with $5 referral credit and pick a template (vLLM, PyTorch, TensorRT-LLM, OpenClaw). B200 is not in our inventory today, so the verified confidential SKUs to choose from are H200, H100 including an eight-GPU node, and RTX PRO 6000 Blackwell. Provisioning a confidential instance takes a few minutes rather than seconds; our last measured run to a working SSH session on an eight-GPU node was six minutes. Billing is per second, and stopping the instance stops the meter.