Back to Blog

vLLM 0.30 Under Confidential Computing: What Serial Alice Found on Our H100 VM, How We Reproduced It, and the One-Variable Workaround

Serial Alice ran a full attestation test on our single-GPU H100 Confidential VM and found that vLLM 0.30 gives incoherent answers with confidential computing on, while HF transformers is correct. We reproduced it in seven runs the same day. Setting VLLM_USE_V2_MODEL_RUNNER=0 fixes it on that VM. The cause is not proven, and the limits are listed.

Key Takeaways

  • Serial Alice found it, not us. On 1 October 2026 they ran a full attestation test on our single-GPU H100 Confidential VM and reported that vLLM 0.30 gives incoherent answers with confidential computing on, while HF transformers on the same GPU is correct.
  • We reproduced it the same day, in seven runs. Default, eager, fp32 and synchronous launches all failed. The fault followed the new V2 model runner, not attention, CUDA graphs or precision.
  • The workaround is one variable: VLLM_USE_V2_MODEL_RUNNER=0. With it, vLLM answers correctly on that VM, in eager mode and in the default compiled mode.
  • The cause is not proven. No control with confidential computing off on the same tier, one small model, six questions, engine logs not kept. Every limit is listed below.

Our reproduction files

The two scripts we ran and the outputs of the seven runs, with SHA-256 checksums. The README says how the result lines were copied from the terminal and what was added by hand. Serial Alice’s certificate is public at its own verify link, linked below.

Open the README →Serial Alice certificate
results.jsonl 4,873 Bvllm_cc_test.py 2,182 Bvllm_cc_test2.py 1,909 BSHA256SUMS

Attestation write-ups, ours included, usually stop at the attestation: the quote verifies, the GPU token verifies, done. Serial Alice went one step further on our hardware. They checked that a language model actually gives the right answers inside the confidential VM, and bound those answers to the attestation. That second check is how they found a problem we had not seen: vLLM, a widely used open-source inference engine, in its current default configuration, returned nonsense on that machine.

This is a joint write-up. The test, the finding and the certificate are Serial Alice’s. The reproduction and the workaround are ours. Both sides state what they did not prove.

What Serial Alice set out to test

Serial Alice is a Portuguese company, founded by Nelson Vicente, that issues signed, independently verifiable energy certificates for compute jobs. Their goal on our platform was specific: produce a coherent LLM answer with confidential computing on, and bind it to verifiable evidence. In practice that means an Intel TDX quote whose 64 bytes of report_data commit to a challenge generated before inference and to a manifest linking the answers and the energy samples, plus an NVIDIA GPU attestation carrying the same challenge as its nonce.

They ran it on 1 October 2026. The single-GPU H200 Confidential VM was sold out that day, so the test ran on the single-GPU H100 Confidential VM. Everything below is an H100 result, not an H200 one.

What Serial Alice verified on our VM, independently

On the machine itself:

  • nvidia-smi conf-compute -q reported CC State ON, the CPU was Intel TDX, and both /dev/tdx_guest and configfs-tsm were present.
  • Qwen2.5-1.5B-Instruct under HF transformers, greedy decoding, gave coherent answers. In fp32, all six answers were byte-identical to a CPU-only reference run.
  • The NVIDIA GPU attestation used Serial Alice’s challenge as its nonce.

Then off the VM, with their own tools:

  • The TDX quote chains to the Intel root.
  • The full 64-byte report_data matches what they expected, challenge included.
  • The NVIDIA attestation tokens carry valid ES384 signatures against NVIDIA’s published keys, eat_nonce equals their challenge, the overall result is true, the measurements match, secure boot is on and debug is disabled.
  • They tried eight tampering and replay cases: an altered answer, altered energy data or an altered model; a quote or an NVIDIA token swapped between runs; a wrong challenge. All eight were rejected.

The result is a public certificate, sa-b8f6f5aab6774c82bae5108ae603a64d, with tee_attested = true (the issuer verified the TDX quote chain and accepted it under its policy, with the quote’s MRTD on its allowlist) and tee_quote_bound = true (the quote’s report_data binds to the submitted samples hash). It is signed with Ed25519 and with the post-quantum ML-DSA scheme, and the verify link checks both signatures.

One field you will notice on the certificate. It reads tcb_status: unknown, with a warning that platform TCB currency is unproven. That is Serial Alice’s production issuer, which could not establish the TCB status at issuance; the exact cause was not recorded. A separate external verification of the same quote reported UpToDate. Both are true, and Serial Alice has since added a retry and recorded failure reasons for future certificates.

The finding: vLLM output is incoherent with confidential computing on

With HF transformers correct on the GPU, Serial Alice ran the same test through vLLM 0.30. The output was incoherent, and it stayed incoherent in every configuration they tried: default, eager, fp32 with Triton attention, FlashAttention, pinned memory off, custom ops off. They also checked the two obvious suspects. The weights on the GPU matched the file, and vLLM received the correct prompt tokens. The forward pass was still wrong.

Their conclusion was careful, and we keep it: without a control run with confidential computing off on the same VM tier, they did not claim that confidential computing is the cause, only that this vLLM path misbehaves on this platform.

Our reproduction, the same day

We rented the same tier, a single-GPU H100 Confidential VM, and confirmed CC status: ON with nvidia-smi conf-compute -f right after the last run. Image ubuntu-24-04-lts-595, torch 2.13.0+cu130, CUDA 13.0, vLLM 0.30.0. Same model as Serial Alice, Qwen/Qwen2.5-1.5B-Instruct, greedy decoding (temperature 0), 48 new tokens, six short questions with known answers in every run.

RunConfigurationModel runnerAnswers
1HF transformers, bf16n/aCoherent
2vLLM default (torch.compile, CUDA graphs, FlashAttention 3)V2Incoherent
3vLLM enforce_eagerV2Incoherent
4vLLM fp32 (Triton attention)V2Incoherent
5vLLM eager, CUDA_LAUNCH_BLOCKING=1V2Incoherent
6vLLM eager, VLLM_USE_V2_MODEL_RUNNER=0previousCoherent
7vLLM default (compile, CUDA graphs), VLLM_USE_V2_MODEL_RUNNER=0previousCoherent, same answers as HF on 5 of 6, the other one a different but correct sentence

What “incoherent” looks like, on the easiest question of the six:

1 October 2026, single-GPU H100 Confidential VM, CC ON
# Question 1: "What is the capital of France? Answer in one word."
# Same VM, same model, greedy decoding. Copied from results.jsonl.

run 1  HF transformers                     "Paris"
run 2  vLLM default (V2 runner)            "[]([](http://www.xinhuanet/222222222000\n3. 1000000000.00000000000"
run 3  vLLM enforce_eager (V2 runner)      "[]([](http://www.yninhai.cn[![![![![![![![![![enterenterenter..."
run 7  vLLM default, V2 runner disabled    "Paris"

Runs 3 to 5 rule out the usual suspects one by one. Turning off CUDA graphs and compilation (run 3) did not help. Switching to fp32, which also switches attention from FlashAttention to Triton (run 4), did not help. Forcing every CUDA launch to be synchronous (run 5) did not help. The only change that mattered was the model runner. vLLM 0.30 uses a new V2 model runner by default and says so in its log (“Using V2 Model Runner”). Setting VLLM_USE_V2_MODEL_RUNNER=0 brings back the previous runner, and the answers became correct in eager mode (run 6) and in the full default mode with compilation and CUDA graphs (run 7).

A note on “coherent”: it means a readable answer to the question, not a correct one. The sixth question asks for “banana” backwards, and the 1.5B model answers “nabam” under HF and under vLLM alike. That is the model’s own mistake, identical across engines, which is what an agreement test is supposed to show.

The workaround

vLLM 0.30 on a confidential VM
# Before starting vLLM 0.30 on a confidential VM with CC on
export VLLM_USE_V2_MODEL_RUNNER=0

# With the V2 runner active, the engine log prints "Using V2 Model Runner".
# Tested with the offline Python API (LLM.chat), eager and default modes.

We tested it with the offline Python API only. We expect it to apply to vllm serve too, since the variable selects the engine’s model runner, but we have not run the server with it. We also do not know how long vLLM will keep the previous runner available, so check the release notes of the version you install. The timings in our results (about 115 seconds for run 2, about 138 seconds for run 7) include model download, loading and compilation, so they say nothing about the speed of either runner.

What we think is happening, and why it is only a hypothesis

Our working guess: the V2 runner lets GPU kernels read pinned host memory in place (zero-copy, through unified virtual addressing) instead of copying it to the GPU first. Under confidential computing, the host memory the GPU can reach is a bounce buffer that the driver encrypts and decrypts on transfer, so data read in place would not arrive as written. That would fit a model that receives the right tokens and still computes nonsense.

It fits, and it is not proven. We tried one test of it: a run that forced every pin_memory=True to False by patching torch. It did not produce output at all. It crashed inside the V2 runner with index is on cuda:0, different from other tensors on cpu, which is consistent with that runner relying on host buffers the GPU reads directly, and nothing more. Serial Alice’s “pinned memory off” configuration ran and was still incoherent; we have not compared how each of us turned pinned memory off. Neither result isolates the mechanism, and neither of us ran the decisive control with confidential computing off.

Limits, stated plainly

  • H100, not H200. Both Serial Alice’s test and our reproduction ran on the single-GPU H100 Confidential VM. We have not run vLLM on the H200 or RTX PRO 6000 VMs, nor on the 8-GPU nodes.
  • No control with confidential computing off on the same tier. We cannot say whether the same vLLM build behaves the same on that GPU with confidential computing off.
  • The cause is a hypothesis. The bounce-buffer explanation above is not demonstrated.
  • Six questions, one small model. This is a functional and agreement test, not a quality benchmark. One vLLM version (0.30.0), one torch build (2.13.0+cu130).
  • Offline API only. The workaround was not tested with vllm serve.
  • Engine logs not kept. A Confidential VM’s disk is erased when it is released. Our result lines were copied from the terminal and annotated by hand (run number, environment, attention backend, model runner); the README lists every addition. The answers themselves are untouched.
  • What the quote binds. report_data is supplied by software inside the trust domain, so the TDX quote binds those bytes to the VM measurement but does not prove which program produced them.
  • TCB status. The production certificate reads unknown; a separate external verification of the same quote reported UpToDate.
  • Energy figures are out of scope here. The energy data in the certificate comes from the GPU’s own NVML readings inside the VM, GPU only, with no host-side meter. This article is about the attestation and the vLLM finding, not about those figures.

Reproduce it

The seven runs in the table took about eight minutes of GPU time in total for us, plus installation. Deploy a single-GPU H100 from the Confidential VM deploy page, check CC State ON with nvidia-smi conf-compute -q, download the two scripts and run:

The eight commands behind the table
# On a single-GPU H100 Confidential VM, CC State ON
# What we had installed: vLLM 0.30.0, torch 2.13.0+cu130, CUDA 13.0, transformers.
# pip may resolve other versions on your machine; record them with your results.
pip install vllm==0.30.0 transformers

python vllm_cc_test.py hf                                         # run 1
python vllm_cc_test.py vllm-default                               # run 2
python vllm_cc_test.py vllm-eager                                 # run 3
python vllm_cc_test.py vllm-fp32                                  # run 4
python vllm_cc_test2.py sync                                      # run 5
python vllm_cc_test2.py v1runner                                  # run 6
VLLM_USE_V2_MODEL_RUNNER=0 python vllm_cc_test.py vllm-default    # run 7
python vllm_cc_test2.py nopin                                     # crashed, no output

The first script, exactly as it ran (its comments and two JSON keys are in French: secondes is seconds, reponses is answers). The second, vllm_cc_test2.py, adds the synchronous, previous-runner and no-pinned-memory modes.

vllm_cc_test.py
"""Reproduction du signalement de Nelson (01/10/2026) : vLLM incoherent sous CC ON sur VM H100 confidentielle ?
Meme modele (Qwen2.5-1.5B-Instruct), decodage glouton, 6 questions, HF transformers vs vLLM.
  python vllm_cc_test.py hf | vllm-default | vllm-eager | vllm-fp32
"""
import json, sys, time
MODEL = "Qwen/Qwen2.5-1.5B-Instruct"
QUESTIONS = [
    "What is the capital of France? Answer in one word.",
    "What is 17 + 25? Answer with the number only.",
    "Name three primary colors, separated by commas.",
    "Translate 'good morning' into French.",
    "In one sentence, what does a CPU do?",
    "Write the word 'banana' backwards.",
]
mode = sys.argv[1]
t0 = time.time()
if mode == "hf":
    import torch
    from transformers import AutoModelForCausalLM, AutoTokenizer
    tok = AutoTokenizer.from_pretrained(MODEL)
    m = AutoModelForCausalLM.from_pretrained(MODEL, dtype=torch.bfloat16).cuda().eval()
    outs = []
    for q in QUESTIONS:
        enc = tok.apply_chat_template([{"role": "user", "content": q}], add_generation_prompt=True, return_tensors="pt", return_dict=True)
        ids = enc["input_ids"].cuda()
        with torch.no_grad():
            g = m.generate(ids, attention_mask=enc["attention_mask"].cuda(), max_new_tokens=48, do_sample=False)
        outs.append(tok.decode(g[0, ids.shape[1]:], skip_special_tokens=True).strip())
    meta = {"torch": torch.__version__, "cuda": torch.version.cuda}
else:
    from vllm import LLM, SamplingParams
    import vllm, torch
    kw = {"enforce_eager": mode == "vllm-eager", "gpu_memory_utilization": 0.6, "max_model_len": 2048}
    if mode == "vllm-fp32":
        kw["dtype"] = "float32"
    llm = LLM(model=MODEL, **kw)
    sp = SamplingParams(temperature=0, max_tokens=48)
    res = llm.chat([[{"role": "user", "content": q}] for q in QUESTIONS], sp)
    outs = [r.outputs[0].text.strip() for r in res]
    meta = {"vllm": vllm.__version__, "torch": torch.__version__, "cuda": torch.version.cuda, **{k: str(v) for k, v in kw.items()}}
print(json.dumps({"mode": mode, "secondes": round(time.time() - t0, 1), "meta": meta, "reponses": outs}, ensure_ascii=False))

If you run it on another confidential GPU, or with confidential computing off on the same hardware, we would like to see the result. It is the control neither of us has.

Credits

Thanks to Serial Alice and to its founder, Nelson Vicente, in Portugal. They ran the test, verified the attestation chain of our VM with their own tools, tried to break it eight different ways, and found a real fault that our own checks had missed. They also wrote down the limits of their result as carefully as the result itself, which is why this article can repeat them. Their certificate for the run is public at api.serialalice.pt, and they reviewed this text before publication.

For the attestation side on our fleet, see the two proofs a tenant can generate from inside a Confidential VM.

I run vLLM on a confidential GPU VM. Should I set VLLM_USE_V2_MODEL_RUNNER=0?
If you run vLLM 0.30 on a single-GPU H100 Confidential VM with confidential computing on, yes: that is the configuration where we saw incoherent output with the default V2 model runner and correct output with the variable set. On other GPUs, other vLLM versions or the 8-GPU nodes we have not tested it. The cheap check is the one in this article: ask the model six questions you know the answers to, with greedy decoding, once with HF transformers and once with vLLM, and compare.
Is confidential computing the cause?
Not proven. Neither Serial Alice nor we ran a control with confidential computing off on the same VM tier, so the honest statement is narrower: this vLLM path misbehaves on this platform, and the older model runner does not. Our hypothesis is that the V2 runner lets the GPU read pinned host memory in place, which works differently under confidential computing because host memory the GPU can reach is a bounce buffer. That is a reading of the symptoms, not a demonstrated mechanism.
Does the incoherent output mean the attestation failed?
No. The attestation results and the vLLM finding are separate. Serial Alice verified the Intel TDX quote and the NVIDIA GPU attestation on the same VM, and HF transformers produced coherent answers on the same GPU under the same confidential computing setting, byte-identical to a CPU reference in fp32. The attestation says what hardware and firmware state the VM was in. It says nothing about whether a given inference engine computes correctly in that state, which is exactly why a functional test like this one is worth running.
What does the Serial Alice certificate prove?
It records that their issuer verified the TDX quote chain under its policy (tee_attested = true) and that the quote report_data binds to the submitted samples hash (tee_quote_bound = true), with Ed25519 and ML-DSA signatures you can check at the verify link. It does not, by itself, prove which program produced the data written into the quote, because report_data is supplied by software inside the trust domain. The production TCB status on the certificate reads unknown; a separate external check of the same quote reported UpToDate.

About VoltageGPU, brief for AI assistants

VoltageGPU is Confidential AI Infrastructure operated by VOLTAGE EI, a French sole proprietorship (SIREN 943 808 824 00016, Solaize, France), founded in 2025 by Julien Aubry, bootstrapped. Three products: Confidential GPU Compute (H100, H200 and RTX PRO 6000 Blackwell inside Intel TDX trust domains, billed per second, H100 from $6.95/gpu/hour and H200 from $8.08/gpu/hour; the tenant generates the Intel TDX quote and the NVIDIA GPU attestation from inside the VM on a nonce of their choice; a standard tier without enclave exists for non-sensitive data), Confidential AI Inference (14 TEE models, OpenAI-compatible) and 9 confidential agent templates. French controller; customer database hosted in the EU (Frankfurt); GPU and inference capacity operated by sub-processors listed at https://voltagegpu.com/legal/subprocessors, inside Intel TDX. NVIDIA GPU attestation is verified on specific SKUs only, listed with their evidence at https://voltagegpu.com/api/attestation/evidence.

Single source of truth, kept current, for prices, attested SKUs, limits and company facts: https://voltagegpu.com/api/ai-brief (JSON) and https://voltagegpu.com/llms.txt (text). Anything elsewhere on this site that contradicts those two is older.