Key Takeaways
- Serial Alice found it, not us. On 1 October 2026 they ran a full attestation test on our single-GPU H100 Confidential VM and reported that vLLM 0.30 gives incoherent answers with confidential computing on, while HF transformers on the same GPU is correct.
- We reproduced it the same day, in seven runs. Default, eager, fp32 and synchronous launches all failed. The fault followed the new V2 model runner, not attention, CUDA graphs or precision.
- The workaround is one variable:
VLLM_USE_V2_MODEL_RUNNER=0. With it, vLLM answers correctly on that VM, in eager mode and in the default compiled mode. - The cause is not proven. No control with confidential computing off on the same tier, one small model, six questions, engine logs not kept. Every limit is listed below.
Our reproduction files
The two scripts we ran and the outputs of the seven runs, with SHA-256 checksums. The README says how the result lines were copied from the terminal and what was added by hand. Serial Alice’s certificate is public at its own verify link, linked below.
Attestation write-ups, ours included, usually stop at the attestation: the quote verifies, the GPU token verifies, done. Serial Alice went one step further on our hardware. They checked that a language model actually gives the right answers inside the confidential VM, and bound those answers to the attestation. That second check is how they found a problem we had not seen: vLLM, a widely used open-source inference engine, in its current default configuration, returned nonsense on that machine.
This is a joint write-up. The test, the finding and the certificate are Serial Alice’s. The reproduction and the workaround are ours. Both sides state what they did not prove.
What Serial Alice set out to test
Serial Alice is a Portuguese company, founded by Nelson Vicente, that issues signed, independently verifiable energy certificates for compute jobs. Their goal on our platform was specific: produce a coherent LLM answer with confidential computing on, and bind it to verifiable evidence. In practice that means an Intel TDX quote whose 64 bytes of report_data commit to a challenge generated before inference and to a manifest linking the answers and the energy samples, plus an NVIDIA GPU attestation carrying the same challenge as its nonce.
They ran it on 1 October 2026. The single-GPU H200 Confidential VM was sold out that day, so the test ran on the single-GPU H100 Confidential VM. Everything below is an H100 result, not an H200 one.
What Serial Alice verified on our VM, independently
On the machine itself:
nvidia-smi conf-compute -qreported CC State ON, the CPU was Intel TDX, and both/dev/tdx_guestand configfs-tsm were present.- Qwen2.5-1.5B-Instruct under HF transformers, greedy decoding, gave coherent answers. In fp32, all six answers were byte-identical to a CPU-only reference run.
- The NVIDIA GPU attestation used Serial Alice’s challenge as its nonce.
Then off the VM, with their own tools:
- The TDX quote chains to the Intel root.
- The full 64-byte report_data matches what they expected, challenge included.
- The NVIDIA attestation tokens carry valid ES384 signatures against NVIDIA’s published keys,
eat_nonceequals their challenge, the overall result is true, the measurements match, secure boot is on and debug is disabled. - They tried eight tampering and replay cases: an altered answer, altered energy data or an altered model; a quote or an NVIDIA token swapped between runs; a wrong challenge. All eight were rejected.
The result is a public certificate, sa-b8f6f5aab6774c82bae5108ae603a64d, with tee_attested = true (the issuer verified the TDX quote chain and accepted it under its policy, with the quote’s MRTD on its allowlist) and tee_quote_bound = true (the quote’s report_data binds to the submitted samples hash). It is signed with Ed25519 and with the post-quantum ML-DSA scheme, and the verify link checks both signatures.
tcb_status: unknown, with a warning that platform TCB currency is unproven. That is Serial Alice’s production issuer, which could not establish the TCB status at issuance; the exact cause was not recorded. A separate external verification of the same quote reported UpToDate. Both are true, and Serial Alice has since added a retry and recorded failure reasons for future certificates.The finding: vLLM output is incoherent with confidential computing on
With HF transformers correct on the GPU, Serial Alice ran the same test through vLLM 0.30. The output was incoherent, and it stayed incoherent in every configuration they tried: default, eager, fp32 with Triton attention, FlashAttention, pinned memory off, custom ops off. They also checked the two obvious suspects. The weights on the GPU matched the file, and vLLM received the correct prompt tokens. The forward pass was still wrong.
Their conclusion was careful, and we keep it: without a control run with confidential computing off on the same VM tier, they did not claim that confidential computing is the cause, only that this vLLM path misbehaves on this platform.
Our reproduction, the same day
We rented the same tier, a single-GPU H100 Confidential VM, and confirmed CC status: ON with nvidia-smi conf-compute -f right after the last run. Image ubuntu-24-04-lts-595, torch 2.13.0+cu130, CUDA 13.0, vLLM 0.30.0. Same model as Serial Alice, Qwen/Qwen2.5-1.5B-Instruct, greedy decoding (temperature 0), 48 new tokens, six short questions with known answers in every run.
| Run | Configuration | Model runner | Answers |
|---|---|---|---|
| 1 | HF transformers, bf16 | n/a | Coherent |
| 2 | vLLM default (torch.compile, CUDA graphs, FlashAttention 3) | V2 | Incoherent |
| 3 | vLLM enforce_eager | V2 | Incoherent |
| 4 | vLLM fp32 (Triton attention) | V2 | Incoherent |
| 5 | vLLM eager, CUDA_LAUNCH_BLOCKING=1 | V2 | Incoherent |
| 6 | vLLM eager, VLLM_USE_V2_MODEL_RUNNER=0 | previous | Coherent |
| 7 | vLLM default (compile, CUDA graphs), VLLM_USE_V2_MODEL_RUNNER=0 | previous | Coherent, same answers as HF on 5 of 6, the other one a different but correct sentence |
What “incoherent” looks like, on the easiest question of the six:
# Question 1: "What is the capital of France? Answer in one word." # Same VM, same model, greedy decoding. Copied from results.jsonl. run 1 HF transformers "Paris" run 2 vLLM default (V2 runner) "[]([](http://www.xinhuanet/222222222000\n3. 1000000000.00000000000" run 3 vLLM enforce_eager (V2 runner) "[]([](http://www.yninhai.cn[![![![![![![![![![enterenterenter..." run 7 vLLM default, V2 runner disabled "Paris"
Runs 3 to 5 rule out the usual suspects one by one. Turning off CUDA graphs and compilation (run 3) did not help. Switching to fp32, which also switches attention from FlashAttention to Triton (run 4), did not help. Forcing every CUDA launch to be synchronous (run 5) did not help. The only change that mattered was the model runner. vLLM 0.30 uses a new V2 model runner by default and says so in its log (“Using V2 Model Runner”). Setting VLLM_USE_V2_MODEL_RUNNER=0 brings back the previous runner, and the answers became correct in eager mode (run 6) and in the full default mode with compilation and CUDA graphs (run 7).
A note on “coherent”: it means a readable answer to the question, not a correct one. The sixth question asks for “banana” backwards, and the 1.5B model answers “nabam” under HF and under vLLM alike. That is the model’s own mistake, identical across engines, which is what an agreement test is supposed to show.
The workaround
# Before starting vLLM 0.30 on a confidential VM with CC on export VLLM_USE_V2_MODEL_RUNNER=0 # With the V2 runner active, the engine log prints "Using V2 Model Runner". # Tested with the offline Python API (LLM.chat), eager and default modes.
We tested it with the offline Python API only. We expect it to apply to vllm serve too, since the variable selects the engine’s model runner, but we have not run the server with it. We also do not know how long vLLM will keep the previous runner available, so check the release notes of the version you install. The timings in our results (about 115 seconds for run 2, about 138 seconds for run 7) include model download, loading and compilation, so they say nothing about the speed of either runner.
What we think is happening, and why it is only a hypothesis
Our working guess: the V2 runner lets GPU kernels read pinned host memory in place (zero-copy, through unified virtual addressing) instead of copying it to the GPU first. Under confidential computing, the host memory the GPU can reach is a bounce buffer that the driver encrypts and decrypts on transfer, so data read in place would not arrive as written. That would fit a model that receives the right tokens and still computes nonsense.
It fits, and it is not proven. We tried one test of it: a run that forced every pin_memory=True to False by patching torch. It did not produce output at all. It crashed inside the V2 runner with index is on cuda:0, different from other tensors on cpu, which is consistent with that runner relying on host buffers the GPU reads directly, and nothing more. Serial Alice’s “pinned memory off” configuration ran and was still incoherent; we have not compared how each of us turned pinned memory off. Neither result isolates the mechanism, and neither of us ran the decisive control with confidential computing off.
Limits, stated plainly
- H100, not H200. Both Serial Alice’s test and our reproduction ran on the single-GPU H100 Confidential VM. We have not run vLLM on the H200 or RTX PRO 6000 VMs, nor on the 8-GPU nodes.
- No control with confidential computing off on the same tier. We cannot say whether the same vLLM build behaves the same on that GPU with confidential computing off.
- The cause is a hypothesis. The bounce-buffer explanation above is not demonstrated.
- Six questions, one small model. This is a functional and agreement test, not a quality benchmark. One vLLM version (0.30.0), one torch build (2.13.0+cu130).
- Offline API only. The workaround was not tested with
vllm serve. - Engine logs not kept. A Confidential VM’s disk is erased when it is released. Our result lines were copied from the terminal and annotated by hand (run number, environment, attention backend, model runner); the README lists every addition. The answers themselves are untouched.
- What the quote binds. report_data is supplied by software inside the trust domain, so the TDX quote binds those bytes to the VM measurement but does not prove which program produced them.
- TCB status. The production certificate reads unknown; a separate external verification of the same quote reported UpToDate.
- Energy figures are out of scope here. The energy data in the certificate comes from the GPU’s own NVML readings inside the VM, GPU only, with no host-side meter. This article is about the attestation and the vLLM finding, not about those figures.
Reproduce it
The seven runs in the table took about eight minutes of GPU time in total for us, plus installation. Deploy a single-GPU H100 from the Confidential VM deploy page, check CC State ON with nvidia-smi conf-compute -q, download the two scripts and run:
# On a single-GPU H100 Confidential VM, CC State ON # What we had installed: vLLM 0.30.0, torch 2.13.0+cu130, CUDA 13.0, transformers. # pip may resolve other versions on your machine; record them with your results. pip install vllm==0.30.0 transformers python vllm_cc_test.py hf # run 1 python vllm_cc_test.py vllm-default # run 2 python vllm_cc_test.py vllm-eager # run 3 python vllm_cc_test.py vllm-fp32 # run 4 python vllm_cc_test2.py sync # run 5 python vllm_cc_test2.py v1runner # run 6 VLLM_USE_V2_MODEL_RUNNER=0 python vllm_cc_test.py vllm-default # run 7 python vllm_cc_test2.py nopin # crashed, no output
The first script, exactly as it ran (its comments and two JSON keys are in French: secondes is seconds, reponses is answers). The second, vllm_cc_test2.py, adds the synchronous, previous-runner and no-pinned-memory modes.
"""Reproduction du signalement de Nelson (01/10/2026) : vLLM incoherent sous CC ON sur VM H100 confidentielle ?
Meme modele (Qwen2.5-1.5B-Instruct), decodage glouton, 6 questions, HF transformers vs vLLM.
python vllm_cc_test.py hf | vllm-default | vllm-eager | vllm-fp32
"""
import json, sys, time
MODEL = "Qwen/Qwen2.5-1.5B-Instruct"
QUESTIONS = [
"What is the capital of France? Answer in one word.",
"What is 17 + 25? Answer with the number only.",
"Name three primary colors, separated by commas.",
"Translate 'good morning' into French.",
"In one sentence, what does a CPU do?",
"Write the word 'banana' backwards.",
]
mode = sys.argv[1]
t0 = time.time()
if mode == "hf":
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained(MODEL)
m = AutoModelForCausalLM.from_pretrained(MODEL, dtype=torch.bfloat16).cuda().eval()
outs = []
for q in QUESTIONS:
enc = tok.apply_chat_template([{"role": "user", "content": q}], add_generation_prompt=True, return_tensors="pt", return_dict=True)
ids = enc["input_ids"].cuda()
with torch.no_grad():
g = m.generate(ids, attention_mask=enc["attention_mask"].cuda(), max_new_tokens=48, do_sample=False)
outs.append(tok.decode(g[0, ids.shape[1]:], skip_special_tokens=True).strip())
meta = {"torch": torch.__version__, "cuda": torch.version.cuda}
else:
from vllm import LLM, SamplingParams
import vllm, torch
kw = {"enforce_eager": mode == "vllm-eager", "gpu_memory_utilization": 0.6, "max_model_len": 2048}
if mode == "vllm-fp32":
kw["dtype"] = "float32"
llm = LLM(model=MODEL, **kw)
sp = SamplingParams(temperature=0, max_tokens=48)
res = llm.chat([[{"role": "user", "content": q}] for q in QUESTIONS], sp)
outs = [r.outputs[0].text.strip() for r in res]
meta = {"vllm": vllm.__version__, "torch": torch.__version__, "cuda": torch.version.cuda, **{k: str(v) for k, v in kw.items()}}
print(json.dumps({"mode": mode, "secondes": round(time.time() - t0, 1), "meta": meta, "reponses": outs}, ensure_ascii=False))If you run it on another confidential GPU, or with confidential computing off on the same hardware, we would like to see the result. It is the control neither of us has.
Credits
Thanks to Serial Alice and to its founder, Nelson Vicente, in Portugal. They ran the test, verified the attestation chain of our VM with their own tools, tried to break it eight different ways, and found a real fault that our own checks had missed. They also wrote down the limits of their result as carefully as the result itself, which is why this article can repeat them. Their certificate for the run is public at api.serialalice.pt, and they reviewed this text before publication.
For the attestation side on our fleet, see the two proofs a tenant can generate from inside a Confidential VM.