vLLM under confidential computing: reproduction and workaround VoltageGPU, 1 October 2026 Context On 1 October 2026 Serial Alice (serialalice.pt) reported that vLLM 0.30 produced incoherent output on our single-GPU H100 Confidential VM with confidential computing ON, while HF transformers on the same GPU was correct. These files are our reproduction, run the same day. Serial Alice's own evidence package is theirs and is not in this folder. Machine Single-GPU H100 Confidential VM (Intel TDX guest), the same tier Serial Alice tested. nvidia-smi conf-compute -f: "CC status: ON" (read on the VM right after run 7). Image ubuntu-24-04-lts-595. torch 2.13.0+cu130, CUDA 13.0, vLLM 0.30.0. Test Qwen/Qwen2.5-1.5B-Instruct, greedy decoding (temperature 0), 48 new tokens, the same 6 short questions in every run. Scripts: vllm_cc_test.py and vllm_cc_test2.py. The comments and some JSON keys in the scripts are in French ("secondes" = seconds, "reponses" = answers). They are published exactly as they were run. Results (results.jsonl, one line per run) 1 HF transformers, bf16 coherent 2 vLLM default (torch.compile + CUDA graphs, FA3) incoherent 3 vLLM enforce_eager incoherent 4 vLLM fp32 (Triton attention) incoherent 5 vLLM eager + CUDA_LAUNCH_BLOCKING=1 incoherent 6 vLLM eager + VLLM_USE_V2_MODEL_RUNNER=0 coherent 7 vLLM default + VLLM_USE_V2_MODEL_RUNNER=0 coherent, same answers as HF on 5 of 6, the other one (the CPU question) is a different but correct sentence "Coherent" means a readable answer to the question, not a correct one. The sixth question ("banana" backwards) gets "nabam" from HF and from vLLM alike: the 1.5B model is wrong on its own, in every coherent run. A run that forced every pin_memory=True to False (vllm_cc_test2.py nopin) crashed inside the V2 runner ("index is on cuda:0, different from other tensors on cpu") and printed no JSON, so it has no line in results.jsonl. We read it as consistent with the V2 runner relying on host buffers that the GPU reads directly, nothing more. Reading The incoherence follows the V2 model runner (the vLLM 0.30 default, logged as "Using V2 Model Runner"), not the attention backend, CUDA graphs, precision or launch ordering. Switching back to the previous runner fixes it on this VM. Our hypothesis, not proven: the V2 runner lets kernels read pinned host memory in place (zero-copy / UVA). Under CC, host memory the GPU can see is a bounce buffer, so data read that way does not arrive correctly. We did not run a CC-OFF control on the same tier either. How these files were made A Confidential VM's disk is erased when the VM is released. Each line of results.jsonl is the JSON the script printed, copied from the terminal, then annotated by hand: we added the "run" and "env" fields and the "attention_backend" and "model_runner" values (from the command line used and from the engine log as it ran), spelled out the mode labels of runs 5 to 7, and dropped the "pin_patches" counter that vllm_cc_test2.py prints (0 outside the nopin mode). The answers ("reponses") are untouched. The full engine logs were not kept. Timings include model download, load and compilation and are not performance figures. Anyone can re-run the two scripts in a few minutes to check. Commands python vllm_cc_test.py hf # run 1 python vllm_cc_test.py vllm-default # run 2 python vllm_cc_test.py vllm-eager # run 3 python vllm_cc_test.py vllm-fp32 # run 4 python vllm_cc_test2.py sync # run 5 python vllm_cc_test2.py v1runner # run 6 VLLM_USE_V2_MODEL_RUNNER=0 python vllm_cc_test.py vllm-default # run 7 python vllm_cc_test2.py nopin # crashed, no output Checksums in SHA256SUMS.