It would be dishonest to write a comparison page against Together AI without acknowledging the categories where they are clearly the better tool. Catalogue breadth is the most obvious: Together serves 200+ open-source models including image generation (FLUX.1 dev, Stable Diffusion XL), video generation (Wan, Hunyuan), audio, embeddings, rerankers, and safety classifiers. VoltageGPU lists 16 TEE-attested text and code models, Qwen 3.x, gemma-class, DeepSeek-class, and frontier MoE variants, and does not currently ship a confidential image-or-video inference product. For any workload that needs FLUX, Wan, or Stable Diffusion at API speed, Together is the right answer and VoltageGPU is not in that category at all today.
On raw per-token economics Together wins outright on the tiny-model and frontier-output tiers. Llama 3.1 8B Instruct Turbo is $0.18 / $0.18 per million tokens on Together, we do not list an 8B TEE SKU, so for high-volume tiny-model inference where the workload is non-confidential, Together is the price-correct choice. More importantly, DeepSeek V3 is $1.25 / $1.25 flat on Together versus our Qwen3.5-397B-A17B-TEE frontier tier at $0.72 input but $4.33 output. The math is exact: on input we are 42% cheaper, on output Together is 3.5x cheaper. For an output-heavy frontier workload that does not need TEE, long-form generation, code synthesis at scale, agent-loop reasoning where output tokens dominate, Together wins the economic comparison and the buyer should choose Together.
Together also ships a dedicated-endpoint product starting at $3.99/hr for an H100 80GB that lets enterprise customers buy predictable throughput rather than burst per-token capacity. That is a category VoltageGPU has not built today on the inference side, our confidential GPU pods exist, but the "dedicated managed inference endpoint with auto-scaling" shape is a Together product feature we do not match. Turbo-tier latency on Llama 3.3 70B at 200–400 tokens per second is industry-leading and not a number we publish on our confidential serving infrastructure.
The honest summary is product-shape, not provider-quality: Together is the breadth-and-speed king on US multi-tenant open-weight inference and remains the right answer for unconstrained workloads on the broadest model catalogue. VoltageGPU is European confidential inference and is the right answer when the regulator, the client contract, or the threat model requires the prompt be hardware-sealed from the operator. A workload that picks the wrong-shape tool will be unhappy with both.