Qwen3.6-27B-FP8 on One RTX 6000 Ada: Fast TTFT, 668 tok/s Peak Throughput [Benchmark]
Qwen/Qwen3.6-27B-FP8 served through vLLM on a single RTX 6000 Ada is a strong and practical configuration for 8K-context chat serving.
Jul 5, 202613 min read171
![Qwen3.6-27B-FP8 on One RTX 6000 Ada: Fast TTFT, 668 tok/s Peak Throughput [Benchmark]](/_next/image?url=https%3A%2F%2Fcdn.hashnode.com%2Fuploads%2Fcovers%2F6a22b1a041d5b05f16273b50%2F8fd36dcb-515c-4f77-8071-9a1aedc1c2ed.png&w=3840&q=75)
Search for a command to run...
Articles tagged with #ai-tools
Qwen/Qwen3.6-27B-FP8 served through vLLM on a single RTX 6000 Ada is a strong and practical configuration for 8K-context chat serving.
![Qwen3.6-27B-FP8 on One RTX 6000 Ada: Fast TTFT, 668 tok/s Peak Throughput [Benchmark]](/_next/image?url=https%3A%2F%2Fcdn.hashnode.com%2Fuploads%2Fcovers%2F6a22b1a041d5b05f16273b50%2F8fd36dcb-515c-4f77-8071-9a1aedc1c2ed.png&w=3840&q=75)
Throughput, latency, and queue depth for Gemma-4 31B served on vLLM under progressive load, from 12 to 24 concurrency The numbers that matter: 1.17k tok/s peak, ~0.7s median TTFT, and tail latency as the one thing to watch.

Open-source AI has crossed an important line. The question is no longer whether open models are good enough to power serious products. The question is how quickly teams can deploy them privately, rel
