Qwen3.6-27B-FP8 on One RTX 6000 Ada: Fast TTFT, 668 tok/s Peak Throughput [Benchmark]
Qwen/Qwen3.6-27B-FP8 served through vLLM on a single RTX 6000 Ada is a strong and practical configuration for 8K-context chat serving.
Jul 5, 202613 min read206
![Qwen3.6-27B-FP8 on One RTX 6000 Ada: Fast TTFT, 668 tok/s Peak Throughput [Benchmark]](https://cdn.hashnode.com/uploads/covers/6a22b1a041d5b05f16273b50/8fd36dcb-515c-4f77-8071-9a1aedc1c2ed.png)
Search for a command to run...
Articles tagged with #vllm
Qwen/Qwen3.6-27B-FP8 served through vLLM on a single RTX 6000 Ada is a strong and practical configuration for 8K-context chat serving.
![Qwen3.6-27B-FP8 on One RTX 6000 Ada: Fast TTFT, 668 tok/s Peak Throughput [Benchmark]](https://cdn.hashnode.com/uploads/covers/6a22b1a041d5b05f16273b50/8fd36dcb-515c-4f77-8071-9a1aedc1c2ed.png)
Throughput, latency, and queue depth for Gemma-4 31B served on vLLM under progressive load, from 12 to 24 concurrency The numbers that matter: 1.17k tok/s peak, ~0.7s median TTFT, and tail latency as the one thing to watch.
