⚡ LLM Tokens / sec

latency · throughput · batch scaling

① Latency → TPS

Tokens/sec160.0
Inter‑token latency 5.96 ms

② Throughput estimator

Memory‑bandwidth bound TPS 957.1
⚡ compute-bound if small batch

③ Batch scaling

Estimated batch TPS 384.0
Efficiency 40.0%

📊 TPS interpretation

Tokens/secRating
> 100⚡ Excellent
30 – 100✅ Good
10 – 30⚠️ Acceptable
< 10🐢 Slow
Cost per 1M tokens (GPU $/h)
@ 50 TPS: $0.56
@ 150 TPS: $0.19
@ 500 TPS: $0.06

assumes $10/GPU-hr · cost = (1e6 / TPS) / 3600 * $10

Recommended by our team

BeLikeNative.com

The #1 AI writing tool for freelancers — perfect grammar in any language, instantly.