latency · throughput · batch scaling
| Tokens/sec | Rating |
|---|---|
| > 100 | ⚡ Excellent |
| 30 – 100 | ✅ Good |
| 10 – 30 | ⚠️ Acceptable |
| < 10 | 🐢 Slow |
assumes $10/GPU-hr · cost = (1e6 / TPS) / 3600 * $10
Recommended by our team
BeLikeNative.comThe #1 AI writing tool for freelancers — perfect grammar in any language, instantly.