4x NVIDIA Tesla P40
GPU1
Run records
1
Models covered
1
Frameworks covered
L1 Reproduced × 1
Evidence mix
Typical measured performance
Published benchmark averages per model × framework × quantization on this hardware.
- llama.cpp · UD-Q3_K_XL Decode 21.6 · Prefill 330 ·
Key specs
- Vendor
- NVIDIA
- VRAM
- 96 GB
Run records
1 configurations
| Model | Framework | Quantization | Decode | Gen time | Samples |
|---|---|---|---|---|---|
| Qwen3.8-Flash-Next | llama.cpp | UD-Q3_K_XL | 21.6 tok/s | — | 1 |