4x NVIDIA Tesla P40

GPU

1

Run records

1

Models covered

1

Frameworks covered

L1 Reproduced × 1

Evidence mix

Typical measured performance

Published benchmark averages per model × framework × quantization on this hardware.

  1. llama.cpp · UD-Q3_K_XL Decode 21.6 · Prefill 330 ·

Key specs

Vendor
NVIDIA
VRAM
96 GB

Run records

1 configurations

ModelFrameworkQuantizationDecodeGen timeSamples
Qwen3.8-Flash-Nextllama.cppUD-Q3_K_XL21.6 tok/s—1