Back to search results

gemma-4-12B-it · IQ2_M

Start a discussion
NVIDIA RTX 509032GBvLLM v0.45L2 Cross-framework verified

This page aggregates 1 real-world runs of gemma-4-12B-it (IQ2_M) on NVIDIA RTX 5090 with vLLM, contributed by 1 independent source platforms; metrics are averages of published measurements.

288.51 tok/s

Decode

Decode speed

940.03 tok/s

Prefill

Prefill speed

1.35 s

TTFT

Time to first token

15.2 GB

VRAM

VRAM usage

L2 Cross-framework verified

1

Measured runs

1

Independent sources

Other

Source platforms

Today

Last verified

Performance

  1. llama.cpp · IQ2_XS Decode 253.43 · Prefill 894.02 ·
  2. llama.cpp · IQ3_S Decode 214.27 · Prefill 831.85 ·
  3. vLLM · IQ1_S Decode 298.32 · Prefill 1152.98 ·
  4. vLLM · IQ2_M (current)Decode 288.51 · Prefill 940.03 ·
  5. Ollama · IQ2_XS Decode 246.6 · Prefill 1083.78 ·
  6. vLLM · Q8_0 Decode 117.2 · Prefill 879 ·
  7. vLLM · Q5_K_M Decode 171.1 · Prefill 1283.3 ·
  8. vLLM · Q4_K_M Decode 208.7 · Prefill 1565.3 ·
  9. llama.cpp · Q8_0 Decode 107.5 · Prefill 806.3 ·
  10. llama.cpp · Q5_K_M Decode 157.1 · Prefill 1178.3 ·
  11. llama.cpp · Q4_K_M Decode 191.5 · Prefill 1436.3 ·

Core figures

Decode (avg)
288.51 tok/s
Prefill (avg)
940.03 tok/s
TTFT (avg)
1.35 s
VRAM (avg)
15.2 GB
MTP acceptance rate
—
TTFB
2.31 GB
Power draw
581.4 W

Configuration

Member-level fields are taken from the most recent run

Model
gemma-4-12B-it
Quantization
IQ2_M
Framework
vLLM
Version
v0.45
Context length
16384 tokens
Batch size
16
GPU layers
999
Flash Attention
Off

Hardware

Nominal and measured figures are shown side by side; whether it runs is the reader's call

GPU
NVIDIA RTX 5090
Nominal VRAM
32 GB
Measured VRAM (avg)
15.2 GB
OS
Windows 11 23H2
Driver
555.85
CUDA
12.8
Power draw
581.4 W

Sources & evidence

1 measured records in total, each traceable to its original source

  1. L2 Cross-framework verifiedOtherOriginal link Verified on 2026-10-10

    v0.45 · Windows 11 23H2 · CUDA 12.8 · 16384 ctx

    288.51 tok/s

    Decode

    940.03 tok/s

    Prefill

    1.35 s

    TTFT

    15.2 GB

    VRAM

    —

    MTP

    581.4 W

    Power draw

    L2 batch seed 2026-10-10: vLLM on hw_id=13 running model_id=65 at IQ2_M