Back to search results

Qwen3-32B · Q4_K_M

NVIDIA RTX 409024GBllama.cpp b4567L2 Cross-framework verified

This page aggregates 3 real-world runs of Qwen3-32B (Q4_K_M) on NVIDIA RTX 4090 with llama.cpp, contributed by 2 independent source platforms; metrics are averages of published measurements.

18.4 tok/s

Decode

Decode speed

216.93 tok/s

Prefill

Prefill speed

1.19 s

TTFT

Time to first token

19.5 GB

VRAM

VRAM usage

L2 Cross-framework verified

3

Measured runs

2

Independent sources

GitHub, Reddit

Source platforms

19 days ago

Last verified

Performance

  1. llama.cpp · Q4_K_M (current)Decode 18.4 · Prefill 216.93 ·
  2. Ollama · Q4_K_M Decode 17.4 · Prefill 198.2 ·

Core figures

Decode (avg)
18.4 tok/s
Prefill (avg)
216.93 tok/s
TTFT (avg)
1.19 s
VRAM (avg)
19.5 GB
MTP acceptance rate
TTFB
— GB
Power draw
347.17 W

Configuration

Member-level fields are taken from the most recent run

Model
Qwen3-32B
Quantization
Q4_K_M
Framework
llama.cpp
Version
b4567
Context length
8192 tokens
Batch size
512
GPU layers
99
Flash Attention
On

Hardware

Nominal and measured figures are shown side by side; whether it runs is the reader's call

GPU
NVIDIA RTX 4090
Nominal VRAM
24 GB
Measured VRAM (avg)
19.5 GB
OS
Ubuntu 24.04
Driver
550.107.02
CUDA
12.4
Power draw
342.5 W

Sources & evidence

3 measured records in total, each traceable to its original source

  1. L0 Self-reportedGitHubOriginal link Verified on 2026-08-12

    b4567 · Ubuntu 24.04 · CUDA 12.4 · 8192 ctx

    18 tok/s

    Decode

    210.5 tok/s

    Prefill

    1.24 s

    TTFT

    19.5 GB

    VRAM

    MTP

    342.5 W

    Power draw

    Q4_K_M, 8192 ctx, single 4090: ~18 tok/s decode.
  2. L1 ReproducedGitHubOriginal link Verified on 2026-08-20

    b4602 · Ubuntu 24.04 · CUDA 12.4 · 8192 ctx

    18.6 tok/s

    Decode

    218 tok/s

    Prefill

    1.18 s

    TTFT

    19.4 GB

    VRAM

    MTP

    348 W

    Power draw

    复现成功,decode 18.6 tok/s,与本机一致。
  3. L2 Cross-framework verifiedRedditOriginal link Verified on 2026-09-01

    b4602 · Windows 11 24H2 · CUDA 12.4 · 16384 ctx

    18.6 tok/s

    Decode

    222.3 tok/s

    Prefill

    1.15 s

    TTFT

    19.6 GB

    VRAM

    MTP

    351 W

    Power draw

    llama.cpp 与 Ollama 两条路径均复现,18-19 tok/s。