Back to search results

Qwen3.8-27B · IQ3_S

Start a discussion
AMD Radeon RX 7900 XTX24GBllama.cpp b11160L0 Self-reported

This page aggregates 2 real-world runs of Qwen3.8-27B (IQ3_S) on AMD Radeon RX 7900 XTX with llama.cpp, contributed by 1 independent source platforms; metrics are averages of published measurements.

26.73 tok/s

Decode

Decode speed

487.5 tok/s

Prefill

Prefill speed

— s

TTFT

Time to first token

22.15 GB

VRAM

VRAM usage

L0 Self-reported

2

Measured runs

1

Independent sources

GitHub

Source platforms

5 days ago

Last verified

Performance

  1. llama.cpp · Q4_K_M Decode 124.1 · Prefill 156 ·
  2. llama.cpp · IQ3_S (current)Decode 26.73 · Prefill 487.5 ·

Core figures

Decode (avg)
26.73 tok/s
Prefill (avg)
487.5 tok/s
TTFT (avg)
— s
VRAM (avg)
22.15 GB
MTP acceptance rate
—
TTFB
— GB
Power draw
— W

Configuration

Member-level fields are taken from the most recent run

Model
Qwen3.8-27B
Quantization
IQ3_S
Framework
llama.cpp
Version
b11160
Context length
229376 tokens
Flash Attention
On

Hardware

Nominal and measured figures are shown side by side; whether it runs is the reader's call

GPU
AMD Radeon RX 7900 XTX
Nominal VRAM
24 GB
Measured VRAM (avg)
22.15 GB
OS
CachyOS (Arch)

Get started

Original model resources

Base model resources may not match this quantization. Use the verified deployment weights above when available.

Startup command

Command from the latest benchmark run
llama-server -m <gguf> -c 229376 -ctk q8_0 -ctv q8_0 -fa on -np 1 --ctx-checkpoints 4 -ngl all --spec-type draft-mtp --spec-draft-n-max 2

Sources & evidence

2 measured records in total, each traceable to its original source

  1. L0 Self-reportedGitHubOriginal link Verified on 2026-09-28

    b11160 · CachyOS (Arch) · 229376 ctx

    22.03 tok/s

    Decode

    422 tok/s

    Prefill

    — s

    TTFT

    22.7 GB

    VRAM

    —

    MTP

    — W

    Power draw

    | 224K (`-c 229376`) | 221,167 | 422 | 22.3 / 25.1 / 18.7 | 22.7 GB | none (0 evicted) |
  2. L0 Self-reportedGitHubOriginal link Verified on 2026-09-28

    b11160 · CachyOS (Arch) · 204800 ctx

    31.43 tok/s

    Decode

    553 tok/s

    Prefill

    — s

    TTFT

    21.6 GB

    VRAM

    —

    MTP

    — W

    Power draw

    | 200K (`-c 204800`, context only) | 127,983 | 553 | 30.6 / 36.0 / 27.7 | 21.6 GB | none |