Back to search results

Qwen3.8-27B · Q6_K

AMD Radeon RX 7900 GRE + NVIDIA GTX 1080Ti27GBllama.cpp 10453L0 Self-reported

This page aggregates 1 real-world runs of Qwen3.8-27B (Q6_K) on AMD Radeon RX 7900 GRE + NVIDIA GTX 1080Ti with llama.cpp, contributed by 1 independent source platforms; metrics are averages of published measurements.

12.53 tok/s

Decode

Decode speed

116.21 tok/s

Prefill

Prefill speed

s

TTFT

Time to first token

23 GB

VRAM

VRAM usage

L0 Self-reported

1

Measured runs

1

Independent sources

Reddit

Source platforms

Today

Last verified

Performance

  1. llama.cpp · Q6_K (current)Decode 12.53 · Prefill 116.21 ·
  2. llama.cpp · Q4_K_M Decode 11.13 · Prefill 71.46 ·

Core figures

Decode (avg)
12.53 tok/s
Prefill (avg)
116.21 tok/s
TTFT (avg)
— s
VRAM (avg)
23 GB
MTP acceptance rate
TTFB
— GB
Power draw
— W

Configuration

Member-level fields are taken from the most recent run

Model
Qwen3.8-27B
Quantization
Q6_K
Framework
llama.cpp
Version
10453

Hardware

Nominal and measured figures are shown side by side; whether it runs is the reader's call

GPU
AMD Radeon RX 7900 GRE + NVIDIA GTX 1080Ti
Nominal VRAM
27 GB
Measured VRAM (avg)
23 GB
OS
Ubuntu

Sources & evidence

1 measured records in total, each traceable to its original source

  1. L0 Self-reportedRedditOriginal link Verified on 2026-09-20

    10453 · Ubuntu

    12.53 tok/s

    Decode

    116.21 tok/s

    Prefill

    s

    TTFT

    23 GB

    VRAM

    MTP

    W

    Power draw

    Radeon plus GTX 1080Ti Q6\_K using RPC |qwen35 27B Q6\_K|21.30 GiB|27.32 B|pp512|116.21 ± 0.55| |qwen35 27B Q6\_K|21.30 GiB|27.32 B|tg128|12.53 ± 0.02| build: 3cb7ffb1a (10453) Almost 23gb VRAM used about 2gb for combined desktop resources