Back to search results

Qwen3.5-0.8B · Q6_K

Start a discussion
NVIDIA RTX 509032GBStrata v0.53L2 Cross-framework verified

This page aggregates 1 real-world runs of Qwen3.5-0.8B (Q6_K) on NVIDIA RTX 5090 with Strata, contributed by 1 independent source platforms; metrics are averages of published measurements.

317.49 tok/s

Decode

Decode speed

1,303.51 tok/s

Prefill

Prefill speed

2.21 s

TTFT

Time to first token

3.78 GB

VRAM

VRAM usage

L2 Cross-framework verified

1

Measured runs

1

Independent sources

Other

Source platforms

Today

Last verified

Performance

  1. Ollama · Q3_K_L Decode 320.91 · Prefill 1586.97 ·
  2. Ollama · Q8_0 Decode 384.44 · Prefill 1342.7 ·
  3. Strata · FP16 Decode 337.54 · Prefill 1202.28 ·
  4. Strata · Q6_K (current)Decode 317.49 · Prefill 1303.51 ·

Core figures

Decode (avg)
317.49 tok/s
Prefill (avg)
1,303.51 tok/s
TTFT (avg)
2.21 s
VRAM (avg)
3.78 GB
MTP acceptance rate
0.67
TTFB
2.54 GB
Power draw
523.1 W

Configuration

Member-level fields are taken from the most recent run

Model
Qwen3.5-0.8B
Quantization
Q6_K
Framework
Strata
Version
v0.53
Context length
4096 tokens
Batch size
256
GPU layers
5
Flash Attention
On

Hardware

Nominal and measured figures are shown side by side; whether it runs is the reader's call

GPU
NVIDIA RTX 5090
Nominal VRAM
32 GB
Measured VRAM (avg)
3.78 GB
OS
Windows 11 23H2
Driver
565.57
CUDA
12.6
Power draw
523.1 W

Sources & evidence

1 measured records in total, each traceable to its original source

  1. L2 Cross-framework verifiedOtherOriginal link Verified on 2026-10-10

    v0.53 · Windows 11 23H2 · CUDA 12.6 · 4096 ctx

    317.49 tok/s

    Decode

    1,303.51 tok/s

    Prefill

    2.21 s

    TTFT

    3.78 GB

    VRAM

    0.67

    MTP

    523.1 W

    Power draw

    L2 batch seed 2026-10-10: Strata on hw_id=13 running model_id=51 at Q6_K