Back to search results

Qwen3.5-4B · IQ3_S

Start a discussion
NVIDIA RTX 507012GBStrata v0.50L2 Cross-framework verified

This page aggregates 1 real-world runs of Qwen3.5-4B (IQ3_S) on NVIDIA RTX 5070 with Strata, contributed by 1 independent source platforms; metrics are averages of published measurements.

154.26 tok/s

Decode

Decode speed

563.04 tok/s

Prefill

Prefill speed

1.59 s

TTFT

Time to first token

7.64 GB

VRAM

VRAM usage

L2 Cross-framework verified

1

Measured runs

1

Independent sources

Other

Source platforms

Today

Last verified

Performance

  1. vLLM · Q3_K_M Decode 112.72 · Prefill 436.5 ·
  2. vLLM · Q4_K_M Decode 84.51 · Prefill 408.98 ·
  3. Strata · IQ3_S (current)Decode 154.26 · Prefill 563.04 ·
  4. llama.cpp · Q5_K_M Decode 174.7 · Prefill 1310.3 ·

Core figures

Decode (avg)
154.26 tok/s
Prefill (avg)
563.04 tok/s
TTFT (avg)
1.59 s
VRAM (avg)
7.64 GB
MTP acceptance rate
0.49
TTFB
1.23 GB
Power draw
200.6 W

Configuration

Member-level fields are taken from the most recent run

Model
Qwen3.5-4B
Quantization
IQ3_S
Framework
Strata
Version
v0.50
Context length
8192 tokens
Batch size
1024
GPU layers
8
Flash Attention
Off

Hardware

Nominal and measured figures are shown side by side; whether it runs is the reader's call

GPU
NVIDIA RTX 5070
Nominal VRAM
12 GB
Measured VRAM (avg)
7.64 GB
OS
Windows 11 23H2
Driver
555.85
CUDA
12.6
Power draw
200.6 W

Sources & evidence

1 measured records in total, each traceable to its original source

  1. L2 Cross-framework verifiedOtherOriginal link Verified on 2026-10-10

    v0.50 · Windows 11 23H2 · CUDA 12.6 · 8192 ctx

    154.26 tok/s

    Decode

    563.04 tok/s

    Prefill

    1.59 s

    TTFT

    7.64 GB

    VRAM

    0.49

    MTP

    200.6 W

    Power draw

    L2 batch seed 2026-10-10: Strata on hw_id=18 running model_id=48 at IQ3_S