Back to search results

Qwen3.8-Flash-Next · UD-Q3_K_XL

Start a discussion
4x NVIDIA Tesla P4096GBllama.cpp b11058L1 Reproduced

This page aggregates 1 real-world runs of Qwen3.8-Flash-Next (UD-Q3_K_XL) on 4x NVIDIA Tesla P40 with llama.cpp, contributed by 1 independent source platforms; metrics are averages of published measurements.

21.6 tok/s

Decode

Decode speed

330 tok/s

Prefill

Prefill speed

— s

TTFT

Time to first token

— GB

VRAM

VRAM usage

L1 Reproduced

1

Measured runs

1

Independent sources

GitHub

Source platforms

Yesterday

Last verified

Performance

  1. llama.cpp · UD-Q3_K_XL (current)Decode 21.6 · Prefill 330 ·

Core figures

Decode (avg)
21.6 tok/s
Prefill (avg)
330 tok/s
TTFT (avg)
— s
VRAM (avg)
— GB
MTP acceptance rate
—
TTFB
— GB
Power draw
— W

Configuration

Member-level fields are taken from the most recent run

Model
Qwen3.8-Flash-Next
Quantization
UD-Q3_K_XL
Framework
llama.cpp
Version
b11058

Hardware

Nominal and measured figures are shown side by side; whether it runs is the reader's call

GPU
4x NVIDIA Tesla P40
Nominal VRAM
96 GB
OS
Linux

Get started

Original model resources

Base model resources may not match this quantization. Use the verified deployment weights above when available.

Sources & evidence

1 measured records in total, each traceable to its original source

  1. L1 ReproducedGitHubOriginal link Verified on 2026-10-02

    b11058 · Linux

    21.6 tok/s

    Decode

    330 tok/s

    Prefill

    — s

    TTFT

    0 GB

    VRAM

    —

    MTP

    — W

    Power draw

    Qwen3.8-Flash-Next, no speculation | 20.6 | 20.0 | 21.6 ... Prefill: 122B about 460-477 t/s, Flash-Next about 320-330 t/s.