Back to search results

Llama-3.1-8B-Instruct · Q5_K_M

NVIDIA RTX 4060 Ti 16GB16GBllama.cpp b4602L1 Reproduced

This page aggregates 1 real-world runs of Llama-3.1-8B-Instruct (Q5_K_M) on NVIDIA RTX 4060 Ti 16GB with llama.cpp, contributed by 1 independent source platforms; metrics are averages of published measurements.

33.8 tok/s

Decode

Decode speed

620.4 tok/s

Prefill

Prefill speed

0.41 s

TTFT

Time to first token

6.2 GB

VRAM

VRAM usage

L1 Reproduced

1

Measured runs

1

Independent sources

GitHub

Source platforms

2026-08-05

Last verified

Performance

  1. llama.cpp · Q5_K_M (current)Decode 33.8 · Prefill 620.4 ·

Core figures

Decode (avg)
33.8 tok/s
Prefill (avg)
620.4 tok/s
TTFT (avg)
0.41 s
VRAM (avg)
6.2 GB
MTP acceptance rate
TTFB
— GB
Power draw
142.7 W

Configuration

Member-level fields are taken from the most recent run

Model
Llama-3.1-8B-Instruct
Quantization
Q5_K_M
Framework
llama.cpp
Version
b4602
Context length
8192 tokens
GPU layers
99
Flash Attention
On

Hardware

Nominal and measured figures are shown side by side; whether it runs is the reader's call

GPU
NVIDIA RTX 4060 Ti 16GB
Nominal VRAM
16 GB
Measured VRAM (avg)
6.2 GB
OS
Windows 11 24H2
Driver
552.22
CUDA
12.4
Power draw
142.7 W

Sources & evidence

1 measured records in total, each traceable to its original source

  1. L1 ReproducedGitHubOriginal link Verified on 2026-08-05

    b4602 · Windows 11 24H2 · CUDA 12.4 · 8192 ctx

    33.8 tok/s

    Decode

    620.4 tok/s

    Prefill

    0.41 s

    TTFT

    6.2 GB

    VRAM

    MTP

    142.7 W

    Power draw

    8B Q5_K_M 在 4060 Ti 16G 上约 34 tok/s。