搜索实测数据

试试这些搜索

找到 6 条相关运行记录

硬件 全部模型 全部框架 全部量化 全部证据等级 全部来源平台 全部

Qwen3-32B · Q4_K_M

L2 跨框架验证

NVIDIA RTX 4090 · 24GB · llama.cpp

GitHub

2 天前

18.4 tok/s

Decode

216.93 tok/s

Prefill

1.19 s

TTFT

19.5 GB

VRAM

配置(最近一次实测):

Ubuntu 24.04 CUDA 12.4 Q4_K_M llama.cpp b4567 Flash Attention

来源:GitHub、Reddit · 3 次实测 · 2 个独立来源

Qwen3-32B · Q4_K_M

L1 已复现

NVIDIA RTX 4090 · 24GB · Ollama

V2EX

2 天前

17.4 tok/s

Decode

198.2 tok/s

Prefill

1.31 s

TTFT

19.8 GB

VRAM

配置(最近一次实测):

Ubuntu 24.04 CUDA 12.4 Q4_K_M Ollama 0.5.7

来源:V2EX · 1 次实测 · 1 个独立来源

Qwen3-32B · Q4_K_M

L1 已复现

NVIDIA RTX 3090 · 24GB · llama.cpp

GitHub

2 天前

12.95 tok/s

Decode

143.35 tok/s

Prefill

1.6 s

TTFT

19.5 GB

VRAM

配置(最近一次实测):

Ubuntu 22.04 CUDA 12.2 Q4_K_M llama.cpp b4567

来源:GitHub、Bilibili · 2 次实测 · 2 个独立来源

Llama-3.1-8B-Instruct · Q5_K_M

L1 已复现

NVIDIA RTX 4060 Ti 16GB · 16GB · llama.cpp

GitHub

2 天前

33.8 tok/s

Decode

620.4 tok/s

Prefill

0.41 s

TTFT

6.2 GB

VRAM

配置(最近一次实测):

Windows 11 24H2 CUDA 12.4 Q5_K_M llama.cpp b4602 Flash Attention

来源:GitHub · 1 次实测 · 1 个独立来源

Llama-3.1-8B-Instruct · Q4_K_M

L1 已复现

Apple M3 Max(MacBook Pro 14) · 36GB · Ollama

Reddit

2 天前

39.15 tok/s

Decode

408.85 tok/s

Prefill

0.51 s

TTFT

5.8 GB

VRAM

配置(最近一次实测):

macOS 15.3 Q4_K_M Ollama 0.5.7

来源:Reddit、Hugging Face · 2 次实测 · 2 个独立来源

DeepSeek-R1-Distill-Qwen-14B · AWQ-INT4

L2 跨框架验证

NVIDIA RTX 4090 · 24GB · vLLM

GitHub

2 天前

63.6 tok/s

Decode

1,192.75 tok/s

Prefill

0.29 s

TTFT

11.1 GB

VRAM

配置(最近一次实测):

Ubuntu 24.04 CUDA 12.4 AWQ-INT4 vLLM 0.7.3

来源:GitHub、Hugging Face · 2 次实测 · 2 个独立来源

共 6 条结果