Qwen3-32B · Q4_K_M
L2 跨框架验证NVIDIA RTX 4090 · 24GB · llama.cpp
GitHub
2 天前
18.4 tok/s
Decode
216.93 tok/s
Prefill
1.19 s
TTFT
19.5 GB
VRAM
配置(最近一次实测):
来源:GitHub、Reddit · 3 次实测 · 2 个独立来源
NVIDIA RTX 4090 · 24GB · llama.cpp
GitHub
2 天前
18.4 tok/s
Decode
216.93 tok/s
Prefill
1.19 s
TTFT
19.5 GB
VRAM
配置(最近一次实测):
来源:GitHub、Reddit · 3 次实测 · 2 个独立来源
NVIDIA RTX 4090 · 24GB · Ollama
V2EX
2 天前
17.4 tok/s
Decode
198.2 tok/s
Prefill
1.31 s
TTFT
19.8 GB
VRAM
配置(最近一次实测):
来源:V2EX · 1 次实测 · 1 个独立来源
NVIDIA RTX 3090 · 24GB · llama.cpp
GitHub
2 天前
12.95 tok/s
Decode
143.35 tok/s
Prefill
1.6 s
TTFT
19.5 GB
VRAM
配置(最近一次实测):
来源:GitHub、Bilibili · 2 次实测 · 2 个独立来源
NVIDIA RTX 4060 Ti 16GB · 16GB · llama.cpp
GitHub
2 天前
33.8 tok/s
Decode
620.4 tok/s
Prefill
0.41 s
TTFT
6.2 GB
VRAM
配置(最近一次实测):
来源:GitHub · 1 次实测 · 1 个独立来源
Apple M3 Max(MacBook Pro 14) · 36GB · Ollama
2 天前
39.15 tok/s
Decode
408.85 tok/s
Prefill
0.51 s
TTFT
5.8 GB
VRAM
配置(最近一次实测):
来源:Reddit、Hugging Face · 2 次实测 · 2 个独立来源
NVIDIA RTX 4090 · 24GB · vLLM
GitHub
2 天前
63.6 tok/s
Decode
1,192.75 tok/s
Prefill
0.29 s
TTFT
11.1 GB
VRAM
配置(最近一次实测):
来源:GitHub、Hugging Face · 2 次实测 · 2 个独立来源
共 6 条结果