Benchmark search

Try these searches

Found 41 run configurations

Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.

ModelQwen3-1.7B

Qwen3-1.7B · IQ2_M

L2 Cross-framework verified

NVIDIA RTX 4090 · 24GB · llama.cpp

Other

Today

323.97 tok/s

Decode

1,188.09 tok/s

Prefill

2.05 s

TTFT

3.96 GB

VRAM

Config (latest run)
Windows 11 23H2 CUDA 12.6 IQ2_M llama.cpp v0.33 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · Q2_K

L2 Cross-framework verified

NVIDIA RTX 4090 · 24GB · llama.cpp

Other

Today

306.65 tok/s

Decode

1,436.24 tok/s

Prefill

1.31 s

TTFT

4.06 GB

VRAM

Config (latest run)
Ubuntu 24.04 CUDA 12.8 Q2_K llama.cpp v0.52 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · IQ1_S

L2 Cross-framework verified

NVIDIA RTX 4090 · 24GB · Strata

Other

Today

287.87 tok/s

Decode

1,066.63 tok/s

Prefill

0.32 s

TTFT

3.6 GB

VRAM

Config (latest run)
Ubuntu 24.04 CUDA 12.8 IQ1_S Strata v0.35 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · Q2_K

L2 Cross-framework verified

NVIDIA RTX 4090 · 24GB · Strata

Other

Today

285.75 tok/s

Decode

1,116.76 tok/s

Prefill

0.6 s

TTFT

4.06 GB

VRAM

Config (latest run)
Windows 11 23H2 CUDA 12.6 Q2_K Strata v0.51 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · Q3_K_L

L2 Cross-framework verified

NVIDIA RTX 3090 · 24GB · llama.cpp

Other

Today

227.45 tok/s

Decode

862.55 tok/s

Prefill

0.46 s

TTFT

4.67 GB

VRAM

Config (latest run)
Ubuntu 22.04 CUDA 12.6 Q3_K_L llama.cpp v0.31 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · Q5_K_S

L2 Cross-framework verified

NVIDIA RTX 3090 · 24GB · llama.cpp

Other

Today

175.42 tok/s

Decode

578.57 tok/s

Prefill

0.86 s

TTFT

5.31 GB

VRAM

Config (latest run)
macOS 15.0 CUDA 12.6 Q5_K_S llama.cpp v0.23

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · Q6_K

L2 Cross-framework verified

NVIDIA RTX 3090 · 24GB · Strata

Other

Today

215.56 tok/s

Decode

902.66 tok/s

Prefill

0.24 s

TTFT

5.67 GB

VRAM

Config (latest run)
Ubuntu 24.04 CUDA 12.6 Q6_K Strata v0.51 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · IQ2_XS

L2 Cross-framework verified

Apple M3 Max · 36GB · vLLM

Other

Today

309.15 tok/s

Decode

1,109.54 tok/s

Prefill

2.17 s

TTFT

3.81 GB

VRAM

Config (latest run)
macOS 15.0 CUDA N/A IQ2_XS vLLM v0.24 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · Q3_K_L

L2 Cross-framework verified

Apple M3 Max · 36GB · vLLM

Other

Today

226.47 tok/s

Decode

1,116.37 tok/s

Prefill

0.18 s

TTFT

4.67 GB

VRAM

Config (latest run)
macOS 15.0 CUDA N/A Q3_K_L vLLM v0.41 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · FP16

L2 Cross-framework verified

Apple M3 Max · 36GB · Ollama

Other

Today

96.85 tok/s

Decode

510.37 tok/s

Prefill

0.44 s

TTFT

8.7 GB

VRAM

Config (latest run)
macOS 15.0 CUDA N/A FP16 Ollama v0.28 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · IQ4_XS

L2 Cross-framework verified

Apple M3 Max · 36GB · Ollama

Other

Today

188.36 tok/s

Decode

704.89 tok/s

Prefill

1.31 s

TTFT

4.75 GB

VRAM

Config (latest run)
macOS 15.0 CUDA N/A IQ4_XS Ollama v0.39 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · Q5_K_S

L2 Cross-framework verified

Apple M3 Max · 36GB · Strata

Other

Today

239.13 tok/s

Decode

849.25 tok/s

Prefill

1.36 s

TTFT

5.31 GB

VRAM

Config (latest run)
macOS 15.0 CUDA N/A Q5_K_S Strata v0.31

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · IQ3_XS

L2 Cross-framework verified

AMD Radeon RX 7900 XTX · 24GB · llama.cpp

Other

Today

234.47 tok/s

Decode

1,079.52 tok/s

Prefill

2.39 s

TTFT

4.32 GB

VRAM

Config (latest run)
Windows 11 23H2 CUDA 12.8 IQ3_XS llama.cpp v0.27 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · Q4_K_M

L2 Cross-framework verified

AMD Radeon RX 7900 XTX · 24GB · llama.cpp

Other

Today

178.72 tok/s

Decode

842.7 tok/s

Prefill

0.98 s

TTFT

5.03 GB

VRAM

Config (latest run)
Ubuntu 24.04 CUDA 12.4 Q4_K_M llama.cpp v0.38

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · Q3_K_M

L2 Cross-framework verified

AMD Radeon RX 7900 XTX · 24GB · Ollama

Other

Today

196.24 tok/s

Decode

705.31 tok/s

Prefill

0.21 s

TTFT

4.53 GB

VRAM

Config (latest run)
Ubuntu 22.04 CUDA 12.4 Q3_K_M Ollama v0.38 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · Q8_0

L2 Cross-framework verified

AMD Radeon RX 7900 XTX · 24GB · Ollama

Other

Today

111.35 tok/s

Decode

362.65 tok/s

Prefill

0.89 s

TTFT

6.21 GB

VRAM

Config (latest run)
Windows 11 23H2 CUDA 12.8 Q8_0 Ollama v0.37 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · Q4_K_S

L2 Cross-framework verified

AMD Radeon RX 7900 XTX · 24GB · Strata

Other

Today

249 tok/s

Decode

1,211.72 tok/s

Prefill

2.07 s

TTFT

4.88 GB

VRAM

Config (latest run)
macOS 15.0 CUDA 12.8 Q4_K_S Strata v0.22 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · Q5_K_M

L2 Cross-framework verified

NVIDIA RTX 5090 · 32GB · llama.cpp

Other

Today

353.77 tok/s

Decode

1,906.09 tok/s

Prefill

1.04 s

TTFT

5.42 GB

VRAM

Config (latest run)
macOS 15.0 CUDA 12.6 Q5_K_M llama.cpp v0.36 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · Q6_K

L2 Cross-framework verified

NVIDIA RTX 5090 · 32GB · llama.cpp

Other

Today

345.13 tok/s

Decode

1,779.25 tok/s

Prefill

0.42 s

TTFT

5.67 GB

VRAM

Config (latest run)
Ubuntu 24.04 CUDA 12.8 Q6_K llama.cpp v0.27 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3-1.7B · FP16

L2 Cross-framework verified

NVIDIA RTX 5090 · 32GB · Ollama

Other

Today

199.27 tok/s

Decode

1,057.11 tok/s

Prefill

2.17 s

TTFT

8.7 GB

VRAM

Config (latest run)
macOS 15.0 CUDA 12.4 FP16 Ollama v0.38

Sources: Other · 1 runs · 1 independent sources

41 results

Related models
Related hardware
Related frameworks
Related quantizations