Benchmark search

Try these searches

Found 18 run configurations

Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.

ModelQwen2.5-Omni-3B

Qwen2.5-Omni-3B · Q8_0

L0 Self-reported

NVIDIA RTX 5060 SUPER · 12GB · llama.cpp

Other

Today

109.2 tok/s

Decode

819 tok/s

Prefill

0.12 s

TTFT

3.7 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q8_0 llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q5_K_M

L0 Self-reported

NVIDIA RTX 5060 SUPER · 12GB · llama.cpp

Other

Today

158.8 tok/s

Decode

1,191 tok/s

Prefill

0.08 s

TTFT

2.7 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q5_K_M llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q4_K_M

L0 Self-reported

NVIDIA RTX 5060 SUPER · 12GB · llama.cpp

Other

Today

194.1 tok/s

Decode

1,455.8 tok/s

Prefill

0.07 s

TTFT

2.3 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q4_K_M llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q8_0

L0 Self-reported

NVIDIA RTX 5070 SUPER · 18GB · llama.cpp

Other

Today

218.4 tok/s

Decode

1,638 tok/s

Prefill

0.06 s

TTFT

3.7 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q8_0 llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q5_K_M

L0 Self-reported

NVIDIA RTX 5070 SUPER · 18GB · llama.cpp

Other

Today

317.7 tok/s

Decode

2,382.8 tok/s

Prefill

0.04 s

TTFT

2.7 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q5_K_M llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q4_K_M

L0 Self-reported

NVIDIA RTX 5070 SUPER · 18GB · llama.cpp

Other

Today

388.3 tok/s

Decode

2,912.3 tok/s

Prefill

0.03 s

TTFT

2.3 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q4_K_M llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q8_0

L0 Self-reported

NVIDIA RTX 3090 · 24GB · llama.cpp

Other

Today

228.2 tok/s

Decode

1,711.5 tok/s

Prefill

0.06 s

TTFT

3.7 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q8_0 llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q5_K_M

L0 Self-reported

NVIDIA RTX 3090 · 24GB · llama.cpp

Other

Today

331.9 tok/s

Decode

2,489.3 tok/s

Prefill

0.04 s

TTFT

2.7 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q5_K_M llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q4_K_M

L0 Self-reported

NVIDIA RTX 3090 · 24GB · llama.cpp

Other

Today

405.6 tok/s

Decode

3,042 tok/s

Prefill

0.03 s

TTFT

2.3 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q4_K_M llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q8_0

L0 Self-reported

NVIDIA RTX 5070 · 12GB · llama.cpp

Other

Today

163.8 tok/s

Decode

1,228.5 tok/s

Prefill

0.08 s

TTFT

3.7 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q8_0 llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q5_K_M

L0 Self-reported

NVIDIA RTX 5070 · 12GB · llama.cpp

Other

Today

238.3 tok/s

Decode

1,787.3 tok/s

Prefill

0.06 s

TTFT

2.7 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q5_K_M llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q4_K_M

L0 Self-reported

NVIDIA RTX 5070 · 12GB · llama.cpp

Other

Today

291.2 tok/s

Decode

2,184 tok/s

Prefill

0.05 s

TTFT

2.3 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q4_K_M llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q8_0

L0 Self-reported

NVIDIA RTX 5090 · 32GB · llama.cpp

Other

Today

436.8 tok/s

Decode

3,276 tok/s

Prefill

0.03 s

TTFT

3.7 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q8_0 llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q5_K_M

L0 Self-reported

NVIDIA RTX 5090 · 32GB · llama.cpp

Other

Today

635.3 tok/s

Decode

4,764.8 tok/s

Prefill

0.02 s

TTFT

2.7 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q5_K_M llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q4_K_M

L0 Self-reported

NVIDIA RTX 5090 · 32GB · llama.cpp

Other

Today

776.5 tok/s

Decode

5,823.8 tok/s

Prefill

0.02 s

TTFT

2.3 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q4_K_M llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q8_0

L0 Self-reported

NVIDIA RTX 4090 · 24GB · llama.cpp

Other

Today

245.7 tok/s

Decode

1,842.8 tok/s

Prefill

0.05 s

TTFT

3.7 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q8_0 llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q5_K_M

L0 Self-reported

NVIDIA RTX 4090 · 24GB · llama.cpp

Other

Today

357.4 tok/s

Decode

2,680.5 tok/s

Prefill

0.04 s

TTFT

2.7 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q5_K_M llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen2.5-Omni-3B · Q4_K_M

L0 Self-reported

NVIDIA RTX 4090 · 24GB · llama.cpp

Other

Today

436.8 tok/s

Decode

3,276 tok/s

Prefill

0.03 s

TTFT

2.3 GB

VRAM

Config (latest run)
Linux CUDA 12.6 Q4_K_M llama.cpp b4200 Flash Attention

Sources: Other · 1 runs · 1 independent sources

18 results

Related models
Related hardware
Related frameworks
Related quantizations