Benchmark search

Try these searches

Found 2 run configurations

Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.

ModelTernary Bonsai 2 27B

Ternary Bonsai 2 27B · PTQ1_0

L1 Reproduced

NVIDIA RTX 5060 Ti 16GB · 16GB · llama.cpp

GitHub

Yesterday

8.1 tok/s

Decode

21.87 tok/s

Prefill

— s

TTFT

— GB

VRAM

Config (latest run)
Windows 11 PTQ1_0 llama.cpp b11047

Sources: GitHub · 1 runs · 1 independent sources

Ternary Bonsai 2 27B · PTQ1_0

L0 Self-reported

NVIDIA RTX 3080 · 10GB · llama.cpp

GitHub

5 days ago

52.18 tok/s

Decode

1,210.6 tok/s

Prefill

— s

TTFT

7.03 GB

VRAM

Config (latest run)
Ubuntu 24.04 CUDA 12.8 PTQ1_0 llama.cpp prism-b10735-842b188 Flash Attention

Sources: GitHub · 1 runs · 1 independent sources

2 results

Related models
Related hardware
Related frameworks
Related quantizations