Benchmark search

Try these searches

Found 5 run configurations

Each row is a run configuration (model + hardware + framework + quantization); repeated runs of the same setup are merged — raw counts appear in each result’s source line.

ModelQwen3.6-35B-A3B

Qwen3.6-35B-A3B · IQ2_XS

L2 Cross-framework verified

Apple M3 Max · 36GB · llama.cpp

Other

Today

89.41 tok/s

Decode

406.89 tok/s

Prefill

0.17 s

TTFT

34.2 GB

VRAM

Config (latest run)
macOS 15.0 CUDA N/A IQ2_XS llama.cpp v0.25 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3.6-35B-A3B · IQ1_S

L2 Cross-framework verified

Apple M3 Max · 36GB · vLLM

Other

Today

102.56 tok/s

Decode

411.25 tok/s

Prefill

1.72 s

TTFT

32.97 GB

VRAM

Config (latest run)
macOS 15.0 CUDA N/A IQ1_S vLLM v0.36 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3.6-35B-A3B · IQ3_XS

L2 Cross-framework verified

Apple M4 Max · 128GB · llama.cpp

Other

Today

121.63 tok/s

Decode

540.05 tok/s

Prefill

2.37 s

TTFT

47.67 GB

VRAM

Config (latest run)
macOS 15.0 CUDA N/A IQ3_XS llama.cpp v0.23

Sources: Other · 1 runs · 1 independent sources

Qwen3.6-35B-A3B · Q4_K_S

L2 Cross-framework verified

Apple M4 Max · 128GB · llama.cpp

Other

Today

111.01 tok/s

Decode

387.61 tok/s

Prefill

2.26 s

TTFT

59.43 GB

VRAM

Config (latest run)
macOS 15.0 CUDA N/A Q4_K_S llama.cpp v0.46 Flash Attention

Sources: Other · 1 runs · 1 independent sources

Qwen3.6-35B-A3B · Q3_K_L

L2 Cross-framework verified

Apple M4 Max · 128GB · Ollama

Other

Today

102.3 tok/s

Decode

395.71 tok/s

Prefill

2.39 s

TTFT

55.02 GB

VRAM

Config (latest run)
macOS 15.0 CUDA N/A Q3_K_L Ollama v0.51

Sources: Other · 1 runs · 1 independent sources

5 results

Related models
Related hardware
Related frameworks
Related quantizations