MiniMax-H3 · Online dynamic FP8
This page aggregates 1 real-world runs of MiniMax-H3 (Online dynamic FP8) on NVIDIA DGX Spark with vLLM-Omni, contributed by 1 independent source platforms; metrics are averages of published measurements.
80.58 s
Gen time
Per image / per clip
89.17 GB
VRAM
VRAM usage
1
Measured runs
1
Independent sources
GitHub
Source platforms
6 days ago
Last verified
Performance
No published records for this metric under the same model + hardware yet
Core figures
- Gen time (avg)
- 80.58 s
- VRAM (avg)
- 89.17 GB
- Power draw
- — W
- Output
- 768x448 · 56 frames · 20 steps · 24fps
Configuration
Member-level fields are taken from the most recent run
- Model
- MiniMax-H3
- Quantization
- Online dynamic FP8
- Framework
- vLLM-Omni
- Version
- vLLM-Omni 0.1.dev2381+g310b4b477
- Batch size
- 1
- Output spec
- 768x448 · 56 frames · 20 steps · 24fps
Hardware
Nominal and measured figures are shown side by side; whether it runs is the reader's call
- GPU
- NVIDIA DGX Spark
- Nominal VRAM
- 128 GB
- Measured VRAM (avg)
- 89.17 GB
- OS
- Linux ARM64
- Driver
- 580.173.02
- CUDA
- 13.0
Get started
Weights for this deployment
Original model resources
Base model resources may not match this quantization. Use the verified deployment weights above when available.
Startup command
Command from the latest benchmark runH3_DIFFUSION_ATTENTION_BACKEND=CUDNN_ATTN H3_EXECUTION_MODE=compile H3_CACHE_BACKEND=cache_dit H3_CACHE_CONFIG='{"Fn_compute_blocks":1,"Bn_compute_blocks":0,"max_warmup_steps":4,"max_cached_steps":-1,"residual_diff_threshold":0.10,"max_continuous_cached_steps":1,"enable_taylorseer":false}' docker compose up -dDeployment guide
https://github.com/joeynyc/MiniMax-H3-DGX-Spark#quick-startSources & evidence
1 measured records in total, each traceable to its original source
vLLM-Omni 0.1.dev2381+g310b4b477 · Linux ARM64 · CUDA 13.0
80.58 s
Gen time
89.17 GB
VRAM
— W
Power draw
单DGX Spark,balanced(cuDNN+compile+Cache-DiT0.10):请求2秒、20步,输出56帧/768×448/24fps。warm两次client mean80.579s;89.1659GiB是模型加载统一内存,非峰值,整机peak113.09GiB。排除首次编译;缓存非无损,full-compute111.373s。