Compile. Benchmark.
Deploy.
We measure Muna-compiled language and image models across artifact size, cold starts, and inference performance. Every number below is derived from raw result files you can download, with the run configuration published alongside it.
FLUX.2 Klein 9B
image@black-forest-labs/flux-2-klein-9b
NVIDIA B200 · August 2026
- compiled code size
- 16MB
- compiled code size
- artifact pull
- 1.6GB · 17 files
- artifact pull
- median cold-start time to first image
- 5.0s
- median cold-start time to first image
- median generation time · 1024×1024
- 426ms
- median generation time · 1024×1024
Gemma 4 26B A4B
chat@google/gemma-4-26b-a4b-it
NVIDIA B200 · May 2026 · vs SGLang
- compiled code size
- 96MB
- compiled code size
- artifact pull, vs 25.8GB · 193,735 files
- 760MB · 5 files
- artifact pull, vs 25.8GB · 193,735 files
- CTTFT @p90 · Muna (bare metal)
- 6.3s
- CTTFT @p90 · Muna (bare metal)
- lower median CTTFT vs SGLang (Modal)
- 18×
- lower median CTTFT vs SGLang (Modal)
- TTFT @16k-token prompts
- 147ms
- TTFT @16k-token prompts
- tok/s across prompt categories
- 206–593
- tok/s across prompt categories