Compile. Benchmark.
Deploy.

We measure Muna-compiled language and image models across artifact size, cold starts, and inference performance. Every number below is derived from raw result files you can download, with the run configuration published alongside it.

FLUX.2 Klein 9B

image

@black-forest-labs/flux-2-klein-9b

NVIDIA B200 · August 2026

compiled code size
16MB
compiled code size
artifact pull
1.6GB · 17 files
artifact pull
median cold-start time to first image
5.0s
median cold-start time to first image
median generation time · 1024×1024
426ms
median generation time · 1024×1024

Gemma 4 26B A4B

chat

@google/gemma-4-26b-a4b-it

NVIDIA B200 · May 2026 · vs SGLang

compiled code size
96MB
compiled code size
artifact pull, vs 25.8GB · 193,735 files
760MB · 5 files
artifact pull, vs 25.8GB · 193,735 files
CTTFT @p90 · Muna (bare metal)
6.3s
CTTFT @p90 · Muna (bare metal)
lower median CTTFT vs SGLang (Modal)
18×
lower median CTTFT vs SGLang (Modal)
TTFT @16k-token prompts
147ms
TTFT @16k-token prompts
tok/s across prompt categories
206–593
tok/s across prompt categories