Muna logomuna

Our compiled models.
Your AI client.

We help AI teams run inference with open models, with a focus on boosting GPU utilization so that they serve more and spend less.

Each model is compiled into a self-contained binary and served by an inference runtime we wrote to pack GPUs: several models co-located on a single GPU (megakernel style), with optimizations up and down the stack.

As a result, tokens cost about 40% less here than on OpenRouter, behind the OpenAI and Anthropic APIs you already use.

Quick start

Point the OpenAI SDK at https://inference.muna.ai/v1, or the Anthropic SDK at https://inference.muna.ai, and call any model below.

from openai import OpenAI
# 💥 Create an OpenAI client with the Muna URL
openai = OpenAI(
base_url="https://inference.muna.ai/v1",
api_key="<your Muna API key>"
)
# 🔥 Create a chat completion
completion = openai.chat.completions.create(
model="@google/gemma-4-26b-a4b-it",
messages=[{ "role": "user", "content": "What is a GPU?" }]
)
# 🚀 Print the output
print(completion.choices[0].message.content)

Models

ModelKindPrice
@google/gemma-4-26b-a4b-it
Stream chat completions from Gemma 4 26B.
chat
$0.065
in / 1M
$0.02
cached / 1M
$0.30
out / 1M
@qwen/qwen-3.8-27b
Dense 27B chat model, 128K context, tool calling.
chat
$0.25
in / 1M
$0.02
cached / 1M
$1.75
out / 1M
35% cheaper than OpenRouter
@nomic/nomic-embed-text-v1.5
Text embeddings, 768 dimensions, up to 8K tokens.
embedding
$0.01
/ 1M
@qwen/qwen-3-embedding-8b
Multilingual text embeddings, 4,096 dimensions, up to 40K tokens.
embedding
$0.02
/ 1M
@deepseek/deepseek-v4-flash
Fast MoE chat model.Coming soon
chat—
@zai/glm-5.3-flash
Fast chat model from Z.ai.Coming soon
chat—

Need a model that is not here? Ask; bringing one up takes days, not quarters.

How it works

01Write a Python function

Each model starts as a Python function: load the model, then stream tokens out. This is an excerpt of the function behind @qwen/qwen-3.8-27b.

from muna import compile
from muna.beta import TorchToSGLangInferenceMetadata
from muna.beta.openai import ChatCompletionChunk, Message
from transformers import Qwen3_5ForCausalLM
# Load the model and request manager
model = Qwen3_5ForCausalLM._from_config(config)
manager = model.init_continuous_batching(...)
@compile(
tag="@qwen/qwen-3.8-27b",
targets=["x86_64-unknown-linux-gnu"],
metadata=[
# Lower the PyTorch model onto our inference engine
TorchToSGLangInferenceMetadata(
model=model,
compute_architecture="sm_100"
)
]
)
def qwen_3_8_27b(
messages: list[Message],
...
) -> Iterator[ChatCompletionChunk]:
"""
Stream chat completions from Qwen 3.8 27B.
"""
inputs = _process(messages)
manager.add_request(
input_ids=inputs.input_ids[0],
request_id=req_id,
streaming=True
)
for output in manager.request_id_iter(request_id=req_id):
yield _chunk(req_id, output)

02Compile it

Muna compiles the function into a self-contained binary that runs on our C++ inference engine. There is no container and no Python interpreter at runtime.

$ muna compile qwen_3_8_27b.py

03Serve it with fast cold starts

The binary is tens of megabytes, so a node downloads it in moments and starts serving as soon as the weights are read off disk. These numbers come from production and refresh every five minutes.

ModelBinaryCold start @p50 (@p90)
@google/gemma-4-26b-a4b-it78.7 MB8.5 s (12.0 s)
@qwen/qwen-3.8-27b54.7 MB9.1 s (12.9 s)
@nomic/nomic-embed-text-v1.524.7 MB1.8 s (3.0 s)
@qwen/qwen-3-embedding-8b24.6 MB3.2 s (4.0 s)

"Binary" is the compiled model library that inference nodes download and execute. Cold starts are load-to-ready times across production loads since September 14 where the model had the node to itself; loads that overlap another model's load share disk bandwidth and take longer. In our Gemma 4 benchmark, the first token arrives 18× sooner than with SGLang.

04Share the GPU

Because a model loads in seconds, a GPU does not have to be dedicated to one model. Our runtime keeps several models on the same GPU, gives each one the GPU when it has requests, and evicts models that go idle. Every GPU-hour serves more tokens, which is why the prices above are lower.

Dedicated: one GPU per model
2 GPUs at $14/hr, 70% of each one unused
30% utilized
30% utilized
Co-located: models take turns
1 GPU at $7/hr, same traffic
60% utilized
Model A
Model B
Idle

05Run it yourself

You can run the exact same compiled models that power our inference endpoints on your own machines. Use the local_* acceleration with the Muna client:

from muna import Muna
# 💥 Create a Muna client (reads MUNA_ACCESS_KEY)
muna = Muna()
# 🔥 Run the same compiled model on your own GPU
completion = muna.beta.openai.chat.completions.create(
model="@qwen/qwen-3.8-27b",
messages=[{ "role": "user", "content": "What is a GPU?" }],
acceleration="local_gpu"
)
# 🚀 Print the output
print(completion.choices[0].message.content)