munaOur compiled models.
Your AI client.
We help AI teams run inference with open models, with a focus on boosting GPU utilization so that they serve more and spend less.
Each model is compiled into a self-contained binary and served by an inference runtime we wrote to pack GPUs: several models co-located on a single GPU (megakernel style), with optimizations up and down the stack.
As a result, tokens cost about 40% less here than on OpenRouter, behind the OpenAI and Anthropic APIs you already use.
Quick start
Point the OpenAI SDK at https://inference.muna.ai/v1, or the Anthropic SDK at https://inference.muna.ai, and call any model below.
from openai import OpenAI
# 💥 Create an OpenAI client with the Muna URLopenai = OpenAI( base_url="https://inference.muna.ai/v1", api_key="<your Muna API key>")
# 🔥 Create a chat completioncompletion = openai.chat.completions.create( model="@google/gemma-4-26b-a4b-it", messages=[{ "role": "user", "content": "What is a GPU?" }])
# 🚀 Print the outputprint(completion.choices[0].message.content)Models
| Model | Kind | Price |
|---|---|---|
@google/gemma-4-26b-a4b-itStream chat completions from Gemma 4 26B. | chat |
|
@qwen/qwen-3.8-27bDense 27B chat model, 128K context, tool calling. | chat |
|
@nomic/nomic-embed-text-v1.5Text embeddings, 768 dimensions, up to 8K tokens. | embedding |
|
@qwen/qwen-3-embedding-8bMultilingual text embeddings, 4,096 dimensions, up to 40K tokens. | embedding |
|
@deepseek/deepseek-v4-flashFast MoE chat model.Coming soon | chat | — |
@zai/glm-5.3-flashFast chat model from Z.ai.Coming soon | chat | — |
Need a model that is not here? Ask; bringing one up takes days, not quarters.
How it works
01Write a Python function
Each model starts as a Python function: load the model, then stream tokens out. This is an excerpt of the function behind @qwen/qwen-3.8-27b.
from muna import compilefrom muna.beta import TorchToSGLangInferenceMetadatafrom muna.beta.openai import ChatCompletionChunk, Messagefrom transformers import Qwen3_5ForCausalLM
# Load the model and request managermodel = Qwen3_5ForCausalLM._from_config(config)manager = model.init_continuous_batching(...)
@compile( tag="@qwen/qwen-3.8-27b", targets=["x86_64-unknown-linux-gnu"], metadata=[ # Lower the PyTorch model onto our inference engine TorchToSGLangInferenceMetadata( model=model, compute_architecture="sm_100" ) ])def qwen_3_8_27b( messages: list[Message], ...) -> Iterator[ChatCompletionChunk]: """ Stream chat completions from Qwen 3.8 27B. """ inputs = _process(messages) manager.add_request( input_ids=inputs.input_ids[0], request_id=req_id, streaming=True ) for output in manager.request_id_iter(request_id=req_id): yield _chunk(req_id, output)02Compile it
Muna compiles the function into a self-contained binary that runs on our C++ inference engine. There is no container and no Python interpreter at runtime.
03Serve it with fast cold starts
The binary is tens of megabytes, so a node downloads it in moments and starts serving as soon as the weights are read off disk. These numbers come from production and refresh every five minutes.
| Model | Binary | Cold start @p50 (@p90) |
|---|---|---|
@google/gemma-4-26b-a4b-it | 78.7 MB | 8.5 s (12.0 s) |
@qwen/qwen-3.8-27b | 54.7 MB | 9.1 s (12.9 s) |
@nomic/nomic-embed-text-v1.5 | 24.7 MB | 1.8 s (3.0 s) |
@qwen/qwen-3-embedding-8b | 24.6 MB | 3.2 s (4.0 s) |
"Binary" is the compiled model library that inference nodes download and execute. Cold starts are load-to-ready times across production loads since September 14 where the model had the node to itself; loads that overlap another model's load share disk bandwidth and take longer. In our Gemma 4 benchmark, the first token arrives 18× sooner than with SGLang.
04Share the GPU
Because a model loads in seconds, a GPU does not have to be dedicated to one model. Our runtime keeps several models on the same GPU, gives each one the GPU when it has requests, and evicts models that go idle. Every GPU-hour serves more tokens, which is why the prices above are lower.
05Run it yourself
You can run the exact same compiled models that power our inference endpoints on your own machines. Use the local_* acceleration with the Muna client:
from muna import Muna
# 💥 Create a Muna client (reads MUNA_ACCESS_KEY)muna = Muna()
# 🔥 Run the same compiled model on your own GPUcompletion = muna.beta.openai.chat.completions.create( model="@qwen/qwen-3.8-27b", messages=[{ "role": "user", "content": "What is a GPU?" }], acceleration="local_gpu")
# 🚀 Print the outputprint(completion.choices[0].message.content)