munaZero Data Retention Policy
Last modified October 9, 2026
Overview
Effective date: September 10, 2026
Applies to: the Muna Inference API at
This policy supplements the Muna Privacy Policy. Where this policy is more protective of Customer Content, this policy controls.
Applies to: the Muna Inference API at
https://inference.muna.aiThis policy supplements the Muna Privacy Policy. Where this policy is more protective of Customer Content, this policy controls.
1. Summary
Muna operates the Inference API on a zero data retention basis. With limited exceptions as described in Section 8, Muna does not store the prompts you send or the outputs the models return. Content submitted to the API is held only in volatile memory for as long as is technically necessary to compute and return a response, and in the in-memory caches described in Section 6.
2. Definitions
"Customer Content" means all data you submit to, and all data returned by, the Inference API, including: prompts, messages, system instructions, tool and function definitions, tool results, model completions, reasoning output, embedding inputs and the resulting vectors, and any images, audio, documents or other files submitted for processing or generated in response.
"Service Metadata" means operational records that do not contain Customer Content. See Section 5.
"Retention" means writing to any persistent medium, including disk, object storage, databases, log aggregation systems, backups, or snapshots.
"Service Metadata" means operational records that do not contain Customer Content. See Section 5.
"Retention" means writing to any persistent medium, including disk, object storage, databases, log aggregation systems, backups, or snapshots.
3. Scope
This policy applies to all models Muna hosts and all inference endpoints, including:
It applies to every model in the serving catalog, including third-party open-weight models Muna hosts on its own infrastructure.
| Endpoint | Covered |
|---|---|
POST /v1/chat/completions | Yes, including tool use and streaming |
POST /v1/messages | Yes, including tool use, thinking, and streaming |
POST /v1/embeddings | Yes |
POST /v1/images/generations | Yes |
GET /v1/models | Yes (contains no Customer Content) |
GET /v1/models/{model} | Yes (contains no Customer Content) |
It applies to every model in the serving catalog, including third-party open-weight models Muna hosts on its own infrastructure.
4. What Muna Does Not Retain
Except in the limited circumstances described in Section 8, Muna does not:
In no circumstances does Muna:
Other than as described in Section 8, Customer Content exists only in volatile memory: in the serving process for the duration of the request, and in the in-memory caches described in Section 6.
- write prompts, completions, embedding inputs or output vectors to persistent storage;
- include Customer Content in application logs, error logs, traces, or crash dumps;
- include Customer Content in backups or disaster-recovery snapshots;
- make Customer Content available to Muna personnel.
In no circumstances does Muna:
- use Customer Content to train, fine-tune, evaluate, or otherwise improve any model (whether Muna's own or a third party's);
- sell, license, or otherwise disclose Customer Content to any third party.
Other than as described in Section 8, Customer Content exists only in volatile memory: in the serving process for the duration of the request, and in the in-memory caches described in Section 6.
5. What Muna Does Retain
To operate, bill for, and secure the service, Muna retains the following. Other than the diagnostic logs described in Section 8.1, none of it includes Customer Content:
Token counts are retained; the tokens themselves are not.
| Category | Examples | Retention |
|---|---|---|
| Request metadata | Timestamp, request ID, model ID, HTTP status, latency | 30 days |
| Usage and billing | Request ID, account ID, model ID, endpoint, prompt / completion / cached token counts | 7 years, as required for billing and tax |
| Rate limiting | Per-account counters | Transient |
| Infrastructure logs | Load-balancer and edge logs (source IP, path, status, bytes) | 30 days |
| Account data | Account holder name, email, billing details | Per the Privacy Policy |
| Diagnostic logs | Error and crash diagnostics, which may incidentally contain fragments of Customer Content (see Section 8.1) | 30 days |
Token counts are retained; the tokens themselves are not.
6. Caching
To avoid redundant work, the Inference API keeps the following caches. All of them reside in volatile memory only, are never written to persistent storage, and are cleared on process restart:
Consistent with industry practice, in-memory caching of this kind is not treated as retention under this policy. If you require inference with no caching whatsoever, please contact us.
- Prefix cache. Serving nodes cache the attention state for the shared prefix of successive requests, and report the number of cached tokens in the
usageblock of a response. The prefix cache is shared across all requests to the same model on a serving node, and entries are evicted least-recently-used as memory is needed. Because it is shared, the cached token count of a response can reflect a prefix that was recently sent in another request. - Image cache. The API gateway caches decoded images from recent requests, so multi-turn conversations that resend the same images do not decode them again. Entries are evicted least-recently-used under a fixed memory budget. These images are used only for request routing.
- Routing index. The API gateway keeps one-way hashes of prompt prefixes to route requests to the node most likely to hold them in its prefix cache. The hashes cannot be reversed into Customer Content.
Consistent with industry practice, in-memory caching of this kind is not treated as retention under this policy. If you require inference with no caching whatsoever, please contact us.
7. Subprocessors and Infrastructure
Muna serves the Inference API from a number of hardware and infrastructure providers. Customer Content transits these providers' networks in encrypted form and is processed in memory on compute instances Muna provisions and operates. Muna only provisions serving nodes in professionally operated datacenters, and never on peer-hosted or "community" capacity, where the owner of the machine would have physical access to it. Muna's current subprocessors for the Inference API are:
Muna does not route Inference API requests to any third-party model APIs. All hosted models run on infrastructure Muna operates or controls.
| Subprocessor | Purpose |
|---|---|
| Fly.io | Hosts the inference API gateway, which processes requests in memory and routes them to serving nodes. |
| Vercel | Hosts the Muna website and platform API; does not receive Inference API requests. |
| Tailscale | Encrypted (WireGuard) network between the API gateway and serving nodes. |
| Amazon Web Services | GPU compute for serving nodes. |
| Crusoe | GPU compute for serving nodes. |
| GPU.ai | GPU compute for serving nodes (secure, non-community capacity only). |
| Lambda | GPU compute for serving nodes. |
| Latitude.sh | GPU compute for serving nodes. |
| Novita AI | GPU compute for serving nodes. |
| Oracle Cloud Infrastructure | GPU compute for serving nodes. |
| RunPod | GPU compute for serving nodes (Secure Cloud only). |
| Spheron | GPU compute for serving nodes. |
| Verda | GPU compute for serving nodes. |
| Voltage Park | GPU compute for serving nodes. |
Muna does not route Inference API requests to any third-party model APIs. All hosted models run on infrastructure Muna operates or controls.
8. Exceptions
Muna may retain Customer Content only where:
Muna does not retain Customer Content for automated abuse scanning or trust-and-safety review.
- Legal obligation. Muna is compelled by valid legal process, or is subject to a litigation hold. Where lawfully permitted, Muna will notify you before complying.
- You ask us to. You explicitly enable a feature that requires storage (for example, a stateful or batch feature), or you send content to Muna support for troubleshooting. Such content is retained only for the stated purpose and deleted afterwards.
Muna does not retain Customer Content for automated abuse scanning or trust-and-safety review.
8.1 Diagnostic and Error Logs (Incidental Capture)
Muna operates the Inference API so that Customer Content is not written to logs. In rare failure conditions, however, a component may emit an error that includes a fragment of the data being processed at the time—for example a malformed value that a parser rejected, or token identifiers from a decoding failure. Muna treats any such capture as follows:
Muna does not consider this incidental capture to be retention of Customer Content within the meaning of Section 4, and it does not create a store of prompts or outputs that can be searched, reconstructed, or associated with a conversation.
- Minimised. Error handlers are designed to report the class of failure rather than the data that triggered it. Where a fragment is unavoidably included, it is truncated to the shortest form that is diagnostically useful.
- Short-lived. Diagnostic logs are retained for no more than 30 days and are then deleted automatically, including from backups.
- Access-controlled. Access is limited to engineering personnel investigating a specific incident, is logged, and is not available for bulk search or export.
- Never used for training. Fragments captured this way are never used to train, fine-tune or evaluate any model, and are never disclosed to third parties.
- Reported and remediated. Where Muna identifies a code path that emits Customer Content into logs, Muna treats it as a defect and removes it.
Muna does not consider this incidental capture to be retention of Customer Content within the meaning of Section 4, and it does not create a store of prompts or outputs that can be searched, reconstructed, or associated with a conversation.
9. Data in Transit and Encryption
All requests must use TLS 1.2 or higher. Customer Content is encrypted in transit between your client and Muna, and between Muna's API gateway and its serving nodes.
10. Deletion and Verification
Because Customer Content is never written to persistent storage, there is nothing to delete on request and no deletion SLA applies. Muna will, on request from an authorised account contact:
- provide a written attestation of the controls described in this policy;
- support a customer-led review of these controls as part of a security assessment.
11. Changes
Muna will give at least 30 days' notice before any change that reduces the protections in this policy. Material changes will be reflected in the effective date above.
12. Contact
Questions, attestation requests and data protection enquiries: hi@muna.ai