Muna logomuna

Zero Data Retention Policy

Last modified October 9, 2026

Overview

Effective date: September 10, 2026
Applies to: the Muna Inference API at https://inference.muna.ai

This policy supplements the Muna Privacy Policy. Where this policy is more protective of Customer Content, this policy controls.

1. Summary

Muna operates the Inference API on a zero data retention basis. With limited exceptions as described in Section 8, Muna does not store the prompts you send or the outputs the models return. Content submitted to the API is held only in volatile memory for as long as is technically necessary to compute and return a response, and in the in-memory caches described in Section 6.

2. Definitions

"Customer Content" means all data you submit to, and all data returned by, the Inference API, including: prompts, messages, system instructions, tool and function definitions, tool results, model completions, reasoning output, embedding inputs and the resulting vectors, and any images, audio, documents or other files submitted for processing or generated in response.

"Service Metadata" means operational records that do not contain Customer Content. See Section 5.

"Retention" means writing to any persistent medium, including disk, object storage, databases, log aggregation systems, backups, or snapshots.

3. Scope

This policy applies to all models Muna hosts and all inference endpoints, including:

EndpointCovered
POST /v1/chat/completionsYes, including tool use and streaming
POST /v1/messagesYes, including tool use, thinking, and streaming
POST /v1/embeddingsYes
POST /v1/images/generationsYes
GET /v1/modelsYes (contains no Customer Content)
GET /v1/models/{model}Yes (contains no Customer Content)

It applies to every model in the serving catalog, including third-party open-weight models Muna hosts on its own infrastructure.

4. What Muna Does Not Retain

Except in the limited circumstances described in Section 8, Muna does not:

  • write prompts, completions, embedding inputs or output vectors to persistent storage;
  • include Customer Content in application logs, error logs, traces, or crash dumps;
  • include Customer Content in backups or disaster-recovery snapshots;
  • make Customer Content available to Muna personnel.

In no circumstances does Muna:

  • use Customer Content to train, fine-tune, evaluate, or otherwise improve any model (whether Muna's own or a third party's);
  • sell, license, or otherwise disclose Customer Content to any third party.

Other than as described in Section 8, Customer Content exists only in volatile memory: in the serving process for the duration of the request, and in the in-memory caches described in Section 6.

5. What Muna Does Retain

To operate, bill for, and secure the service, Muna retains the following. Other than the diagnostic logs described in Section 8.1, none of it includes Customer Content:

CategoryExamplesRetention
Request metadataTimestamp, request ID, model ID, HTTP status, latency30 days
Usage and billingRequest ID, account ID, model ID, endpoint, prompt / completion / cached token counts7 years, as required for billing and tax
Rate limitingPer-account countersTransient
Infrastructure logsLoad-balancer and edge logs (source IP, path, status, bytes)30 days
Account dataAccount holder name, email, billing detailsPer the Privacy Policy
Diagnostic logsError and crash diagnostics, which may incidentally contain fragments of Customer Content (see Section 8.1)30 days

Token counts are retained; the tokens themselves are not.

6. Caching

To avoid redundant work, the Inference API keeps the following caches. All of them reside in volatile memory only, are never written to persistent storage, and are cleared on process restart:

  • Prefix cache. Serving nodes cache the attention state for the shared prefix of successive requests, and report the number of cached tokens in the usage block of a response. The prefix cache is shared across all requests to the same model on a serving node, and entries are evicted least-recently-used as memory is needed. Because it is shared, the cached token count of a response can reflect a prefix that was recently sent in another request.
  • Image cache. The API gateway caches decoded images from recent requests, so multi-turn conversations that resend the same images do not decode them again. Entries are evicted least-recently-used under a fixed memory budget. These images are used only for request routing.
  • Routing index. The API gateway keeps one-way hashes of prompt prefixes to route requests to the node most likely to hold them in its prefix cache. The hashes cannot be reversed into Customer Content.

Consistent with industry practice, in-memory caching of this kind is not treated as retention under this policy. If you require inference with no caching whatsoever, please contact us.

7. Subprocessors and Infrastructure

Muna serves the Inference API from a number of hardware and infrastructure providers. Customer Content transits these providers' networks in encrypted form and is processed in memory on compute instances Muna provisions and operates. Muna only provisions serving nodes in professionally operated datacenters, and never on peer-hosted or "community" capacity, where the owner of the machine would have physical access to it. Muna's current subprocessors for the Inference API are:

SubprocessorPurpose
Fly.ioHosts the inference API gateway, which processes requests in memory and routes them to serving nodes.
VercelHosts the Muna website and platform API; does not receive Inference API requests.
TailscaleEncrypted (WireGuard) network between the API gateway and serving nodes.
Amazon Web ServicesGPU compute for serving nodes.
CrusoeGPU compute for serving nodes.
GPU.aiGPU compute for serving nodes (secure, non-community capacity only).
LambdaGPU compute for serving nodes.
Latitude.shGPU compute for serving nodes.
Novita AIGPU compute for serving nodes.
Oracle Cloud InfrastructureGPU compute for serving nodes.
RunPodGPU compute for serving nodes (Secure Cloud only).
SpheronGPU compute for serving nodes.
VerdaGPU compute for serving nodes.
Voltage ParkGPU compute for serving nodes.

Muna does not route Inference API requests to any third-party model APIs. All hosted models run on infrastructure Muna operates or controls.

8. Exceptions

Muna may retain Customer Content only where:

  • Legal obligation. Muna is compelled by valid legal process, or is subject to a litigation hold. Where lawfully permitted, Muna will notify you before complying.
  • You ask us to. You explicitly enable a feature that requires storage (for example, a stateful or batch feature), or you send content to Muna support for troubleshooting. Such content is retained only for the stated purpose and deleted afterwards.

Muna does not retain Customer Content for automated abuse scanning or trust-and-safety review.

8.1 Diagnostic and Error Logs (Incidental Capture)

Muna operates the Inference API so that Customer Content is not written to logs. In rare failure conditions, however, a component may emit an error that includes a fragment of the data being processed at the time—for example a malformed value that a parser rejected, or token identifiers from a decoding failure. Muna treats any such capture as follows:

  • Minimised. Error handlers are designed to report the class of failure rather than the data that triggered it. Where a fragment is unavoidably included, it is truncated to the shortest form that is diagnostically useful.
  • Short-lived. Diagnostic logs are retained for no more than 30 days and are then deleted automatically, including from backups.
  • Access-controlled. Access is limited to engineering personnel investigating a specific incident, is logged, and is not available for bulk search or export.
  • Never used for training. Fragments captured this way are never used to train, fine-tune or evaluate any model, and are never disclosed to third parties.
  • Reported and remediated. Where Muna identifies a code path that emits Customer Content into logs, Muna treats it as a defect and removes it.

Muna does not consider this incidental capture to be retention of Customer Content within the meaning of Section 4, and it does not create a store of prompts or outputs that can be searched, reconstructed, or associated with a conversation.

9. Data in Transit and Encryption

All requests must use TLS 1.2 or higher. Customer Content is encrypted in transit between your client and Muna, and between Muna's API gateway and its serving nodes.

10. Deletion and Verification

Because Customer Content is never written to persistent storage, there is nothing to delete on request and no deletion SLA applies. Muna will, on request from an authorised account contact:

  • provide a written attestation of the controls described in this policy;
  • support a customer-led review of these controls as part of a security assessment.

11. Changes

Muna will give at least 30 days' notice before any change that reduces the protections in this policy. Material changes will be reflected in the effective date above.

12. Contact

Questions, attestation requests and data protection enquiries: hi@muna.ai