FinanceGadget
Guide

Zero Data Retention in AI Tools: What It Means and Who Offers It

The short answer

Zero Data Retention (ZDR) is an enterprise API capability where a model provider processes your prompt in volatile memory (RAM) during inference and purges it immediately upon sending the response. Under a verified ZDR agreement, the vendor does not write prompts, completions, embeddings, or fine-tuning artifacts to persistent disk storage or 30-day abuse monitoring queues. ZDR is currently offered by OpenAI (for eligible enterprise API endpoints), Anthropic (via API ZDR requests), AWS Bedrock, and Google Cloud Vertex AI.

Why default API access is not zero retention

When an organization connects to a commercial AI model API using standard pay-as-you-go credentials, data is rarely discarded immediately. Most major providers maintain a 30-day abuse monitoring retention window by default.

Under standard commercial API terms:

  1. Input Ingestion: Your HTTP request (containing code, context, or customer data) is decrypted at the API gateway.
  2. Short-Term Disk Storage: Prompts and completions are logged to encrypted storage volumes for up to 30 days to facilitate automated abuse detection filters and post-hoc security investigations.
  3. Authorized Human Review: If an automated filter flags a request for safety violations, authorized vendor security personnel may inspect the cached prompt text.

For enterprise legal and security teams handling financial records, health data (GDPR special categories), or strict client non-disclosure agreements (NDAs), a 30-day vendor storage window creates an unacceptable data liability.

How Zero Data Retention (ZDR) works under the hood

When ZDR is configured and approved for an enterprise endpoint, the execution path changes fundamentally.

Standard API Flow:
Client ---> API Gateway ---> Volatile Inference (GPU) ---> Response
                                  |
                                  +---> 30-Day Disk Log (Abuse Queue)

ZDR Execution Flow:
Client ---> API Gateway ---> Volatile Inference (GPU) ---> Response
                                  |
                                  X (Logging bypassed / RAM purged)

Under a ZDR architecture:

  • Volatile Processing: Prompts are decrypted into GPU/TPU high-bandwidth memory (HBM) strictly for tensor computation.
  • Immediate Memory Purge: As soon as token generation concludes and the HTTP response stream closes, the memory buffer allocated to that request context is overwritten.
  • Disk Bypass: The API logging pipeline is programmatically bypassed. No record of the prompt payload or completion text is written to disk, database tables, or central log aggregators.
  • Metadata-Only Logging: Vendors maintain only non-sensitive HTTP transport metadata (timestamp, HTTP status code, token count billed, and client IP address) for billing and rate-limiting.

Provider comparison: Who actually provides ZDR?

Not all ZDR claims are identical. Vendors enforce different eligibility criteria, request processes, and endpoint restrictions.

1. OpenAI API

  • ZDR Availability: Available for enterprise API customers upon application.
  • Scope: Applies to select endpoints (e.g. gpt-4o, gpt-4o-mini, embeddings).
  • Exceptions: Audio, vision, and custom fine-tuning endpoints may retain separate logging windows. Standard consumer ChatGPT Free, Plus, and Pro plans do not support ZDR.
  • Reference: See our detailed OpenAI Provider Dossier.

2. Anthropic API (Claude)

  • ZDR Availability: Available for API commercial customers via explicit Zero Data Retention agreements.
  • Scope: Prompts and completions processed via the Anthropic API are not stored on persistent disk when ZDR is active.
  • Exceptions: Requests originating from the consumer claude.ai web interface or unauthenticated free tiers are excluded.
  • Reference: See our Anthropic Provider Dossier.

3. AWS Bedrock

  • ZDR Availability: Built into the core infrastructure posture by default.
  • Scope: Under AWS Service Terms Section 50, AWS Bedrock does not store or log customer prompts or responses unless the customer explicitly configures Amazon S3 log delivery or CloudWatch logging.
  • Reference: See our AWS Bedrock Dossier.

4. Google Cloud Vertex AI

  • ZDR Availability: Standard feature for Cloud Vertex AI customers.
  • Scope: Customer prompts, generated text, and fine-tuning datasets are stored only within the customer’s allocated Google Cloud project resources. Vertex AI API endpoints do not retain prompt payloads outside customer control.
  • Reference: See our Google Cloud Dossier.

What to verify before relying on ZDR

Before declaring ZDR compliance to your internal security board or external auditors, confirm the following technical details:

  1. Verify Endpoint Scope: Confirm whether ZDR applies universally across all model endpoints or only to specific text models. Image generation, speech-to-text, and third-party web search plugins frequently route through separate pipelines without ZDR guarantees.
  2. Review Third-Party Sub-processors: Check whether the vendor uses external cloud hosts (e.g. specialized GPU cloud providers). Ensure the vendor’s sub-processor contracts pass ZDR requirements down to infrastructure partners.
  3. Implement Client-Side DLP: ZDR prevents vendor storage, but it does not stop malicious prompt injection or unintended data disclosure during live session interaction. Combine ZDR with client-side ignore rules using our AI Extension Ignore File Generator.
Advertisement