Live · claude-sonnet-4-6 · 86.9% cache hit rate

Your LLM API bill is twice what it needs to be.

promptzip is a drop-in proxy that automatically caches prompts, compresses JSON payloads, and deduplicates repeated queries. Two-line integration. Real savings from the first request.

Without promptzip
$74
per day · 10,000 requests
2,440 tokens × 10k × $3.00/1M
51% less
With promptzip
$36
per day · same 10,000 requests
320 uncached × $3.00/1M
2,120 cached × $0.30/1M

Measured live · claude-sonnet-4-6 · 32 requests · 2,440-token system prompt · June 2026

Integration

Two lines of code. That's the entire migration.

Change the base URL and swap in your pz- key. The proxy intercepts every request, injects cache_control, substitutes your real LLM key, and forwards it transparently.

Prompt cache injection
Detects system prompts and adds Anthropic cache_control breakpoints. Every repeat request pays 10× less.
JSON compression
Minifies large JSON blocks embedded in messages — RAG chunks, tool outputs — before they reach the model.
Semantic dedup
Paraphrased queries (cosine ≥ 0.92) return cached responses. No inference, no cost.
your_app.py
# before
import anthropic
client = anthropic.Anthropic("sk-ant-...")
# after — only these two lines change
import anthropic
client = anthropic.Anthropic(
api_key = "pz-...",
base_url = "https://proxy.promptzip.io"
)
# req 1 → cache write $0.01099
# req 2+ → ✓ cache read $0.00257
How it works

Three layers. All transparent.

1
Prompt cache injection
Every system prompt gets cache_control breakpoints inserted before the request reaches Anthropic. The model caches the prefix — repeat calls pay the cached rate ($0.30/1M vs $3.00/1M).
Up to 90% reduction on repeat traffic
2
JSON compression
Scans user messages for embedded JSON blobs — retrieved documents, tool outputs, structured payloads — and minifies them in-place. Identical content, fewer tokens billed.
20–40% reduction on data-heavy prompts
3
Semantic response cache
Incoming queries are embedded and checked against pgvector. When a past response scores ≥ 0.92 cosine similarity, it's returned directly. No model call, no API cost.
Eliminates redundant inference entirely
Benchmark

Real test. Real numbers.

claude-sonnet-4-6 · 32 requests · 2,440-token system prompt · measured live today

proxy.infrajump.com
# input write read out cost
────────────────────────────────────────────
1 14 2440 0 120 $0.01099 return policy for software licenses
2 13 0 ✓2440 120 $0.00257 how long does EU shipping take?
3 15 0 ✓2440 120 $0.00258 what payment methods do you accept?
4 15 0 ✓2440 120 $0.00258 how do I track my order?
5 14 2440 0 120 $0.01099 return policy (cache expired, rewrite)
6 13 0 ✓2440 120 $0.00257 how long does EU shipping take?
...
32 17 0 ✓2440 120 $0.00258 API rate limit errors
────────────────────────────────────────────
total: 78,639 tokens · 68,320 cached (86.9%) · $0.116 actual
Direct to Anthropic
0%
tokens cached
2,440 × 10,000 × $3.00/1M
= $73.20 / day
Through promptzip
86.9%
tokens cached
320 uncached × $3.00/1M
2,120 cached × $0.30/1M
= $36.00 / day
$22,515 / year saved
$1,876/mo · typical support bot · 10k req/day
Start for free →
Pricing

Pay 25% of what you save. Nothing if you save nothing.

25%

of measured savings, billed monthly

$2,000 saved$500 billed· you keep $1,500
You save $500/mo$125/mo
You save $2,000/mo$500/mo
You save $10,000/mo$2,500/mo
You save $0$0

What's included

OpenAI, Anthropic, Gemini, Groq
Prompt cache injection — automatic
JSON compression
Semantic response cache (pgvector)
Real-time savings dashboard
Monthly invoice via Stripe
No contracts — cancel any time
First $500 saved is free
Get started free

No credit card required

Stop overpaying for LLM APIs.

Two lines of code. Savings in the first hour. We charge only when you save.