Live · claude-sonnet-4-6 · 86.9% cache hit rate
Your LLM API bill is twice what it needs to be.
promptzip is a drop-in proxy that automatically caches prompts, compresses JSON payloads, and deduplicates repeated queries. Two-line integration. Real savings from the first request.
Without promptzip
$74
per day · 10,000 requests
2,440 tokens × 10k × $3.00/1M
51% less
With promptzip
$36
per day · same 10,000 requests
320 uncached × $3.00/1M
2,120 cached × $0.30/1M
2,120 cached × $0.30/1M
Measured live · claude-sonnet-4-6 · 32 requests · 2,440-token system prompt · June 2026
Integration
Two lines of code. That's the entire migration.
Change the base URL and swap in your pz- key. The proxy intercepts every request, injects cache_control, substitutes your real LLM key, and forwards it transparently.
Prompt cache injection
Detects system prompts and adds Anthropic cache_control breakpoints. Every repeat request pays 10× less.
JSON compression
Minifies large JSON blocks embedded in messages — RAG chunks, tool outputs — before they reach the model.
Semantic dedup
Paraphrased queries (cosine ≥ 0.92) return cached responses. No inference, no cost.
your_app.py
# before
import anthropic
client = anthropic.Anthropic("sk-ant-...")
# after — only these two lines change
import anthropic
client = anthropic.Anthropic(
api_key = "pz-...",
base_url = "https://proxy.promptzip.io"
)
# req 1 → cache write $0.01099
# req 2+ → ✓ cache read $0.00257▍
How it works
Three layers. All transparent.
1
Prompt cache injection
Every system prompt gets cache_control breakpoints inserted before the request reaches Anthropic. The model caches the prefix — repeat calls pay the cached rate ($0.30/1M vs $3.00/1M).
Up to 90% reduction on repeat traffic
2
JSON compression
Scans user messages for embedded JSON blobs — retrieved documents, tool outputs, structured payloads — and minifies them in-place. Identical content, fewer tokens billed.
20–40% reduction on data-heavy prompts
3
Semantic response cache
Incoming queries are embedded and checked against pgvector. When a past response scores ≥ 0.92 cosine similarity, it's returned directly. No model call, no API cost.
Eliminates redundant inference entirely
Benchmark
Real test. Real numbers.
claude-sonnet-4-6 · 32 requests · 2,440-token system prompt · measured live today
proxy.infrajump.com
# input write read out cost
────────────────────────────────────────────
1 14 2440 0 120 $0.01099 return policy for software licenses
2 13 0 ✓2440 120 $0.00257 how long does EU shipping take?
3 15 0 ✓2440 120 $0.00258 what payment methods do you accept?
4 15 0 ✓2440 120 $0.00258 how do I track my order?
5 14 2440 0 120 $0.01099 return policy (cache expired, rewrite)
6 13 0 ✓2440 120 $0.00257 how long does EU shipping take?
...
32 17 0 ✓2440 120 $0.00258 API rate limit errors
────────────────────────────────────────────
total: 78,639 tokens · 68,320 cached (86.9%) · $0.116 actual
Direct to Anthropic
0%
tokens cached
2,440 × 10,000 × $3.00/1M
= $73.20 / day
= $73.20 / day
Through promptzip
86.9%
tokens cached
320 uncached × $3.00/1M
2,120 cached × $0.30/1M
= $36.00 / day
2,120 cached × $0.30/1M
= $36.00 / day
$22,515 / year saved
$1,876/mo · typical support bot · 10k req/day
Pricing
Pay 25% of what you save. Nothing if you save nothing.
25%
of measured savings, billed monthly
$2,000 saved$500 billed· you keep $1,500
You save $500/mo$125/mo
You save $2,000/mo$500/mo
You save $10,000/mo$2,500/mo
You save $0$0
What's included
OpenAI, Anthropic, Gemini, Groq
Prompt cache injection — automatic
JSON compression
Semantic response cache (pgvector)
Real-time savings dashboard
Monthly invoice via Stripe
No contracts — cancel any time
First $500 saved is free
No credit card required
Stop overpaying for LLM APIs.
Two lines of code. Savings in the first hour. We charge only when you save.