Every question makes the next one cheaper. The Ocean is a vector memory that sits in front of inference: embeddings are searched first, hits are served straight from memory, misses are answered fresh β and then remembered, hash-chained, witnessed.
Every answer carries its provenance: β‘ served from the ocean, or π§ fresh inference β now remembered. Every call gets a hash-chained witness receipt.
The most recent calls across the whole ocean, rendered as a tide. Each dot's position is deterministic β derived from its timestamp and prompt. Hover any dot.
Every prompt is embedded at the edge. The question becomes a vector before it becomes an answer β the same embedding model every time, so the ocean stays one geometry.
The vector is searched against every memory the ocean holds. Top hit's similarity is compared against the serve threshold β Vectorize does this in milliseconds.
Above threshold: the remembered answer is served straight back β no model call, tokens saved. Below: a real model answers, and the exchange is written into the ocean so the next asker gets the cheap path.
Every call β hit or miss β mints a hash-chained witness receipt. The chain is append-only and verifiable; you can audit exactly what the ocean served and when.
# ask β the serve decision comes back with the answer curl -X POST https://quilt-cloudflare.superinstance.workers.dev/api/ocean/ask \ -H 'Content-Type: application/json' \ -d '{"prompt":"what is a quilt cell?"}' # β {"answer":"β¦","served":"ocean","similarity":0.94,"witness":"a1b2c3d4β¦", β¦} # live counters curl https://quilt-cloudflare.superinstance.workers.dev/api/ocean/stats # the recent tide curl https://quilt-cloudflare.superinstance.workers.dev/api/ocean/recent