Semantic Caching
How it works#
- Embed. Your query is embedded into a high-dimensional vector using the configured model.
- Search. A similarity search finds the most relevant cached embeddings using approximate nearest-neighbor lookup.
- Match. If similarity exceeds your threshold (default 0.92), the cached response returns in under 1 ms.
- Miss. On a miss, the request passes through to your origin. The response is cached for next time.
API example#
POST /api/v1/cache/semantic
Authorization: Bearer <API_KEY>
Content-Type: application/json
{
"query": "What is the capital of France?",
"threshold": 0.92,
"ttl": 3600
}
A semantically equivalent query like "France's capital city?" hits the cache.
Multimodal support#
Vector Caching v2 extends semantic caching to images, audio, and mixed media. Each modality gets its own embedding model and vector space.
- Text — natural language queries and completions
- Images — CLIP-based embeddings for visual similarity
- Audio — Whisper-based transcription plus embedding
POST /api/v1/cache/semantic
Content-Type: application/json
{
"modality": "image",
"data": "<base64-encoded-image>",
"threshold": 0.88,
"ttl": 7200
}
SDK snippets#
Python
from cachly import Cachly
client = Cachly(api_key="ck_...")
# Semantic cache lookup
hit = client.semantic.get("What is the capital of France?")
if hit:
print("Cache hit:", hit.value)
else:
# Compute and store
answer = call_llm("What is the capital of France?")
client.semantic.set("What is the capital of France?", answer, ttl=3600)TypeScript
import { Cachly } from "@cachly/sdk";
const client = new Cachly({ apiKey: "ck_..." });
const hit = await client.semantic.get("What is the capital of France?");
if (hit) {
console.log("Cache hit:", hit.value);
}Configuration#
| Parameter | Default | Description |
|---|---|---|
threshold |
0.92 |
Cosine similarity threshold for a cache hit |
ttl |
3600 |
Time-to-live, in seconds |
modality |
text |
Embedding modality: text, image, audio |
model |
auto |
Embedding model override (per-modality defaults) |
Related#
Semantic caching runs on the same Cache Engines that back the AI memory layer. For distributed deployments, see Cluster Mode.