What is Prompt caching?
Prompt caching is a provider-billed discount for reusing a stable prompt prefix. What it caches, what it costs, and where the breakpoints go.
Ideas, collected
Essays, guides and observations. Find something worth sitting with.
Prompt caching is a provider-billed discount for reusing a stable prompt prefix. What it caches, what it costs, and where the breakpoints go.
Prefix caching reuses stored attention work across requests that share the same opening tokens. How it hits, and how builders keep it hot.
KV cache, prefix caching, prompt caching, and semantic cache compared: what each reuses, what it saves, and how local runners configure it.
Jev vs Laya vs Cua-S1-Forms: when scoring fixed options beats sampling text, and which open model to self-host.
KV cache reuses past attention keys and values so long chats and agents run faster. What it costs, and how builders shrink it.
Open weights vs open source, explained for builders: what you can download, the VRAM math, and which licences let you self-host.
AI news digest Sep 19: Astra for Law legal index, TypeSafe Jev model, 706K-param forms planner, kill-switch order, 27B in 5.9GB.
Build retrieve-then-write locally: chunking, a resident embeddings model, a JSONL index to Chroma/Qdrant, cited grounded answers.
From 16-bit weights to 4-bit quants: what quantization does, the VRAM math per level, and a practical which-quant-for-which-card table for 4 to 16 GB GPUs.
Kilo Code vs self-hosted Coder Agents vs Copilot agent mode: which local-friendly coding agent fits solo builders and teams.
Which embedding models fit beside a chat quant on a 6GB card, real VRAM costs, and three local runners: Ollama, llama.cpp, Python.
Six builder stories: launches, safeguards, open models, agents.
What really runs in 6 GB VRAM on an RTX 2060: quantized 7–8B models, partial offload for 13B, and keeping embeddings hot on the GPU.
Try a different word, or return to the full collection.