What is Benchmark?
Benchmark: a fixed test suite scoring models on the same tasks so results compare. What good suites measure and where they mislead.
Plain words first
Short definitions of the ideas this journal keeps returning to. 15 of 34 terms published; a new one lands every weekday.
Benchmark: a fixed test suite scoring models on the same tasks so results compare. What good suites measure and where they mislead.
A context window is the stretch of tokens a model can read at once. What sets its size, what happens past the edge, and how builders fit work inside it.
Embeddings turn text, images, or rows into number lists placing similar meanings together. How vectors are made and used for search.
Fine-tuning keeps training a ready model on focused examples for one task or style. What changes, what it costs, prompting vs tuning.
Hallucination: a fluent model answer that sounds true but is not grounded in sources. Why models invent facts and how builders stop it.
Inference turns a trained model plus a fresh prompt into an answer. How serving differs from training, latency and cost budgets.
KV cache reuses past attention keys and values so long chats and agents run faster. What it costs, and how builders shrink it.
Prefix caching reuses stored attention work across requests that share the same opening tokens. How it hits, and how builders keep it hot.
Prompt caching is a provider-billed discount for reusing a stable prompt prefix. What it caches, what it costs, and where the breakpoints go.
Prompt injection hides hostile instructions in data a model reads, hijacking its task. How the trick works and builder defenses.
Quantization stores model weights in fewer bits to fit small GPUs and laptops. Accuracy-vs-speed tradeoffs and when builders use it.
RAG grounds a model in documents retrieved at query time instead of memory alone. How the retrieve-then-write loop works for builders.
A system prompt is the hidden instruction block that sets a model's role, rules, and tone. What it controls and how builders shape it.
Temperature is the sampling dial for token randomness. What low and high settings do, and where builders set it for fact vs flair.
A vector database stores embedding vectors so search finds meaning, not keywords. How similarity indexes work and how builders use them.
Try a different word, or return to the full glossary.