> ## Documentation Index
> Fetch the complete documentation index at: https://daily.dev/llms.txt
> Use this file to discover all available pages before exploring further.

---
title: "Glossary of Key Terms"
url: https://daily.dev/agentic-ai-hub/glossary/
description: "A jargon-buster for the terms used throughout this handbook. Where a full section exists, terms link to it."
lastUpdated: "2026-07-22"
---

A jargon-buster for the terms used throughout this handbook. Where a full section exists, terms link to it.

| Term | Definition |
| :---- | :---- |
| **AEO / GEO (Answer Engine / Generative Engine Optimization)** | Getting content surfaced and *cited* inside AI assistants' answers (ChatGPT, Perplexity, Google's AI answers) rather than ranked in classic search; the AI-era successor to SEO. The emerging [llms.txt](https://llmstxt.org/) convention (a Markdown file pointing models at a site's key content) is one tactic; the term GEO comes from [arXiv:2311.09735](https://arxiv.org/abs/2311.09735). |
| **Agent** | An LLM given a goal, tools, and a loop. It plans, calls tools, observes results, and iterates until done. See [Agentic Engineering](/agentic-ai-hub/agentic-engineering-core-ideas/). |
| **Attention** | The transformer mechanism that lets each token weigh the relevance of every other token in context. |
| **Chain-of-thought (CoT)** | Prompting or training a model to produce intermediate reasoning steps before its final answer, improving accuracy on hard problems. |
| **Context window** | The maximum number of tokens a model can consider at once (prompt + output). Modern models range from ~128K to 1M+ tokens. See [RAG](/agentic-ai-hub/rag-retrieval/). |
| **Distillation** | Training a smaller "student" model to mimic a larger "teacher," which transfers capability at lower cost. |
| **DPO (Direct Preference Optimization)** | A simpler, RL-free alternative to RLHF that tunes a model directly on preferred-vs-rejected response pairs. |
| **Embedding** | A dense numeric vector representing the meaning of text (or image/audio); similar content sits close together in vector space. |
| **Eval** | A structured test measuring model quality on a task or benchmark, used to compare models or guard against regressions. |
| **Few-shot** | Including a handful of examples in the prompt to steer behavior. "Zero-shot" = no examples. |
| **Fine-tuning** | Further training a base model on task- or domain-specific data to specialize its behavior. |
| **Function calling / tool call** | The model emits a structured request (name + JSON args) that your code executes, returning the result to the model. |
| **GGUF** | A file format for storing quantized model weights, used widely by llama.cpp and local-inference tooling. |
| **Guardrails** | Input/output filters and policies that constrain a model by blocking unsafe content, enforcing formats, or checking for injection. |
| **Hallucination** | Confident, fluent output that is factually wrong or fabricated. The core reliability problem of LLMs. |
| **Harness** | The scaffolding around a model: the loop, prompt assembly, tool wiring, and control flow that turns a model into an application or agent. |
| **Inference** | Running a trained model to generate output (as opposed to training it). |
| **Jailbreak** | A prompt crafted to bypass a model's safety training and elicit restricted behavior. |
| **KV cache** | Cached key/value attention tensors from prior tokens, which let generation reuse work instead of recomputing it. This is the main memory cost of long contexts. |
| **Latency** | Time to produce a response. Often split into **TTFT** (time to first token) and per-token generation speed. |
| **Logits** | The raw, unnormalized scores a model assigns to each possible next token before they're turned into probabilities. |
| **LLM-as-judge** | Using a (often stronger) model to grade another model's outputs against criteria. Scalable but imperfect evaluation. |
| **LoRA (Low-Rank Adaptation)** | A parameter-efficient fine-tuning method that trains small adapter matrices instead of full weights. **QLoRA** does this on a quantized base model for even lower memory. |
| **MCP (Model Context Protocol)** | An open standard for connecting models to external tools, data, and resources through a uniform interface. See [MCP](/agentic-ai-hub/mcp-tool-context-ecosystem/). |
| **MoE (Mixture of Experts)** | An architecture that routes each token to a few specialized sub-networks ("experts"). This gives large capacity at lower per-token compute. |
| **Multimodal** | A model that handles more than one modality, e.g. text plus images, audio, or video. |
| **Parameter** | A learned weight in a model. Count (e.g. "70B") roughly indicates capacity and cost. |
| **PEFT** | Parameter-Efficient Fine-Tuning. Umbrella term for methods (LoRA, adapters, prefix-tuning) that adapt a model by training few weights. |
| **Posttraining** | Everything done after pretraining to shape behavior: SFT, RLHF/RLAIF, DPO, etc. |
| **Pretraining** | The initial, compute-heavy phase where a model learns from a massive unlabeled corpus by predicting the next token. |
| **Prompt injection** | An attack where malicious instructions hidden in input (or fetched content) hijack the model's behavior. |
| **Quantization** | Reducing weight precision (e.g. 16-bit → 4-bit) to shrink memory and speed inference, at some quality cost. |
| **RAG (Retrieval-Augmented Generation)** | Retrieving relevant documents at query time and feeding them into the prompt so the model answers from grounded context. See [RAG](/agentic-ai-hub/rag-retrieval/). |
| **Reasoning / test-time compute** | Spending extra compute at inference (longer internal "thinking") to improve answers on hard tasks. This is the basis of reasoning models. |
| **RLHF / RLAIF** | Reinforcement Learning from Human (or AI) Feedback. Aligns a model to preferences using a reward signal from human raters or another model. |
| **SFT (Supervised Fine-Tuning)** | Fine-tuning on curated input→output demonstrations, typically the first posttraining step. |
| **Speculative decoding** | Using a small fast "draft" model to propose tokens that the large model verifies in bulk. This accelerates generation. |
| **Structured output** | Constraining a model to emit valid JSON (or a schema/grammar) so outputs are machine-parseable. |
| **Temperature** | A sampling knob controlling randomness: low = focused/deterministic, high = diverse/creative. |
| **Throughput** | Total work per unit time (e.g. tokens/second across all requests). A serving/capacity metric, distinct from latency. |
| **Token** | The atomic unit a model reads and writes: a word-piece, roughly ¾ of a word in English. Pricing and limits are per-token. |
| **Tokenizer** | The component that splits text into tokens (and back), defining the model's vocabulary. |
| **Top-p (nucleus sampling)** | Sampling only from the smallest set of tokens whose cumulative probability exceeds *p*, which trims the unlikely tail. |
| **TTFT (Time To First Token)** | Latency until the first output token appears; key for perceived responsiveness in streaming UIs. |
| **Vector database** | A store optimized for similarity search over embeddings; the retrieval backbone of most RAG systems. |
| **Zero-shot** | Asking a model to perform a task with no examples, relying purely on its pretrained/instruction-tuned knowledge. |