30 AI Terms to Know: MCP, Skills, Agents, Workflows and More

"We gave the agent an MCP server for Jira, wrote a skill for release notes, moved the search into a subagent, and added a hook so it can't touch .env."

That sentence has four AI terms in it, and they are four different kinds of things. One is a connection. One is a folder with a Markdown file. One is a second model run with its own memory. One is a shell script that has nothing to do with the model.

Most glossaries list these words alphabetically, which is how people end up thinking that MCP and skills are competitors. The words make more sense in one picture. Nearly everything an AI coding tool or agent framework does is the same loop: text goes into the model, the model either answers or asks to call something, your code runs that call, and the result goes back in as more text.

Diagram of the agent loop: context window, model, tool call, application code, result returned to the context window
One loop, three kinds of things. Text goes into the window, the model answers or asks for a tool, and plain code runs the call and feeds the result back.

When a new AI word shows up, I sort it into one of three boxes:

  • Text that gets put in front of the model: a prompt, a rules file, a skill, RAG results.
  • A function the model may ask to call: a tool, an MCP server.
  • Plain code that runs around the model: a workflow, a hook, a guardrail, an eval.

The model only knows what is in its window

LLM (large language model). A program that takes text and predicts what text comes next, one piece at a time. ChatGPT, Claude, and Gemini are products built on LLMs. Think of autocomplete that has read a very large library. The longer version is in What Is a Large Language Model.

Token. The piece of text the model actually reads and writes: a short word, part of a long word, a punctuation mark. Limits and prices are counted in tokens, not in characters or words.

Context window. Everything the model can see while producing one answer: the instructions, the conversation so far, the files it read, the tool results. It has a fixed size, measured in tokens. It is the model's desk, not its memory. Whatever is not on the desk right now does not exist for the model, and the model keeps nothing between calls: the application sends the whole conversation again every time.

Diagram of a context window filled with a system prompt, rules file, conversation, tool results and free space
The context window is a fixed-size box. System prompt, rules file, conversation, and tool results all compete for the same space.

Prompt and system prompt. A prompt is the text you send. The system prompt is the part written by whoever built the application: the role, the rules, the tone. You type "fix this bug". The system prompt already said, "You are a coding assistant; never push to main."

Hallucination. A confident answer that is wrong: a library method that never existed, a citation to a paper nobody wrote. The model is not lying. It produces the most plausible next text, and plausible is not the same as true.

Temperature. A setting for how much randomness the model uses when it picks the next token. Low temperature gives steadier, more repetitive output. High temperature adds variety and surprises. Low is not a promise of the same answer twice.

Reasoning model. A model (or a mode of one) that spends extra compute working through the problem before it gives the final answer. In products, it usually shows up as "thinking". Slower and more expensive, but better at multi-step problems.

Knowledge: put it in the window or bake it into the model

A model knows its training data and nothing about your wiki, your codebase, or last week. There are two ways to fix that, and the first three terms here all belong to the same one.

Embedding. A list of numbers that stands for the meaning of a text, built so that texts about the same thing get similar numbers. "Reset password" and "I forgot my login" land close together even though they share no words.

Vector database. A database that stores embeddings and answers one question quickly: which stored texts are closest in meaning to this one?

RAG (retrieval-augmented generation). Search first, then answer. The application finds the relevant documents and pastes them into the context window next to the question. It is an open-book exam: the model did not learn your docs; it is reading them right now. The name comes from a 2020 paper by Lewis and co-authors, and today it is used for almost any "search, then paste into the prompt" setup.

Diagram of RAG: question, search in a vector database, documents added to the context window, answer
RAG does not change the model. It searches first and puts what it found into the window next to the question.

Fine-tuning. More training for an existing model on your own examples, so that its behavior changes: a tone, an output format, a narrow task. RAG changes what the model reads. Fine-tuning changes the model. For "answer questions about our documents", RAG is the usual first step, and how fine-tuning works is a separate topic.

Tools: the model asks, your code runs

Tool (also called function calling). A function you describe to the model: a name, a description, and its parameters. The model cannot run anything. All it can do is answer with a request like this one:

{ "name": "get_weather", "input": { "city": "Warsaw" } }

Your application runs the function and sends the result back as more text. That round trip is the whole trick behind "AI that does things".

MCP (Model Context Protocol). A standard plug for tools. Without it, every AI application needs its own integration with every service. With MCP, you wrap a service once as an MCP server, and any application that speaks MCP (the host, for example, Claude Code or VS Code) connects through an MCP client and gets its tools. A server can offer three things: tools (actions), resources (data to read), and prompts (ready-made templates). Local servers talk over stdio, remote ones over HTTP. MCP doesn't make the model smarter, and it doesn't decide anything. It is wiring.

Read more: What is the Model Context Protocol (MCP) covers the protocol and how to build a server, and the MCP section of AI-Related .NET Interview Questions has the C# SDK and the security risks.

Sequence diagram of a tool call between model, application, MCP client, MCP server and an external service
The model only asks. The application runs the call, and MCP is the standard plug between the application and the service.

A2A (Agent2Agent). A protocol for agents talking to other agents, started by Google and now run under the Linux Foundation. Its own documentation gives the shortest comparison: MCP is for agent-to-tool communication; A2A is for agent-to-agent communication.

Workflow or agent: who picks the next step

Workflow. Model calls arranged by your code on a fixed path. Step one summarizes the ticket, step two classifies it, step three drafts a reply. The model fills in each step, and the code decides the order. Anthropic's definition is "systems where LLMs and tools are orchestrated through predefined code paths".

Agent. A model using tools in a loop until the job is done. You give it a goal ("make the failing test pass"), and it decides to read a file, run the tests, edit the code, and run the tests again. Here the model picks the next step. That makes an agent more flexible and less predictable than a workflow, and more expensive, because every loop turn is another model call.

Side-by-side diagram of a workflow with fixed steps and an agent looping over tools
Same model, same tools. In a workflow, the code picks the next step. In an agent, the model picks it.

Subagent. A helper agent that the main agent starts for a side task, with its own fresh context window. It reads forty files and hands back ten lines. The point is a clean desk for the main agent, not extra intelligence.

Multi-agent (orchestrator and workers). Several agents on one job. Usually one orchestrator splits the task, hands the pieces to workers, and merges what comes back. Every worker is its own loop with its own bill.

Rules files, skills, hooks: three ways to say "always do it like this"

Rules file (AGENTS.md, CLAUDE.md, Cursor rules). A Markdown file in the repository that the tool loads into the context at the start of every session: build commands, conventions, "use pnpm, not npm". AGENTS.md calls itself "a README for agents" and works across many tools. CLAUDE.md is the Claude Code version. Because the file is always loaded, every line in it costs tokens on every request. Keep it short.

Skill. A folder with a SKILL.md file: a name, a one-line description, and the instructions for one kind of task, plus optional scripts and reference files.

release-notes/
├── SKILL.md      # name, description, steps
└── scripts/      # optional helpers

At startup, the agent sees only the name and the description. It reads the rest when a task matches. Picture a box of recipe cards: you can read the label on every card, and you pull one out only when you cook that dish. The format started at Anthropic, was released as an open standard, and is supported by Claude Code, Cursor, GitHub Copilot, Gemini CLI, and others.

Hook. A script that the tool, not the model, runs at a fixed moment: before a tool call, after a file edit, when a session ends. The model does not get a vote. The Claude Code docs draw the line well: an instruction like "never edit .env" in a rules file "is a request, not a guarantee", and a hook that blocks the edit is enforcement. Hooks by this name are a Claude Code feature. Other tools have their own versions.

Plugin. A package that bundles skills, hooks, subagents, and MCP servers, so that a team installs the same setup in one go. Also Claude Code vocabulary.

Diagram showing when a rules file, a skill, a hook and a subagent are loaded or run
A rules file is always in the window, a skill enters on demand, a hook runs outside the model on an event, and a subagent works in a window of its own.

The pairs worth keeping apart:

PairThe difference
Tool vs MCPA tool is one function. MCP is the standard way to deliver tools from a server to any application.
MCP vs skillMCP gives the agent access (to Jira). A skill gives it know-how (how your team writes a Jira ticket).
Rules file vs skillA rules file is always loaded. A skill is loaded when the task needs it.
Skill vs subagentA skill adds instructions to the current window. A subagent does the work in a separate window.
Skill vs hookA skill is read and interpreted by the model. A hook is run by the tool, every time.
Workflow vs agentIn a workflow, the code picks the next step. In an agent, the model does.
RAG vs fine-tuningRAG changes what the model reads. Fine-tuning changes the model.

The words for keeping an agent in check

The more the model decides on its own, the more of this vocabulary you need.

Context engineering. Deciding what goes into the window and what stays out. Prompt engineering is about the wording of the instructions. Context engineering is about the whole contents: which files, which tool results, how much history. It exists because a fuller window is not a smarter model. Anthropic calls the effect "context rot": as the number of tokens grows, the model gets worse at recalling what is in there.

Compaction. When a conversation nears the window limit, the tool summarizes it and continues from the summary. You keep working, and some detail is gone. The other ways to handle long conversations are in AI Conversation History: 4 Strategies.

Prompt caching. The provider keeps the already-processed beginning of your prompt (system prompt, tool definitions, a long document), so repeated requests that start the same way are faster and cheaper. On Anthropic's API, a cached read is billed at about a tenth of the normal input price, and the cache lives five minutes by default. That is one vendor's price list as of October 2026, so check the current numbers.

Prompt injection. Text that tricks the model into following somebody else's instructions. The direct kind is a user typing "ignore your rules". The indirect kind hides the instruction in a web page, an email, or a file that the agent reads. To the model, it is all text in one window, and it cannot reliably tell data from orders. That's why you shouldn't give an agent more permissions than the task needs.

Guardrails. Checks in ordinary code around the model: filter the input, validate the output, block the dangerous command. Code, not polite requests in a prompt.

Human-in-the-loop. A person approves before the agent's action takes effect: a pull request review, a confirmation before a command runs. How much authority to hand over at each stage is what AI Across the SDLC, Part 2 is about.

Evals. Tests for AI behavior. The same prompt can produce different answers, so instead of one assert, you run a set of example tasks, score the results, and rerun the set whenever the prompt or model changes.

Vibe coding. Andrej Karpathy's term from February 2025 for describing what you want and accepting what the model writes, to the point where you "forget that the code even exists". It described weekend prototyping. The phrase now gets used for any AI-assisted coding, which is a different activity when someone still reads the diff.

AI terms Cheat Sheet

AI terms Cheat Sheet

Further reading


Tags:


Comments:

Please log in to be able add comments.