What Is a Large Language Model? A Plain Explanation
A large language model is a program trained on enormous amounts of text to predict what comes next, one small piece at a time. How that produces useful answers, why it sometimes produces confident wrong ones, and what tools, agents and open models add.
7 min read
A large language model, or LLM, is a computer program trained on a vast amount of text to do one thing: given some text, predict what comes next. It writes an answer by making that prediction over and over, one small piece at a time. Everything else people use it for, from drafting an email to fixing code, comes out of that one skill applied at enormous scale. ChatGPT, Claude, Gemini and Microsoft Copilot are apps built on LLMs; the model is the engine, and the app adds the chat window, your files, memory and connections to other software.
LLM meaning, word by word
- Large: the model has a very large number of internal settings, called parameters, adjusted during training. Google’s machine learning course puts the difference plainly: LLMs contain far more parameters than earlier language models and take in far more context.
- Language: it was trained mostly on text, and it reads and writes text. Many current models also take images, audio or documents, but language is still the core.
- Model: a mathematical function learned from data. It does not store its training text like a library; it stores patterns learned from it.
Tokens: the pieces it reads and writes
An LLM does not see words the way you do. Text is split into tokens, which can be whole words, parts of words, single characters or punctuation. Anthropic’s glossary (opens in a new tab) says a Claude token is roughly 3.5 English characters on average, varying by language. Tokens matter because limits and usage are counted in them, and because the model predicts one token at a time, not one word or one idea.
How a large language model is trained
Training happens in stages. First comes pretraining: the model reads a huge body of text and learns to predict the next word from what came before. Anthropic’s glossary is candid that a model at this stage is not good at answering questions or following instructions. It can continue a document, but it is not yet an assistant.
Then the model is trained further to behave like one. Fine-tuning trains it on more specific examples. Reinforcement learning from human feedback, or RLHF, has people rank sample answers so the model learns to prefer responses like the higher-ranked ones. The result is the helpful, instruction-following assistant you chat with. Training ends on a date, so the model’s built-in knowledge stops there, which is why assistants add web search for anything more recent.
Next-token prediction, with an example
Give a model the text “The capital of Texas is” and it scores every possible next token. “Austin” scores very high, “Houston” lower, “a” lower still. It picks one, adds it to the text, and scores again for the token after that. A thousand-word answer is a thousand or so of those steps.
What makes this work is attention. Google’s machine learning crash course (opens in a new tab) describes self-attention as the model asking, for each token, how much every other token in the input affects its meaning. That is how “bank” comes to mean a riverbank in one sentence and a savings bank in the next, and how an instruction at the top of a long prompt still shapes the last line of the answer.
Picking among likely tokens also explains why the same question can get different answers. A setting called temperature controls how adventurous the choice is, and Anthropic notes that even at the lowest setting, results are not fully identical from one call to the next.
The context window: what it can see right now
An LLM can only take into account the text in front of it at that moment: your instructions, the conversation, any files and tool results. That is its context window, measured in tokens and cleared at the start of every new session. It is why an assistant forgets yesterday’s conversation unless something puts it back. The full picture, including why a bigger window is not the whole answer, is in AI context windows explained.
Why LLMs make things up
An LLM predicts plausible text; it does not look facts up unless it is given a tool to do so. When the plausible answer and the true answer differ, it can state the plausible one with full confidence. People call this hallucination. The US National Institute of Standards and Technology uses the word confabulation, defining it in its generative AI risk profile (opens in a new tab) as confidently stated but erroneous or false content that can mislead users. Google’s course lists it first among the problems with LLMs, alongside their heavy use of computing power and electricity and the bias they can pick up from their training data.
You cannot switch it off, but you can make it rarer and easier to catch: give the model the source material instead of relying on its memory, ask it to quote the passage behind each claim, tell it what to write when something is not in the material, and check anything that matters. How to write prompts that do this is in prompt engineering, and how to review the results is in verifying AI-generated work.
Tools and agents: when an LLM can act
On its own, an LLM only produces text. To search the web, read a spreadsheet or update a task, it is given tools. Anthropic’s documentation on tool use (opens in a new tab) describes the mechanism: the model decides from your request and each tool’s description when to call one, and returns a structured request that the surrounding application carries out, then reads the result. The model asks; the software acts. MCP, the Model Context Protocol, is the common way to plug tools into many different assistants.
Put that in a loop, where the model calls a tool, reads the result and decides the next step until a goal is met, and you have an AI agent. What that means and where it is used is covered in what is agentic AI.
LLM models: open and closed
Some LLMs are closed: you use them only through the vendor’s app or API, and the trained weights stay with the vendor. Claude, Gemini and the GPT models behind ChatGPT work this way. Others are open-weight: the weights can be downloaded and run on your own computer or server, as with OpenAI’s gpt-oss and Google’s Gemma. The Commerce Department’s National Telecommunications and Information Administration studied these in its report on open model weights (opens in a new tab), describing them as models whose weights are open to the public to download, and recommended monitoring their risks rather than restricting them for now.
The trade is control against convenience. An open-weight model keeps data on your hardware, but you run, update and secure it, and you get a model rather than a finished app. How the main assistants and open-weight options compare for work is in ChatGPT alternatives.
A checklist for using an LLM at work
- Give it the material. Paste or connect the source instead of asking from memory.
- Say what to do when something is missing, such as “write not stated”.
- Ask for quotes or links behind factual claims, and open a few.
- Assume nothing carries over between sessions unless you wrote it down somewhere it will read.
- Check whether your account lets the vendor train on your chats, and use a work account for work.
- Give it the fewest tools and permissions the job needs.
- Have a person check anything that is sent, published, paid or deployed.
- Keep the list of what is being done outside the chat, where people and assistants can both see it.
What an LLM will not keep for you
Because each session starts empty, an LLM is a poor place to keep the state of a project. fenbs has no language model of its own; it is the board an assistant reads and writes. Connected over MCP, an assistant calls tools such as fenbs_get_context to read the team’s AI context notes and fenbs_list_items to see the tasks in To Do, Next Up, In Progress and Completed, so it starts each session from what is written down rather than from memory. Every change it makes is recorded in History under its name. fenbs does not check an assistant’s facts for it; the test status and test notes on each task are where a person or an assistant records how the work was checked.
Related
What an AI agent is, in one paragraph: the glossary. How to connect an assistant to a board: Claude and fenbs. Why long sessions degrade: context rot.