How large language models work, in plain English
Understand tokens, training, inference and context, and why a fluent AI answer can be wrong. A plain-English guide with a worked example.

A large language model generates text using patterns learned during training and information available in its current input. That combination can produce a clear explanation or a plausible mistake. Fluency tells you how the answer reads; it does not establish whether the answer is true.
The practical distinction is between changing the model and giving an existing model information to work with. Once you understand that, familiar behaviours become easier to interpret: remembering a name, quoting a document or correcting an answer are not necessarily signs of retraining.
Training changes numerical parameters
Text is divided into tokens: chunks that may be words, parts of words or punctuation. A model turns tokens into numerical representations and performs calculations with learned parameters, the numbers that govern its behaviour.
Training adjusts those parameters against an objective. A common language-model objective is to predict the next token from preceding text. Further training can shape instruction-following and other behaviours. This explanation is a starting point; modern training can involve multiple objectives and stages.
Many language models use a transformer architecture. Its attention mechanism computes relationships between parts of the available input. “Attention” names a calculation; it is not evidence of awareness. The original transformer paper describes that architecture in technical detail.
The training process does not create a neat, searchable catalogue of every reliable fact. Learned patterns can support reasoning and explanation while still producing incorrect names, dates or conclusions.
Inference uses the trained model
When you ask a question, the system performs inference: it runs the trained model on the available input. The model produces scores for possible next tokens, a decoding process selects a token, and generation continues.
The available context can include your message, instructions, selected conversation history, documents and tool results. It has limits. A long conversation may not remain available in full, and a product may select or summarise earlier material.
Correcting a name can influence the next answer because your correction is in context. A product may also save a memory. Neither behaviour proves that the underlying model's parameters changed. Data retention and future training use are separate questions about product policy.
A source can be present and still be misread
Consider this fictional library policy:
Ordinary Monday hours: 9am–5pm. Closed on bank holidays.
You ask whether the library opens on a bank-holiday Monday. The model answers “Yes, 9am–5pm” and cites the policy.
The citation exists, but the answer missed the applicable exception. A second answer saying “I am certain” would not fix it. The supported answer is that the library is closed on bank holidays. The policy does not establish the next reopening time.
This small example explains why checking the link is only one step. You must also check that the source supports the claim and that the relevant condition has been applied.
Prompting, retrieval and fine-tuning do different jobs
- 01Prompting: instructions and context
- 02Retrieval: selected source material
- 03Fine-tuning: changed parameters
Prompting supplies instructions and context. A clearer prompt might say to apply exceptions and flag missing information.
Retrieval finds selected source material and supplies it to the model. Retrieval-augmented generation, often shortened to RAG, combines that source-finding step with generation. It can make relevant documents available; it cannot guarantee correct interpretation.
Fine-tuning changes model parameters through additional training. It can adapt behaviour, but it is not automatically the right way to maintain a changing policy library.
For the fictional library, first make sure the right policy reaches the model and test whether it handles exceptions. Buying more complex technology before locating the failure may leave the underlying problem untouched.
The product is larger than the model
A chat application may add web search, calculations, file reading, stored memory or actions in other software. Those surrounding features affect what it can answer and what it can change. Check the actual product, account and permissions.
A model generating a proposed email is different from a connected system sending it. Good writing does not establish that the recipient, attachments or permission to send are correct.
For a practical next step, try the research-brief project. Keep the sources beside the output and explain each correction. Return to how to learn AI for the wider progression, or explore Ampliflow AI Edge.
For the engineering route, Stanford CS336 goes much further into building language models. Its programming and mathematical prerequisites make it a later step for many beginners, not a prerequisite for checking a draft responsibly.