The direct answer
When people say "GPT," the T stands for transformer. A transformer is a type of neural network — software that learns patterns from data — introduced in a 2017 Google research paper called "Attention Is All You Need." Almost every major AI assistant today, from ChatGPT to Claude to Gemini, runs on this architecture.
The core idea is simple: instead of reading text one word at a time, a transformer looks at the whole passage at once and works out which words relate to which. That one design choice made it possible to train AI on enormous amounts of text — and everything since has built on it.
How it works
Text enters a transformer as tokens — small chunks of words. Each token is converted into a list of numbers (an embedding) that captures its meaning. Then comes the key step: self-attention. For every token, the model computes a score against every other token, measuring how much each one should influence the others.
These scores are computed by multiple "attention heads" working in parallel — different heads can track different kinds of relationships, like grammar, meaning, or word order — and the results pass through dozens of layers, each refining the picture. Because attention compares all words at once instead of one after another, the work can be split across thousands of GPU chips simultaneously. That parallelism is what made training giant models practical in the first place.
At the end, the model predicts the next token, adds it to the sequence, and repeats — generating one chunk at a time until the answer is complete.
A simple example
As an illustration, take this sentence: "The bank refused the loan, so he sat on the bank of the river." A transformer deciding what "bank" means in each spot assigns strong attention between the first "bank" and "loan," and between the second "bank" and "river." Same word, two meanings, sorted out purely by relationships between words.
This is exactly what the researchers demonstrated in their own write-up, where the word "bank" in "I arrived at the bank after crossing the river" attends directly to "river" — resolving the meaning in a single step, without reading the words in between one by one.
Why it matters
Before transformers, language AI read text sequentially — like a person reading with a finger under each word. That was slow to train and bad at connecting ideas far apart in a passage. Transformers process everything in parallel, so they could be trained on vastly more text in reasonable time. That jump is what made capable AI assistants possible.
The same design turned out to work beyond text: researchers have since applied it to images, audio, and video. It also made the model's behavior partly visible — attention scores can be visualized, showing which words the model "looked at" when translating a sentence. That inspectability is part of why researchers keep studying and trusting the architecture.
The common misunderstanding
The common one: that a transformer understands text the way you do. Self-attention computes similarity scores between chunks of text — it is arithmetic, not comprehension. The model has no awareness of what a river actually is; it has learned that the token "river" tends to appear near certain other tokens.
That is powerful enough to produce fluent, useful answers, but it is pattern-matching at scale, not understanding. Keep that distinction in mind and a great deal of AI hype becomes legible.
What changed recently
Since 2017, the transformer has gone from research curiosity to default infrastructure. The big shifts since then are efficiency and range: engineers keep finding ways to make attention cheaper so models can handle much longer documents, and the architecture now routinely processes images and audio alongside text — Google describes its Gemini models as built multimodal from the ground up.
The core idea hasn't changed, though. A transformer from today would still be recognizable to its 2017 inventors: the same attention mechanism, just bigger, faster, and fed more kinds of input.
Try it on PlainLogic
Want to feel how a transformer reacts to wording? The AI Lab on PlainLogic opens up how models process language — and the Prompt Repair game lets you experiment with how small changes in phrasing change what an AI produces.