Most prompting guides are lists of tricks for one vendor at one point in time. Put the instruction last. Wrap it in XML tags. Say “you are an expert.” The tricks are not wrong, but they have a short shelf life, because they describe the surface of a particular model in a particular quarter, and the surface moves. A guide written for GPT-3 in 2022 reads today like advice on how to crank a car. I want to do the opposite, and tell the story of where prompting came from, because once you see the one idea underneath it, you can write prompts that work across the models you can actually reach in 2026, whether that is Claude, GPT, DeepSeek V4, or GLM 5.2, and you can write them the way the people building these systems do, from the current research rather than from a cookbook that went stale a year ago. The model names will keep changing; the idea will not.
The idea is attention, in the precise technical sense. And the best way to understand it is to watch it get invented.
THE BOTTLENECK
In 2014, Ilya Sutskever and his colleagues at Google published a way to translate with neural networks that became known as sequence-to-sequence learning. The shape of it was simple. One network, the encoder, reads the source sentence word by word and compresses everything it has read into a single fixed-length vector. A second network, the decoder, takes that vector and unrolls it into the translated sentence. Picture a translator who is allowed to read the source sentence once, then has to set the page face down and produce the translation from memory alone. The whole meaning of the input has to survive in that single held thought.
This worked, and it had an obvious weakness. The longer the sentence, the more that single held thought had to carry, and the more it dropped on the floor. A short sentence was no trouble; a long one was half forgotten by the time the translator reached its end. The researchers who studied it found that translation quality fell off sharply as input length grew, which is the kind of result that tells you the architecture, and not the training, is the thing in the way.
LEARNING WHERE TO LOOK
The fix arrived the same year, from Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. Their move was to let the translator keep the page in view. Instead of forcing the whole sentence through one held thought, the encoder keeps a representation of every input position, and at each step of producing the translation, the decoder looks back over all of them and decides, for this particular output word, which input words matter. They called it learning to align and translate jointly. We call the mechanism attention.
The word is a good one, because that is what it does. For each thing the model is about to say, it computes how much to weight each thing it has read and reads more from the parts relevant to the current decision. When translating the verb of a German sentence, it learns to look at the end of the sentence, where German verbs are placed. The bottleneck is gone, not because the vector got bigger, but because the model stopped relying on a single vector at all and started looking things up on demand.
This is the hinge of the whole story. Attention is the model deciding what to look at. Everything that came after, including the part where you sit down to write a prompt, is downstream of that one decision.
ATTENTION IS ALL YOU NEED
For three years, attention rode atop recurrent networks, an enhancement stitched onto an older engine. Then, in 2017, a group at Google asked an obvious question with a famous answer. If the looking-up is what does the work, do you need the recurrent network underneath at all? Their paper, Attention Is All You Need, threw the recurrence away and kept only attention, stacked in layers, with the model attending to its own input. That architecture is the Transformer, and every model named at the top of this piece is one. DeepSeek V4 and GLM 5.2 add their own tricks to make attention cheaper across a million tokens, but the underlying engine is the same.
So here is the first invariant, the thing that does not rot. When you write a prompt, you are not casting a spell, and you are not negotiating with a mind. You are arranging the text that a Transformer will attend to. What the model can do with your request is bounded by what its attention lands on. Good prompting is almost entirely about getting the right things in front of that attention and keeping the wrong things away from it.
FROM LOOKING TO FOLLOWING
One more step of history, kept short, because it explains why prompting feels the way it does today.
The Transformer made it cheap to train very large language models, and around 2020, the GPT line showed something unexpected: a big enough model, shown a few examples of a task inside the prompt itself, would infer the task and continue it. No fine-tuning, no gradient updates, just examples in the context. This is in-context learning, and it is genuinely strange that it works. It was also brittle. You had to supply the examples, the model was sensitive to their order, and asking for a task in plain words without examples often produced nonsense.
The fix, again, came as a response to the limitation. Researchers at Google showed with FLAN that if you fine-tune a model on many tasks, each phrased as a natural-language instruction, it learns the general pattern of following instructions and can then follow new ones it never saw during training. OpenAI pushed the same idea further with human feedback in InstructGPT, and that model family eventually became ChatGPT. The order matters for understanding what you are doing: few-shot prompting came first, and instruction-following was built afterward to fix how fragile few-shot was. When you type a plain request today, and the model simply does it, you are using the latter invention. When you paste in three examples and let it continue the pattern, you are using the earlier one. Both still work, and a good prompt often uses both at once.
THE METHOD, WHICH IS JUST THE INVARIANT APPLIED
Everything from here is a corollary of a single sentence: the model acts on what its attention reaches within the context window. Each technique below is a way of managing that.
In-context examples, the underrated lever
Few-shot examples are treated as a beginner’s tool, which is a mistake. They are the most direct control you have, because an example does not describe the behavior you want, it demonstrates it, and demonstration lands on attention more cleanly than description. If you want output in a particular shape, two or three worked examples of that exact shape will beat a paragraph of adjectives about it every time.
The craft is in choosing them. For instance, say you are classifying support tickets into “bug”, “billing”, and “feature request”. The temptation is to give three clean, obvious examples. The better move is to spend your examples on the hard cases: the ticket that mentions a charge but is really reporting a bug, the feature request phrased as a complaint. You are not teaching the model what a bug is. You are showing it where the boundaries sit, and the boundaries are where it will otherwise guess. Keep the examples consistent in format, because the model attends to their shape as much as their content, and watch their order, since the research on prompt format sensitivity found that example ordering alone can swing accuracy by a wide margin on smaller models. The stronger 2026 models are less fragile here, but the principle holds: examples are not decoration, they are the part of the prompt the model leans on hardest.
There is a second gear here that most people never shift into. When context windows were small, you could afford a handful of examples, so few-shot meant three or five. With windows now running to a million tokens, you can hand the model hundreds or thousands, and the research on many-shot in-context learning found that accuracy keeps climbing as you do, often past the point where you would otherwise have reached for fine-tuning. Few-shot is showing a new hire three worked invoices before turning them loose on the pile. Many-shot is handing them the last two thousand the team has already processed and letting the pattern teach itself, and the model picks up the boundaries from sheer volume in a way three handpicked cases cannot convey; given enough examples, it will even override a habit it formed during pretraining. Two variants keep this affordable when you do not have thousands of labeled answers lying around: let the model generate the worked reasoning for each example and keep the ones that come out right, or drop the answers entirely and show it only the inputs, which works more often than it has any right to. If this seems to contradict what comes next, that a longer context makes the model worse, hold the thought; the resolution is the whole point.
Structure, and the honest truth about XML
The longer your prompt, the more it helps to give it a visible structure, and the reason is the invariant again. Structure draws clean boundaries, and clean boundaries tell attention where one thing ends and the next begins. A wall of run-on instructions forces the model to infer the seams. Marked sections hand them over for free.
The framework I use is three labeled blocks. Context gives the background to the problem you are solving. Task states what you want in numbered steps when the work has an order. Constraints specify the conditions the answer must satisfy. It is a plain structure, and it has served me well across very different jobs. The shape is easiest to see in the before-and-after. The unstructured version of a request runs everything together:
Look at this support ticket and tell me whether it is a bug, a billing issue, or a feature request, and bear in mind we treat anything mentioning a refund as billing unless it is clearly describing something broken, and give me just one word.
The model can answer that, but first it has to untangle the task from the rule from the output format, all of which are braided into one sentence. The same request, split into Context, Task, and Constraints, gives each its own slot:
Context: We triage incoming support tickets into three queues. A refund mention usually means billing, unless the ticket is really describing something broken, in which case it is a bug.
Task: Classify the ticket below as bug, billing, or feature_request.
Constraints: Answer with one word and nothing else.
Ticket: [...]
Nothing in the second version is cleverer, and the words are almost the same. What changed is that the boundaries are now visible, so the model can spend its attention on the classification instead of working out where the instruction stops and the rule begins.
I should be fair about the part everyone fixes on, which is whether to wrap the blocks in XML tags or in Markdown headers. The honest answer is that it depends on the model, and not in a way you can read off a chart. The study I linked above found one model swinging by up to 76 accuracy points purely due to formatting, and, more usefully, found that the best format for one model is a poor predictor of the best format for another. Anthropic has said its models are tuned to respect XML tags, which is a real fact about Claude and a weak basis for a universal law. So treat the tag syntax as a model-specific surface, and treat the thing it stands for, clear delineation between context and task, and constraints, as the invariant. If you are choosing between frameworks, the named ones in circulation (CO-STAR and its relatives) are all variations on the same move. The move transfers between models; the brackets do not.
One technique to retire: telling the model, “You are an expert lawyer,” to make it answer better. A systematic study of personas in system prompts tested 162 roles across thousands of factual questions and found that adding the persona did not reliably improve accuracy, and sometimes lowered it. A role still shapes tone and vocabulary, which is a fine reason to keep it. It is not a reliability lever, and it was always a slightly magical-thinking one.
CONTEXT ENGINEERING: THE SAME IDEA, STRETCHED OVER A CONVERSATION
So far, the context has been a prompt you write once. In an agent or a long chat, it is a growing transcript, and now the invariant starts to bite in a new way, because attention does not stay sharp as the context grows.
Chroma’s Context Rot report tested 18 current models and found that all of them become less reliable as the input gets longer, even on tasks that remain trivially easy. Two failure modes are worth naming, because they tell you what to do. The first is distractors. The report draws a careful line between a distractor, which is content that is topically close to what you want but does not actually answer it, and merely irrelevant content, which is off-topic. Irrelevant content, the model mostly ignores. Distractors actively pull, because they look relevant to attention. If you ask for the founding date of a company and the context is full of other dates for other companies, the near-misses are what hurt you, and a second distractor hurts more than the first. The second mode is position: the classic lost-in-the-middle result, where models attend well to the start and end of a long context and let the middle go soft.
This is also the resolution I promised in the many-shot puzzle. Two thousand relevant examples and two thousand tokens of unrelated chatter are both long contexts, and they pull in opposite directions, because the examples are signals the model can lock onto, and the chatter is noise it has to push past. Follow-up work on many-shot found that the gain comes mainly from the model picking out the relevant demonstrations and ignoring the rest, which is the distractor result wearing different clothes. A longer context is not automatically worse; a noisier one is, and the job is to make sure that when the context grows, it grows with signal.
This is what context engineering is for, and it is the same job as prompting, performed over time: keep attention pointed at what matters and keep the distractors out. The basic moves follow directly. Compaction summarises the older turns into a short running brief and keeps only the last several turns verbatim, on the reasoning that recent turns deserve full attention and a twenty-turn-old exchange deserves a sentence. Memory systems go further. Mem0, for instance, runs the conversation through two phases: it extracts salient facts from each exchange, using the latest turn, a rolling summary, and the last few messages, and then it decides for each candidate fact whether to add it, update an existing one, delete a contradicted one, or do nothing. At the next turn, it pulls back only the stored facts relevant to the current question. The paper reports large savings in tokens and latency against stuffing the whole history in, which is the expected result once you accept that a longer context is not a richer context but a noisier one. None of this is glamorous, and it is not shiny like demoing a new chatbot. It is the work that makes a long-running agent stay coherent.
REASONING MODELS CHANGE THE JOB
Here is what is genuinely new in 2026, and where a guide written two years ago goes wrong.
The advanced prompting technique of 2022 was chain-of-thought: append “let’s think step by step” and watch multi-step accuracy jump, because the model now spends tokens reasoning before it answers (Wei et al.). That trick still describes something true about how these models work. What changed is that the major labs took the trick inside the model. DeepSeek showed with R1 that you can train the reasoning in directly, and the current generation ships it as a setting rather than a prompt. DeepSeek V4 has its thinking mode on by default. GLM 5.2 exposes a High and a Max reasoning effort. Claude has extended thinking; the OpenAI o-series hides the reasoning entirely. The model is already thinking step by step. You no longer have to ask, and on these models, asking can hurt, because you are hand-writing a worse version of a procedure the model would have run better on its own, and crowding its context while you do it.
So the job moves up a level. With a reasoning model, you do not supply the steps, you supply a clear brief and let it find the steps: state the goal rather than the procedure, give it the context and constraints, name the audience and the shape of the output, and set the reasoning budget to fit the problem. A hard architectural question wants max effort or extended thinking. A one-line lookup does not, and spending a large reasoning budget on it just buys you latency and the occasional model that talks itself out of the right answer, which is a real failure mode the reasoning models introduced. The instinct to write longer, more procedural prompts gets this generation exactly the wrong way round. The better prompt for a thinking model is often shorter and more declarative than the one you would have written for the same task in 2023.
The older reasoning techniques have not vanished, they have moved underneath. Self-consistency, from Wang et al., samples several independent reasoning paths and takes the majority answer, on the sound intuition that a single greedy chain can wander, whereas several chains rarely wander to the same wrong place. ReAct interleaves reasoning with tool calls, so the model can look something up instead of guessing. Tree of Thoughts lets it branch and evaluate several lines before committing. You can still build these by hand, and sometimes you should. But increasingly they are what the lab has already wired into the thinking mode you toggled on, which is the same pattern as instruction-following: a prompting trick discovered by users, then absorbed into the model.
One technique deserves a warning rather than a recommendation because it is the one everybody reaches for, and the research has been unkind to it. The intuitive move is to let the model check its own work: produce an answer, then ask it to find and fix its mistakes. A pointed study from 2024 found that when a model tries to correct its own reasoning with nothing but its own judgment to go on, it does not reliably improve, and it sometimes makes things worse, talking itself out of an answer that was right the first time. The diagnosed reason is almost funny: the hard part is not fixing the error, it is noticing it, and a model is confidently blind in precisely the places it was already wrong, like a student grading their own exam with no answer key and ticking every box. What rescues the idea is an outside signal. Give the model a test suite, a calculator, a search result, a verifier, anything it did not generate itself, and the loop starts working, which is the quiet reason ReAct earns its keep: the tool call is the external check the model cannot produce from inside its own head.
WHERE THE FRONTIER ACTUALLY IS: OPTIMISING THE PROMPT
Everything so far assumes you write the prompt by hand, read the output, and adjust. That is how most people still work, and it is not how the strongest teams work anymore. Inside a frontier lab, nobody ships a hand-tuned system prompt and calls it finished, for the same reason no serious mechanic tunes a carburetor by ear when there is a dyno in the next room: human intuition is a fine first guess and a poor optimizer. The discipline that grew up around this in 2025 treats the prompt as something you fit against a measurement rather than something you compose.
The tool that made the idea concrete is DSPy, which lets you declare what each step of a model program should do and then hands the wording to an optimizer. The result that made people pay attention is GEPA, from the middle of 2025: it runs your system, reads back its own failed traces in plain language, works out what went wrong, proposes a better prompt, and keeps the variants that survive on a held-out set, evolving the wording the way you would breed a plant rather than engineer a part. The paper reports that it beats a reinforcement-learning baseline by a clear margin while requiring far fewer trial runs, and beats the previous optimizer by double digits. You do not have to adopt the framework to take the lesson. The lesson is that the moment you have an evaluation set, even a small one of a few dozen labeled examples, the prompt stops being a sentence you craft and becomes a parameter you tune, and your job moves from writing the prompt to writing the test the prompt has to pass. Andrej Karpathy gave this its name in 2025 when he called natural-language programming "Software 3.0," and the part worth keeping from the slogan is mundane and correct: if the prompt is the program, it deserves what every other program gets: a test and a way to improve against it.
This is the widest gap between someone who has read a prompting guide and someone who builds with these models for a living: the first polishes wording by hand and trusts their taste, the second writes an evaluation and lets a loop out-write that taste, then spends their own attention on the evaluation, which is the part a model cannot yet do for them.
WRITING FOR 2026, NOT FOR ONE VENDOR
The reason to learn the history is that it hands you the parts that stay still while everything visible moves. Bahdanau’s insight that a model should decide what to look at is doing the same job in a model reading a million tokens today that it did in a translator reading one sentence in 2014. The Transformer that turned attention into the whole engine is still the engine under every model named here. The fact that a longer context is a noisier context, and that your job is to keep attention on the signal, did not change when the context windows grew to a million tokens; it got more important.
The next GLM and the next GPT will ship before this piece is old, and they will come with a new set of vendor-specific tips, most of which will be a particular phrasing of something here. If you write to the model in front of you, you relearn how to prompt every quarter. If you write to the invariant, arrange the context, spend your examples on the hard boundaries, keep the distractors out, let a thinking model think, and once you can measure the task, hand the wording to an optimizer, you write prompts that survive the next release. That is the real advantage of acting on first principles, and it is worth building a habit on. It also happens to be how the people who build these models prompt them, which means the gap between the cookbook reader and the lab engineer was never about secret tricks; it was about working from the mechanism instead of the surface, and in this article, I gave my best to help you gain intuition on the underlying mechanism.


