Prompt Engineering vs Context Engineering
Prompt engineering shapes the instructions sent to an LLM; context engineering shapes everything else in its input window. How the two differ.
Prompt engineering is the practice of wording instructions to get a better response from a single call to a model; context engineering is the broader practice of deciding what information — documents, tool outputs, conversation history, examples — surrounds that prompt inside the model’s input window. Prompt engineering asks “how do I phrase this?” Context engineering asks “what does the model even need to see to answer this?” As applications built on top of large language models have grown from single Q&A prompts into multi-step agents, the second question has become the harder and more consequential one.
What prompt engineering covers
Prompt engineering is about the instruction itself: word choice, structure, examples, and formatting constraints, all aimed at a single call to the model. Familiar techniques include:
- Few-shot examples — showing the model two or three input-output pairs before asking it to handle a new one.
- Chain-of-thought prompting — asking the model to reason step by step before giving a final answer.
- Role and format instructions — “respond only in valid JSON matching this schema,” or “answer as a senior security engineer.”
- System prompts — persistent instructions that apply across a whole conversation rather than a single message.
These techniques matter because the same underlying model can produce meaningfully different output depending on how a question is framed — not because the model “knows more” with a better prompt, but because framing changes what part of its training the phrasing activates and how strictly it follows a requested format.
What context engineering covers
Context engineering treats the entire input window as a resource to allocate, not just the instruction at the top of it. That includes:
- Retrieved documents — the passages a retrieval-augmented generation pipeline pulls in before the model sees the question.
- Tool and function results — outputs from a database query, a web search, or an API call that get inserted back into the conversation for the model to reason over.
- Conversation history — prior turns, which compete for the same limited context window as everything else.
- Ordering and structure — where information sits in the window and how it’s delimited, since models don’t treat all positions in a long input identically.
A context-engineering decision might be: should this agent see the full document, or a summary? Should the last ten conversation turns be included verbatim, or compressed? Should a tool’s raw JSON response go into the prompt, or should it be filtered down to the three fields that matter? None of these are about wording — they’re about curation and budget.
Why the distinction matters more for agents
A single-shot prompt — “summarize this paragraph” — is almost entirely a prompt-engineering problem: there’s one instruction and one piece of input, and getting the phrasing right does most of the work. A multi-step agent is a different problem. It might call a search tool, read the results, call a database, read those results, and only then answer — and at each step, the accumulated history, tool outputs, and retrieved documents are all competing for the same fixed-size context window.
Get context engineering wrong in an agent and the failure modes are distinctive: the model loses track of an instruction given many turns ago because it scrolled out of relevant attention, or it hallucinates because retrieved documents were truncated mid-sentence, or it repeats a tool call because the previous result wasn’t carried forward clearly. These aren’t fixed by better wording — they’re fixed by deciding what stays in the window and what gets summarized, dropped, or fetched again on demand.
Comparison
| Prompt engineering | Context engineering | |
|---|---|---|
| Unit of work | A single instruction | The full input window across a session |
| Typical techniques | Few-shot examples, chain-of-thought, format constraints | Retrieval, summarization, tool-result filtering, history management |
| Where it matters most | Single-shot completions, classification, extraction | Multi-step agents, long conversations, RAG pipelines |
| Failure mode when done poorly | Wrong tone, wrong format, weak reasoning | Lost instructions, hallucination from truncated context, redundant tool calls |
| Cost sensitivity | Marginal — instructions are usually short | Significant — irrelevant context is wasted tokens on every call |
They aren’t competing disciplines
In practice, the two layer on top of each other rather than replacing one another. A well-engineered prompt still needs the right context to act on; a well-curated context window still needs clear instructions to be used correctly. A RAG system, for instance, is fundamentally a context-engineering problem — deciding what to retrieve and how to chunk it — but the prompt that tells the model how to use the retrieved passages (cite sources, don’t answer if the passages don’t contain the answer) is still prompt engineering layered on top.
Cost is one place the two interact directly. Every token of context — retrieved documents, tool output, conversation history — is a token the model has to process and, in most pricing models, a token you pay for. Our LLM token cost calculator is a quick way to see how context size translates into per-call cost at different context lengths, which makes the trade-off between “include more context for accuracy” and “include less for cost and speed” concrete rather than abstract.
The takeaway
Prompt engineering optimizes what you say to the model in a single call; context engineering optimizes what the model can see across an entire session. Short, single-shot tasks live and die by prompt wording. Anything with retrieval, tools, or multiple turns — most production agents — lives and dies by what gets kept in the context window, what gets summarized, and what gets left out entirely. Building a reliable LLM application means treating both as first-class design problems rather than assuming a better-worded prompt will fix a context problem, or vice versa.
Tagged
Keep reading
Chisato · · 4 min read Tree of Thought vs. Chain of Thought Prompting
Chain of thought asks an LLM to reason in a straight line; tree of thought lets it explore, evaluate, and backtrack across multiple branches.
Chisato · · 6 min read Gemini 3.8 Flash: Benchmarks, Price, and a Cyber Sibling
Google shipped Gemini 3.8 Flash, its third Flash release in six weeks, holding the $0.75 input price and adding a locked-down 3.8 Flash Cyber variant.
Chisato · · 6 min read OpenAI Cuts GPT-5.6 Sol Price 20%: New API Rates
OpenAI cut flagship GPT-5.6 Sol to $4/$20 per million tokens for three months — its first cut to the top tier — to counter Anthropic and Chinese models.