Articles

What Is a Foundation Model in AI?

A foundation model is a large model pretrained on broad data, then adapted for many downstream tasks via fine-tuning, RAG, or prompting alone.

Chisato Chisato · · 4 min read
Abstract purple neural network fibers

A foundation model is a model trained on a broad swath of data at large scale, general enough that it can be adapted to many different downstream tasks rather than being built for one job from scratch. The term describes a role, not an architecture: a large transformer trained on internet-scale text is a foundation model, and so is a vision model pretrained on billions of images. What makes it “foundational” is that other systems get built on top of it instead of starting over.

From narrow models to general-purpose bases

Before foundation models became the norm, machine learning teams typically trained a separate model per task: one for sentiment analysis, another for translation, another for image classification, each with its own labeled dataset and architecture. That approach scales poorly — every new task means collecting fresh labeled data and training from a blank slate.

Foundation models flip that. A single large model is pretrained once, at significant cost, on a broad and largely unlabeled corpus. The resulting model captures general patterns in language, code, or images that turn out to transfer well to tasks the model was never explicitly trained on. Instead of training many small models, teams adapt one large one.

Pretraining: where the generality comes from

Pretraining is typically self-supervised: the model learns by predicting missing or next pieces of its own training data, without needing human-labeled examples. A text model predicts the next token; the sheer scale and diversity of the corpus is what makes the resulting representations useful for far more than “predict the next word.”

This is also why what an LLM is and what a foundation model is overlap so heavily — a large language model is the most common example of a foundation model, but the category is broader. Vision-language models, audio models, and multimodal systems that handle several input types at once are foundation models too; see what a vision-language model is for one variant.

Adapting a foundation model to a specific task

A foundation model on its own is a general-purpose base — useful, but not necessarily aligned to any particular product. Three common ways to specialize it:

  • Prompting. Give the model instructions and examples directly in the input, with no weight updates at all. This works because pretraining already exposed the model to enough varied text that it can follow new instructions through in-context learning rather than retraining.
  • Fine-tuning. Continue training the model on a smaller, task-specific dataset, updating its weights so it specializes toward that task or that output style. This is more expensive than prompting but produces more consistent, narrower behavior.
  • Retrieval augmentation. Leave the model’s weights untouched and instead feed it relevant documents at query time, so its answers are grounded in a specific knowledge base. See retrieval-augmented generation for how that pipeline works.

These aren’t mutually exclusive — a production system might combine a fine-tuned model with a retrieval step and careful prompting on top.

Foundation model, LLM, and small language model

Foundation modelSmall language model
ScaleTrained on broad, large-scale dataTrained or distilled to run with far fewer parameters
GoalMaximum generality and transferEfficiency for a narrower set of tasks
DeploymentOften via API, or self-hosted with heavy hardwareCan run on modest hardware, sometimes on-device
Typical useGeneral assistants, base for further tuningFocused tasks where latency or cost matters more than breadth

A small language model is usually produced by distilling or otherwise shrinking a foundation model down, trading some generality for speed and lower cost. Both are still “foundation models” in the sense of being reused across tasks rather than trained per task — the distinction is mostly about scale and deployment footprint, not the underlying training approach.

Why the metaphor — and its risks — matter

Calling these models “foundational” is deliberate: it signals that flaws in the base model propagate to everything built on it. A bias or blind spot introduced during pretraining doesn’t stay contained — it shows up in every fine-tuned variant and every application built on top, unless specifically corrected downstream. That’s one reason model documentation matters: an AI model card is meant to disclose what a foundation model was trained on, how it was evaluated, and where its known limitations are, so downstream teams aren’t discovering them in production.

It’s also why a small number of foundation models end up underpinning a large share of AI applications. Training one from scratch requires enormous compute and data; most organizations adapt an existing foundation model rather than building a new one. That concentration is efficient, but it means an outage, a licensing change, or a safety issue with one widely used foundation model has an outsized ripple effect across everything built on it.

The takeaway

A foundation model is any large model pretrained on broad data so it generalizes across many downstream tasks, rather than being built for just one. LLMs are the most visible example, but the category includes vision and multimodal models too. What you do with a foundation model — prompt it directly, fine-tune it, or pair it with retrieval — depends on how much specialization your task needs and how much you’re willing to spend getting it. The tradeoff to keep in mind: because so much gets built on top of a small number of these models, their limitations don’t stay isolated — they inherit downstream, which is exactly why understanding what went into one matters before you build on it.

Chisato Chisato · · 4 min read

Precision vs Recall, Explained

Precision measures how many of a model's positive predictions were correct; recall measures how many actual positives it found. Why you can't max both.

#AI #Machine Learning #LLMs