Memia

Memia

Memia Strategy Notes

🤖Anatomy of an AI agent (mid-2026 edition)

Models, tokens, harnesses, agents, skills, tools — making sense of a dynamic vocabulary

Ben Reid's avatar
Ben Reid
Jun 09, 2026
∙ Paid

Models, tokens, context windows, harnesses, agents, skills, tools… the vocabulary around AI keeps changing rapidly in 2026 — and as a result discussions about “how does it work?” often… lack precision.

In this short note I’ve put down my own evolving mental model (honed with the assistance of AI, naturally) of how the pieces actually fit together. Hopefully useful!

(Version 1.0 - let me know what I’ve missed / if you agree / disagree on any points in the comments below…)

Key takeaway

With the rapid advancement of AI agents, the system *around* the model now matters as much as the model itself. Once a frontier model crosses a capability threshold, much of the remaining gains in real, long-horizon work come not from a bigger brain but from better scaffolding around it.1

The stack

Here is my current understanding of the architecture / entity relationships behind an AI agent. The software runtime which orchestrates an Agentic AI system is the Harness and everything else operates within that information boundary. The User (either a human or a further-up-the-chain agent) instructs the Agent with a prompt and the harness carries all of the information and constraints on what to do from there. As the agent progresses on its work, it can feed back information to the User until(/…if) it completes its work.

The key logical entities:

  1. Tokens

  2. Context window

  3. Model

  4. Agent

  5. Tools & MCP

  6. Skills

  7. Memory

  8. Harness

  9. Orchestration

We can think of these concepts as nested layers: the harness contains and runs everything; the agent is the reasoning loop; the model is the brain it calls; and tokens are the substrate each layer is measured in.

In one sentence:

Tokens are the substrate, the context window is the budget, the model is the brain, the agent is the loop that gives it agency, tools (via MCP) are its hands, skills and memory are its loadable knowledge — and the harness is the runtime that conducts the whole orchestra.

Drilling down into each in turn:

1. Tokens

A token is the smallest chunk of language a model processes — a word, a fragment of a word, or a single character, depending on the tokeniser. (As a rough rule, one token is about three-quarters of an English word, or four characters)2.

User's avatar

Continue reading this post for free, courtesy of Ben Reid.

Or purchase a paid subscription.
© 2026 Ben Reid (CC BY 4.0) · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture