How AI Works
$definetoken--plain-english
TLDRThe unit AI reads and writes in.
Ask a model how many R's are in "strawberry" and watch it fumble.
Ask enough times and one will look you dead in the eye and swear there are two. The smartest agent you have ever used, tripped by a word a kid can spell. Once you understand tokens, that stops being a mystery and starts being obvious.
The model never sees your sentence the way you do. Before it reads a single thing you typed, your words get chopped into chunks called tokens. A token is sometimes a whole word, sometimes a piece of one. cat is one token. strawberry becomes a few. Rough rule of thumb:
Think of them as the LEGO bricks of language.
You handed the model a sentence. It sees a pile of little bricks, and its entire job is predicting which brick most likely snaps on next. It is not pondering your meaning. It is playing the world's most sophisticated game of "what comes after this."
That one fact is hiding behind three things you have probably already bumped into.
1. The bill. Call a model through an API and you do not pay a flat monthly fee for usage. You pay by the token — the ones you send in and the ones it sends back. A short question is cheap. Pasting the whole repo into one chat is not. Every brick has a price tag.
2. The room. The context window is measured in tokens, not pages or messages. When someone says a model has a "200K context window," that is 200,000 bricks it can hold at once. Fill the room and the early bricks get shoved toward the door.
3. The blind spot. Back to strawberry. The model is not looking at s-t-r-a-w-b-e-r-r-y. It is looking at two or three bricks that, stacked together, mean strawberry to it. The individual letters got swallowed when the word became bricks. Counting R's is a letter question, and there are no letters left to count.
So, practically:
The bricks are not just what you pay for. They are what the model is thinking with. Hand it a cleaner pile and you get a cheaper bill and a sharper answer from the exact same model.