hackhaton-space-stackHackhaton Space Stack
AcademyBuilder
Get Started
All terms
The Basics
  • Stack
  • Frontend & Backend
  • CLI
  • Monorepo
  • Server
  • Localhost
How AI Works
  • Context Window
  • Hallucination
  • Token
  • Prompt Caching
  • Session
  • Compaction
  • Embedding
  • Vector Database
  • RAG
  • Fine-tuning
  • Temperature
  • Inference
  • Reasoning
  • Multimodal
Building With AI
  • Agent
  • MCP
  • System Prompt
  • Skill
  • CLAUDE.md
  • Slash Command
  • Harness
  • Computer Use
  • Agents SDK
  • Voice Agents
  • OAuth
  • Vibe Coding
  • Permission Scope
  • Tool Calling
  • Prompt Injection
  • Eval
  • Guardrails
  • Sandbox
  • Progressive Disclosure
Code & Collaboration
  • Git
  • Commit
  • Branch
  • GitHub
  • Pull Request
  • Open Source
  • Markdown
  • Dependency
  • Merge
  • Fork
APIs & Connections
  • API
  • Auth
  • Database
  • ORM
  • SDK
  • Webhook
  • Endpoint
  • REST
  • HTTP Methods
  • Env File
  • Schema
  • JSON
  • YAML
  • Secret
  • Rate Limit
  • CORS
  • Cookie
  • Encryption
Shipping & Running
  • Deploy
  • Headless
  • Cron
  • DNS
  • CDN
  • Object Storage
  • Serverless
  • Edge
  • Worker
  • Runtime
  • Process
  • Daemon
  • Queue
  • Job
  • State
  • Cache
  • SSH
  • Build
  • Staging
  • Rollback
  • Docker
  • Feature Flag
  • Test
  • CI/CD
  • The Cloud
Debugging & Errors
  • Trace
  • Type Error
  • Stack Trace
  • Log
  • Bug
  • Patch
  • Latency
How Developers Think
  • DRY
  • YAGNI
  • KISS
  • Refactoring
  • Technical Debt
  • Async
← All terms

Type-safe, modern TypeScript scaffolding for full-stack web development

ThreadsGitHub

Info

  • Academy
  • Docs

Legal

  • Terms of Service
  • Privacy Policy

© 2026 Dzulhelmy Nazri

How AI Works

$definetoken--plain-english

Token

TLDRThe unit AI reads and writes in.

Ask a model how many R's are in "strawberry" and watch it fumble.

Ask enough times and one will look you dead in the eye and swear there are two. The smartest agent you have ever used, tripped by a word a kid can spell. Once you understand tokens, that stops being a mystery and starts being obvious.

The model never sees your sentence the way you do. Before it reads a single thing you typed, your words get chopped into chunks called tokens. A token is sometimes a whole word, sometimes a piece of one. cat is one token. strawberry becomes a few. Rough rule of thumb:

  • One token is about three-quarters of a word.
  • So 100 tokens is roughly 75 words.
  • A million tokens is around 750,000 words — a small library, not a tweet.

Think of them as the LEGO bricks of language.

You handed the model a sentence. It sees a pile of little bricks, and its entire job is predicting which brick most likely snaps on next. It is not pondering your meaning. It is playing the world's most sophisticated game of "what comes after this."

That one fact is hiding behind three things you have probably already bumped into.

1. The bill. Call a model through an API and you do not pay a flat monthly fee for usage. You pay by the token — the ones you send in and the ones it sends back. A short question is cheap. Pasting the whole repo into one chat is not. Every brick has a price tag.

2. The room. The context window is measured in tokens, not pages or messages. When someone says a model has a "200K context window," that is 200,000 bricks it can hold at once. Fill the room and the early bricks get shoved toward the door.

3. The blind spot. Back to strawberry. The model is not looking at s-t-r-a-w-b-e-r-r-y. It is looking at two or three bricks that, stacked together, mean strawberry to it. The individual letters got swallowed when the word became bricks. Counting R's is a letter question, and there are no letters left to count.

So, practically:

  • Tighten your prompts. Rambling costs tokens and clutters the room. Say what you mean.
  • Do not trust AI on spelling, character counts, or "how many letters." Those are letter questions. The model only ever sees bricks. Verify those yourself.
  • Convert files before pasting. A messy PDF burns tokens decoding layout. Plain text spends those bricks on the actual thinking instead. Same reason a giant JSON blob feels expensive — it is mostly punctuation turned into bricks.

The bricks are not just what you pay for. They are what the model is thinking with. Hand it a cleaner pile and you get a cheaper bill and a sharper answer from the exact same model.

Related

  • Context Window
  • Prompt Caching
  • Inference
  • Hallucination
PrevHallucination

How AI Works

NextPrompt Caching