hackhaton-space-stackHackhaton Space Stack
AcademyBuilder
Get Started
All terms
The Basics
  • Stack
  • Frontend & Backend
  • CLI
  • Monorepo
  • Server
  • Localhost
How AI Works
  • Context Window
  • Hallucination
  • Token
  • Prompt Caching
  • Session
  • Compaction
  • Embedding
  • Vector Database
  • RAG
  • Fine-tuning
  • Temperature
  • Inference
  • Reasoning
  • Multimodal
Building With AI
  • Agent
  • MCP
  • System Prompt
  • Skill
  • CLAUDE.md
  • Slash Command
  • Harness
  • Computer Use
  • Agents SDK
  • Voice Agents
  • OAuth
  • Vibe Coding
  • Permission Scope
  • Tool Calling
  • Prompt Injection
  • Eval
  • Guardrails
  • Sandbox
  • Progressive Disclosure
Code & Collaboration
  • Git
  • Commit
  • Branch
  • GitHub
  • Pull Request
  • Open Source
  • Markdown
  • Dependency
  • Merge
  • Fork
APIs & Connections
  • API
  • Auth
  • Database
  • ORM
  • SDK
  • Webhook
  • Endpoint
  • REST
  • HTTP Methods
  • Env File
  • Schema
  • JSON
  • YAML
  • Secret
  • Rate Limit
  • CORS
  • Cookie
  • Encryption
Shipping & Running
  • Deploy
  • Headless
  • Cron
  • DNS
  • CDN
  • Object Storage
  • Serverless
  • Edge
  • Worker
  • Runtime
  • Process
  • Daemon
  • Queue
  • Job
  • State
  • Cache
  • SSH
  • Build
  • Staging
  • Rollback
  • Docker
  • Feature Flag
  • Test
  • CI/CD
  • The Cloud
Debugging & Errors
  • Trace
  • Type Error
  • Stack Trace
  • Log
  • Bug
  • Patch
  • Latency
How Developers Think
  • DRY
  • YAGNI
  • KISS
  • Refactoring
  • Technical Debt
  • Async
← All terms

Type-safe, modern TypeScript scaffolding for full-stack web development

ThreadsGitHub

Info

  • Academy
  • Docs

Legal

  • Terms of Service
  • Privacy Policy

© 2026 Dzulhelmy Nazri

How AI Works

$defineinference--plain-english

Inference

TLDRThe moment the model actually answers.

Training already happened. You were not in the room.

Someone spent weeks and a warehouse of chips teaching a model to predict the next token. That bill is not on your card. Inference is the live pass: you send a prompt, GPUs walk those tokens, new tokens come back. Hit send in Cursor. Call an API. Let an agent take a step. That is inference. You pay for it. You wait for it. "The model is down" means nobody is running that pass.

This is why the words stream. You are watching tokens land one at a time, not a paragraph appearing from a drawer. A longer answer is a longer pass. A fatter prompt is a fatter pass. That is also why latency feels like "thinking" even when the model is not doing extra reasoning — it is just cooking tonight's plate.

Once you see the live pass, a few product choices stop being vibes:

  • Loops have a price. "Call this a thousand times" is a thousand plates. Demo-night agents that retry forever are a surprise invoice.
  • Prompt caching is leftover prep. The stable prefix gets paid once, then reused.
  • A rate limit is the kitchen capping how many tickets it will take this minute.
  • Reasoning models spend extra inference on a private draft before the sentence you asked for. Same kitchen, longer ticket.

The runtime on your laptop is not the same bill. Local models still do inference — they just burn your fan instead of your API key.

What this unlocks

You budget the thing you actually buy. Not "AI" as a lump. Turns, tokens, retries, tool calls. A hackathon agent that looks smart in one chat can go broke in a loop. Design the loop like it has a meter.

Related

  • Token
  • Reasoning
  • Runtime
PrevTemperature

How AI Works

NextReasoning