hackhaton-space-stackHackhaton Space Stack
AcademyBuilder
Get Started
All terms
The Basics
  • Stack
  • Frontend & Backend
  • CLI
  • Monorepo
  • Server
  • Localhost
How AI Works
  • Context Window
  • Hallucination
  • Token
  • Prompt Caching
  • Session
  • Compaction
  • Embedding
  • Vector Database
  • RAG
  • Fine-tuning
  • Temperature
  • Inference
  • Reasoning
  • Multimodal
Building With AI
  • Agent
  • MCP
  • System Prompt
  • Skill
  • CLAUDE.md
  • Slash Command
  • Harness
  • Computer Use
  • Agents SDK
  • Voice Agents
  • OAuth
  • Vibe Coding
  • Permission Scope
  • Tool Calling
  • Prompt Injection
  • Eval
  • Guardrails
  • Sandbox
  • Progressive Disclosure
Code & Collaboration
  • Git
  • Commit
  • Branch
  • GitHub
  • Pull Request
  • Open Source
  • Markdown
  • Dependency
  • Merge
  • Fork
APIs & Connections
  • API
  • Auth
  • Database
  • ORM
  • SDK
  • Webhook
  • Endpoint
  • REST
  • HTTP Methods
  • Env File
  • Schema
  • JSON
  • YAML
  • Secret
  • Rate Limit
  • CORS
  • Cookie
  • Encryption
Shipping & Running
  • Deploy
  • Headless
  • Cron
  • DNS
  • CDN
  • Object Storage
  • Serverless
  • Edge
  • Worker
  • Runtime
  • Process
  • Daemon
  • Queue
  • Job
  • State
  • Cache
  • SSH
  • Build
  • Staging
  • Rollback
  • Docker
  • Feature Flag
  • Test
  • CI/CD
  • The Cloud
Debugging & Errors
  • Trace
  • Type Error
  • Stack Trace
  • Log
  • Bug
  • Patch
  • Latency
How Developers Think
  • DRY
  • YAGNI
  • KISS
  • Refactoring
  • Technical Debt
  • Async
← All terms

Type-safe, modern TypeScript scaffolding for full-stack web development

ThreadsGitHub

Info

  • Academy
  • Docs

Legal

  • Terms of Service
  • Privacy Policy

© 2026 Dzulhelmy Nazri

Building With AI

$defineeval--plain-english

Eval

TLDRA test that tells you if the change helped or quietly hurt.

An eval is a practice exam you wrote yourself — then graded after every tweak.

You keep a set of questions with answers you like, or a rubric, or a second model that scores. You change the system prompt. You swap a skill. You run the set again. The score goes up or it does not. Feelings are not a score.

A benchmark is a famous shared exam so two models can be compared on the same yardstick. Useful as a rough signal. Easy to cram for. Trust your exam on your job more.

This is how you catch the quiet regression

You fix one annoying case, ship it, and never notice you broke five others. An eval is the difference between "seems better" and "scored better on the same fifty, every time."

For this product, the ugly cases are obvious: "Next + Hono + Postgres + Better Auth" must plan, then create, not hallucinate a combo the builder would grey out. "Add PWA" must hit ds_plan_addons, not invent a second monorepo tool. Compatibility rules in the engine are a kind of eval the CLI already runs. Prompt-level evals are how you know the plugin still steers toward MCP after you "improved" the wording.

Start with ten cases that already failed once. Tests for code. Evals for fuzzy employees. Without them, every prompt tweak is a superstition.

What this unlocks

You can tell a lucky demo from a real improvement. Vibe coding without a score is steering on vibes twice.

A demo is a story. An eval is a receipt.

Related

  • Hallucination
  • Test
  • Agent
  • Fine-tuning
PrevPrompt Injection

Building With AI

NextGuardrails