▶ Stepthrough Courses All tutorials Blog Glossary Prompts Videos Visual guides Cheat sheets Comparisons Start Learning Free

What are AI evals?

Quick answer

AI evals, short for evaluations, are tests that run an AI system on a fixed set of example inputs and compare the outputs to what a good answer should be. They show whether a prompt, model or workflow change made results better or worse, instead of guessing from one try.

Last updated

Updated · By Robert Breen

Why it matters for a small business

Trying a prompt once and liking the answer tells you very little. The next input may be messier, the model may be updated, or a small wording change to the system message may fix one case and break another. An eval set catches that, because you run the same examples every time and compare.

An eval for office work does not need special software. It can be a sheet with twenty real examples, the answer you expect for each, and a column for pass or fail. Some checks are exact, like the right category or date. Others need a person, or a second AI pass, to judge whether the tone and facts are right. Teams that keep an eval set can switch models or edit prompts with far more confidence.

In a real lesson: Build an AI Agent That Categorizes Business Expenses

The AI Expense Categorizer Agent lesson ends with a tiny eval hiding in plain sight. You build an agent for Maple Street Bookkeeping, a made-up bookkeeping firm, that sorts expenses into Office Supplies, Software & Subscriptions, Meals, Travel, Utilities or Needs Review.

The test sends three expenses with obvious right answers: Corner Office Supply, $64.18, printer paper and toner; CloudLedger, $45.00, monthly accounting software; Harbor Street Cafe, $38.50, lunch with a client. The agent returns Office Supplies, Software & Subscriptions and Meals, and the amounts match exactly. That is a pass on three cases.

To turn that into a real eval, you would keep those three and add harder ones from your own books, written down with the expected category: a charge that could be Travel or Meals, a vendor with no description, an amount that looks wrong. The system message says unclear items should go to Needs Review and that the agent must "not change amounts or invent details," so an expected answer of Needs Review, with the amount untouched, is a fair test. Rerun the set after every prompt edit.

n8n AI Agent node with a system message written for Maple Street Bookkeeping
n8n AI Agent node with a system message written for Maple Street Bookkeeping

Try this lesson free or read the step-by-step guide.

Common confusions

Evals vs benchmarks

Public benchmarks compare models on general tasks. Your own evals test your task, with your data and your idea of a good answer. A model that tops a benchmark can still fail your expense rules.

Evals vs self-checking

Asking the model to check its own draft, as in self-verification, improves one answer. An eval measures many answers over time, so you can see whether the whole setup is getting better.

Tips

  • Start with 10 to 20 real examples, including the awkward ones that went wrong before.
  • Write the expected answer before you run the test, not after.
  • Keep test data free of real client details, or remove them first.
  • Rerun the same set whenever you change the prompt, the model or the tools.

More AI basics terms

Where you use it: free lessons

Prompt templates that use it

Frequently asked questions

How many examples does an eval need?
A small team can start with 10 to 20 well-chosen cases. Add every real mistake you find as a new case, so the set grows to cover what actually goes wrong.
Can AI grade its own evals?
A second AI pass can grade things like tone or whether each claim matches a source, but check a sample of its grades by hand. Exact checks, like amounts and categories, are better done with simple comparisons.

All AI glossary terms, A to Z · Free prompt templates