What are AI evals?
Quick answer
AI evals, short for evaluations, are tests that run an AI system on a fixed set of example inputs and compare the outputs to what a good answer should be. They show whether a prompt, model or workflow change made results better or worse, instead of guessing from one try.
Last updated
Updated · By Robert Breen
Why it matters for a small business
Trying a prompt once and liking the answer tells you very little. The next input may be messier, the model may be updated, or a small wording change to the system message may fix one case and break another. An eval set catches that, because you run the same examples every time and compare.
An eval for office work does not need special software. It can be a sheet with twenty real examples, the answer you expect for each, and a column for pass or fail. Some checks are exact, like the right category or date. Others need a person, or a second AI pass, to judge whether the tone and facts are right. Teams that keep an eval set can switch models or edit prompts with far more confidence.
In a real lesson: Build an AI Agent That Categorizes Business Expenses
The AI Expense Categorizer Agent lesson ends with a tiny eval hiding in plain sight. You build an agent for Maple Street Bookkeeping, a made-up bookkeeping firm, that sorts expenses into Office Supplies, Software & Subscriptions, Meals, Travel, Utilities or Needs Review.
The test sends three expenses with obvious right answers: Corner Office Supply, $64.18, printer paper and toner; CloudLedger, $45.00, monthly accounting software; Harbor Street Cafe, $38.50, lunch with a client. The agent returns Office Supplies, Software & Subscriptions and Meals, and the amounts match exactly. That is a pass on three cases.
To turn that into a real eval, you would keep those three and add harder ones from your own books, written down with the expected category: a charge that could be Travel or Meals, a vendor with no description, an amount that looks wrong. The system message says unclear items should go to Needs Review and that the agent must "not change amounts or invent details," so an expected answer of Needs Review, with the amount untouched, is a fair test. Rerun the set after every prompt edit.

Try this lesson free or read the step-by-step guide.
Common confusions
Evals vs benchmarks
Public benchmarks compare models on general tasks. Your own evals test your task, with your data and your idea of a good answer. A model that tops a benchmark can still fail your expense rules.
Evals vs self-checking
Asking the model to check its own draft, as in self-verification, improves one answer. An eval measures many answers over time, so you can see whether the whole setup is getting better.
Tips
- Start with 10 to 20 real examples, including the awkward ones that went wrong before.
- Write the expected answer before you run the test, not after.
- Keep test data free of real client details, or remove them first.
- Rerun the same set whenever you change the prompt, the model or the tools.
Related terms
More AI basics terms
Where you use it: free lessons
- Build an AI Agent That Categorizes Business Expenses (n8n, 12 min)
- AI Receipt Extractor: Receipts to Google Sheets with n8n (n8n, 12 min)
- Reply Faster to Claim-Status Emails: ChatGPT for Independent Insurance Agencies (ChatGPT, 10 min)
Prompt templates that use it
Frequently asked questions
- How many examples does an eval need?
- A small team can start with 10 to 20 well-chosen cases. Add every real mistake you find as a new case, so the set grows to cover what actually goes wrong.
- Can AI grade its own evals?
- A second AI pass can grade things like tone or whether each claim matches a source, but check a sample of its grades by hand. Exact checks, like amounts and categories, are better done with simple comparisons.