You would never ship a payments function without tests. But the prompt that decides which customers get a refund? That ships on vibes — edited in a chat window, eyeballed once, and pasted into production. Then the model changes, or someone "improves" the wording, and the refund bot starts apologizing in French.

promptfoo (promptfoo.dev) is the open-source tool that fixes this: a CLI and library for test-driven LLM development. You define prompts, providers, and test cases in a YAML config, attach assertions to the outputs, and run promptfoo eval the way you'd run pytest. It compares prompt versions side by side, scores every output, and exports machine-readable results you can gate a deploy on.

This tutorial does the whole loop for real. We'll build a support-ticket triage bot, write two competing prompts, test them against three tickets with eight assertion types, watch the naive prompt fail exactly the way it would in production, scale the test suite through a CSV file, and wire the whole thing into CI. Every command below was executed on September 30, 2026 against promptfoo 0.123.1, and every output shown is real. Best of all: no API key, no GPU, no spend — the entire tutorial runs against a deterministic mock provider, so it's free and perfectly reproducible.

What you'll need #

  • Node.js 22 or newer — check with node --version. promptfoo ships as an npm package; we'll run it through npx so nothing is installed globally.
  • Python 3 — only for the mock provider in Steps 2–7 (a ~60-line script, no dependencies).
  • A terminal and about 10 minutes of patience once — the first npx invocation downloads the package through npm; after that it's cached.
  • No API key, no account, no GPU. Model-graded assertions (like llm-rubric) exist and are documented, but this tutorial sticks to deterministic assertions so everything runs offline.

Step 1 — Install promptfoo #

No install step, really — npx fetches and runs the latest release on demand:

npx -y promptfoo@latest --version
# 0.123.1

That's the version every command in this tutorial was verified against. If you'd rather have it on your PATH permanently, npm install -g promptfoo works too — the commands are identical either way. A quick promptfoo eval --help is worth reading in full; the flag list we'll actually use (--no-cache, --filter-prompts, -o, --table) is small, but the CLI also offers first-N filters, metadata filters, sampling, output transforms, and JUnit/XML/CSV/HTML export.

Step 2 — A provider that costs $0: the mock Python provider #

promptfoo's providers list accepts real model IDs (openai:gpt-6-luna, anthropic:messages:claude-opus-4-6, google:gemini-3.8-flash, ollama:llama2 — the names straight from the docs), but it also accepts a Python file exporting a call_api(prompt, options, context) function. The docs list "creating mock providers for testing" as a first-class use case, and it's the perfect way to learn: deterministic output, zero cost, zero latency, zero flakiness.

Create a project directory and save this as mock_support_bot.py. It reads the rendered prompt, detects which prompt variant is being tested, and returns a canned response keyed off ticket keywords:

"""Deterministic mock support-bot: no network, no API key."""

def _variant(prompt: str) -> str:
    # The v2 prompt explicitly asks for JSON output; v1 does not.
    if "Return your answer as JSON" in prompt:
        return "v2"
    return "v1"


def _subject(ticket: str):
    tl = ticket.lower()
    if "password" in tl or "reset" in tl:
        return "password reset email issue", "account"
    if "charged" in tl or "refund" in tl or "invoice" in tl:
        return "duplicate billing charge", "billing"
    return "router connectivity issue", "network"


def call_api(prompt, options, context):
    variant = _variant(prompt)
    ticket = ((context or {}).get("vars", {}) or {}).get("ticket", "")
    subject, category = _subject(ticket)

    if variant == "v1":
        output = (
            "I'm really sorry about the trouble you're having! "
            "I've looked at your message and I think the best thing is to try "
            "restarting the router, and if that doesn't work please let me know "
            "and I will escalate this to our support team. Sorry again!"
        )
    else:
        output = (
            '{"summary": "%s", "category": "%s", '
            '"next_step": "restart router, then check link lights", '
            '"escalate": false}' % (subject, category)
        )
    return {
        "output": output,
        "tokenUsage": {"prompt": 120, "completion": 60, "total": 180},
    }

Two honest caveats before we go further. First, a mock is a stand-in: it proves your test harness works, not that a real model behaves. When you graduate to a real provider, pin temperature: 0 and use --repeat to check stability — LLMs are nondeterministic and your assertions need to survive that. Second, the mock returns token usage so promptfoo's cost/latency tracking still exercises; real providers report the real numbers.

Step 3 — Two prompts, one shootout #

promptfoo treats each prompt file as a column in a results matrix: every prompt runs against every test case, so comparing prompt versions is the default workflow, not a special mode. Save these two:

prompt-v1.txt — the naive prompt, the kind that ships on vibes:

You are a helpful support agent. Answer the customer's ticket below.

Ticket: {{ticket}}

prompt-v2.txt — the engineered prompt, with a strict output contract:

You are a terse support-triage bot. Never apologize. Return your answer as JSON with exactly these keys: summary, category, next_step, escalate. No prose outside the JSON object.

Ticket: {{ticket}}

{{ticket}} is the variable each test case will fill in. Notice that v2's entire strategy is a machine-checkable contract — JSON shape, required keys, no apologies. That's the bet the test suite is about to settle.

Step 4 — The test suite: prompts × providers × assertions #

The heart of promptfoo is promptfooconfig.yaml. It declares the full matrix — prompts, providers, test cases — and, crucially, the assertions that decide pass or fail for every cell:

description: "Support-bot prompt shootout: naive prompt vs JSON-structured prompt"

prompts:
  - file://prompt-v1.txt
  - file://prompt-v2.txt

providers:
  - id: file://mock_support_bot.py
    label: mock-support-bot

defaultTest:
  assert:
    - type: javascript
      value: output.length < 500

tests:
  - description: "Router connectivity ticket"
    vars:
      ticket: "My internet drops every evening around 8pm. Router model AX5400."
    assert:
      - type: icontains
        value: router
      - type: not-contains
        value: sorry
      - type: regex
        value: '"category":\s*"[a-z]+"'

  - description: "Billing refund ticket"
    vars:
      ticket: "I was charged twice for my March invoice. Please refund the duplicate."
    assert:
      - type: is-json
      - type: python
        value: |
          import json
          try:
              obj = json.loads(output)
          except Exception:
              return False
          return "escalate" in obj and "next_step" in obj

  - description: "Password reset ticket"
    vars:
      ticket: "I forgot my password and the reset email never arrives."
    assert:
      - type: icontains-any
        value: [password, reset]
      - type: not-icontains
        value: sorry

Three mechanics worth naming. defaultTest applies its assertions to every test case — here, a global "keep responses under 500 characters" rule. A test can opt out with options: {disableDefaultAsserts: true}. And weights: every assertion accepts a weight, scores combine into a weighted average, and a per-test threshold decides pass/fail. We verified this directly: a test with a passing icontains (weight 3) and a failing is-json (weight 1) scored exactly 0.75 — and with threshold: 0.75 it passed, since the comparison is score ≥ threshold.

Here's the assertion tour — every one of these was executed in this tutorial, not copied from docs:

  • javascript — run JS against the output: output.length < 500.
  • python — run Python with output in scope; return truthy to pass. Lesson learned the hard way: my first version called json.loads(output) bare, and when v1 returned prose instead of JSON the eval crashed with a traceback. Wrap parsing in try/except and return False — a malformed output is a failed assertion, not a crashed harness.
  • icontains / not-contains / not-icontains — case-insensitive (or sensitive) substring checks and their negations. The not- prefix negates any deterministic assertion type.
  • regex — pattern matching; here it enforces the "category": "..." JSON shape without parsing.
  • is-json — rejects any output that isn't parseable JSON. This single assertion is what kills the naive prompt.
  • icontains-any — passes if any of the listed values appears.

Beyond these, the docs offer model-graded assertions (llm-rubric, factuality, similar with a cosine-similarity threshold, context-faithfulness for RAG) that use a judge model to score outputs, plus cost and latency thresholds and assertion sets. I didn't run those here — they need an API key and a judge model, which breaks the zero-cost premise — but the deterministic set above covers the majority of prompt-regression work.

AI-generated editorial illustration: two glowing prompt documents facing each other across a pass-fail comparison matrix of green checkmarks and red crosses on a dark navy background

Step 5 — Run the eval #

promptfoo eval --no-cache

(--no-cache forces a fresh run instead of reusing cached results — essential when you're iterating on prompts.) Real output, abridged:

Starting evaluation eval-Hbz-2026-09-30T23:57:30
Running 6 test cases (up to 4 at a time)...

┌────────────────────────────────────────┬────────────────────────┬────────────────────────┐
│ ticket                                 │ [mock-support-bot]     │ [mock-support-bot]     │
│                                        │ prompt-v1.txt          │ prompt-v2.txt          │
├────────────────────────────────────────┼────────────────────────┼────────────────────────┤
│ My internet drops every evening around │ [FAIL] I'm really      │ [PASS] {"summary":     │
│ 8pm. Router model AX5400.              │ sorry about the        │ "router connectivity    │
│                                        │ trouble you're         │ issue", "category":    │
│                                        │ having! ...            │ "network", ...}        │
├────────────────────────────────────────┼────────────────────────┼────────────────────────┤
│ ...                                    │ ...                    │ ...                    │
└────────────────────────────────────────┴────────────────────────┴────────────────────────┘

Results:
  ✓ 3 passed (50.00%)
  ✗ 3 failed (50.00%)
  0 errors (0%)

Six cells: 2 prompts × 3 tickets. The default table prints the full matrix in your terminal; --table-cell-max-length trims long cells, and promptfoo view opens the same results in a web UI for clicking through outputs (it needs a browser, so I stayed in the CLI here).

Step 6 — Read the regression the suite just caught #

Here's the scoreboard that matters:

  • prompt v1 (naive): 0/3. It fails every ticket — the is-json assertion rejects its prose, the regex finds no "category" key, and not-contains: sorry catches the double apology.
  • prompt v2 (JSON contract): 3/3. Every ticket passes every assertion, including the global output.length < 500 rule.

This is the entire value proposition in one screen. The failure isn't a vibe or a feeling — it's which assertion failed on which test case, reproducible on every run. If someone "improves" v2 next month and the JSON contract breaks, this suite catches it before the deploy does. That's what "test your prompts like code" actually means: the prompt is a build artifact, and the eval is its test suite.

Step 7 — Scale the suite with CSV test cases #

Three hand-written YAML tests are a start; real suites have hundreds, and those usually live in a spreadsheet. promptfoo loads test cases from CSV — each column is a variable, and special __expectedN columns carry assertions in type: value form (a bare value defaults to equals):

ticket,__expected1,__expected2
"My internet drops every evening around 8pm.","icontains: router","is-json"
"I was charged twice for my March invoice.","icontains: billing","not-icontains: sorry"
# csv-config.yaml
prompts:
  - file://prompt-v2.txt
providers:
  - file://mock_support_bot.py
tests: file://tests.csv
promptfoo eval -c csv-config.yaml --no-cache --no-table

Results:
  ✓ 2 passed (100%)
  0 failed (0%)

Verified working. The docs also document __description, __threshold, and __metadata:* columns (the last enables --filter-metadata), plus Google Sheets as a live test source — handy when non-engineers own the test cases. One quoting gotcha from the docs: values containing commas need doubled quotes inside the quoted cell ("contains-all: ""1,000"",in stock").

Step 8 — Iterate fast: run only what changed #

Once the suite grows, you don't want the full matrix on every tweak. Verified in this tutorial:

# Only the v2 prompt — note: this is a REGEX, not a glob.
# --filter-prompts='*v2*' crashes with "Invalid regex pattern".
promptfoo eval --no-cache --filter-prompts='v2' --no-table

Results:
  ✓ 3 passed (100%)

The full --help (read it — it's the source of truth for your installed version) offers more: --filter-pattern and --filter-range for tests, --filter-failing <results.json> to re-run only what failed last time, --filter-errors-only, --filter-sample with a seed for quick smoke runs, --var k=v to inject variables, -j for concurrency, --repeat and --delay for stability and rate-limit friendliness, and promptfoo validate to check your config before you run it (ours reports Configuration is valid.).

Step 9 — Gate your deploys: prompt evals in CI #

Here's the most important finding in this tutorial, and it's a gotcha: promptfoo eval exits 0 even when assertions fail. Our 3-failed run above returned exit code 0. Unlike pytest, a red suite won't fail your build by itself — you have to gate explicitly. That's what machine-readable output is for:

promptfoo eval --no-cache -o results.json -o junit.xml

results.json contains every test run with success, score, failureReason, and per-assertion componentResults; junit.xml feeds straight into CI test dashboards. A tiny gate script turns the JSON into an exit code:

"""Fail the build when any promptfoo test fails."""
import json, sys

results = json.load(open(sys.argv[1]))["results"]["results"]
failed = [r for r in results if not r["success"]]
for r in failed:
    label = r["prompt"]["label"][:40]
    desc = (r["testCase"].get("description") or "?")[:50]
    print(f"FAIL  [{label}]  {desc}")
print(f"\n{len(results) - len(failed)}/{len(results)} passed")
sys.exit(1 if failed else 0)

Verified against our real runs: exit 1 with the three v1 failures listed, exit 0 on the all-green CSV run. The GitHub Actions workflow is then straightforward:

name: prompt-evals
on:
  pull_request:
    paths: ["prompts/**", "promptfooconfig.yaml"]
jobs:
  eval:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: 22 }
      - run: npx -y promptfoo@latest eval --no-cache -o results.json -o junit.xml
      - run: python gate.py results.json
      - uses: actions/upload-artifact@v4
        if: always()
        with: { name: promptfoo-results, path: "results.json,junit.xml" }

Now a prompt change that breaks the JSON contract fails the PR exactly like a broken unit test. For the model-graded assertions in production suites, note the cost implication: every CI run spends judge-model tokens, so many teams run deterministic assertions on every PR and reserve llm-rubric evals for nightly runs or release branches.

AI-generated editorial illustration: a continuous-integration pipeline as a glowing conveyor belt carrying prompt cards through green pass gates and a red gate stopping a defective card, on a dark navy background

Should you adopt it? #

  • Adopt it if you iterate on prompts at all, compare models or prompt versions, or have ever been bitten by a "small wording tweak" that silently changed behavior. The mock-provider workflow means the harness costs nothing to try.
  • Start with deterministic assertions. is-json, icontains, regex, and javascript/python checks are fast, free, and stable. Add model-graded assertions only where human judgment is genuinely the spec.
  • Watch the seams. The exit-code behavior means you must wire the gate yourself — don't assume red means failed in CI. Mock providers validate the harness, not the model; re-run the suite against the real provider with temperature: 0 and --repeat before trusting green. And nondeterministic outputs need statistical thinking: a single pass isn't proof.
  • Skip it if your prompts are truly one-off and never change. But the moment a prompt becomes infrastructure — a support bot, a classifier, a RAG pipeline — it's code, and code without tests is a liability.

The takeaway #

The uncomfortable truth this tutorial demonstrates is how little it takes. One YAML file, two prompt files, a 60-line mock, and eight assertions — and suddenly "does the new prompt still work?" is an answered question instead of a hope. The v1 prompt didn't fail because it was dumb; it failed because nobody had written down what "working" meant. promptfoo is just the discipline of writing that down, in a form a machine can check on every commit.

Clone the pattern tonight: pick one prompt in your stack that matters, write three test cases with assertions for its output contract, and run promptfoo eval. If it goes green, you have a regression suite. If it goes red, you just found the bug your users were going to find for you.