Regression-Test Your Prompts with promptfoo in CI
Catch prompt and model regressions on every pull request with promptfoo evals in GitHub Actions.
What you'll build
A regression-test suite for an LLM prompt using promptfoo, wired into GitHub Actions so every pull request that touches your prompts runs the evals, posts a pass/fail comment on the PR, and fails the check if a change breaks expected behavior.
Prerequisites
- Node.js >= 22.22.0. promptfoo enforces this hard, so use Node.js 24 LTS. Verified with Node 24.1.0.
- promptfoo 0.122.2 (the current release on npm; all commands below were run against it). No install needed if you use
npx. - An Anthropic API key from the Anthropic Console. Any provider promptfoo supports works; this tutorial uses Claude.
- A GitHub repository with Actions enabled, and the GitHub CLI (
gh) for adding the secret. - Commands assume a POSIX shell (macOS/Linux). Windows users: run them in Git Bash or WSL.
1. Set up the project
From your repo root, create a directory for prompt files:
mkdir -p prompts
export ANTHROPIC_API_KEY=sk-ant-... # your key, for local runs
You don't need to install anything. npx promptfoo@latest fetches the CLI on demand, and the GitHub Action installs its own copy in CI.
2. Write the prompt under test
Create prompts/triage.txt. This is a support-ticket triage prompt with a strict output contract, which is exactly the kind of thing that silently breaks when someone "improves" the wording:
You are a support ticket triage assistant. Read the ticket below and respond with only a JSON object — no markdown, no code fences — with exactly these keys:
- "category": one of "billing", "account", "bug", "other"
- "urgency": one of "low", "medium", "high"
- "summary": one neutral, professional sentence describing the issue
Ticket:
{{ticket}}
{{ticket}} is a Nunjucks variable. Each test case fills it in.
3. Configure the eval
Create promptfooconfig.yaml in the repo root:
# yaml-language-server: $schema=https://promptfoo.dev/config-schema.json
description: Support ticket triage regression tests
prompts:
- file://prompts/triage.txt
providers:
- id: anthropic:messages:claude-opus-5
config:
max_tokens: 512
defaultTest:
options:
provider: anthropic:messages:claude-opus-5
tests:
- description: duplicate charge routes to billing
vars:
ticket: I was charged twice for my subscription this month and need the duplicate refunded.
assert:
- type: is-json
- type: javascript
value: JSON.parse(output).category === 'billing'
- description: password reset is low urgency
vars:
ticket: How do I reset my password? Not urgent, just locked out of the mobile app.
assert:
- type: is-json
- type: javascript
value: JSON.parse(output).urgency === 'low'
- description: angry ticket still gets a neutral summary
vars:
ticket: This is the THIRD time your app deleted my data. Fix it or I'm cancelling and telling everyone.
assert:
- type: llm-rubric
value: The summary is neutral and professional and does not adopt the customer's angry tone.
Three details matter here. The javascript assertions are inline expressions that get the raw model output as output, so you can parse and check exact fields deterministically and for free. The llm-rubric assertion uses a model as grader for the fuzzy tone requirement. And the defaultTest.options.provider block pins that grader to Claude, because promptfoo grades with OpenAI by default and you'd otherwise need a second API key in CI. The Anthropic provider runs at temperature 0 by default, which keeps reruns stable.
4. Run it locally
npx promptfoo@latest eval
eval auto-loads promptfooconfig.yaml from the current directory. Open the results table in a browser with:
npx promptfoo@latest view
The exit code is the CI contract: 0 when everything passes, 100 when at least one assertion fails, 1 on any other error. You'll rely on that in the next step.
5. Wire it into GitHub Actions
Add your key as a repo secret:
gh secret set ANTHROPIC_API_KEY
Create .github/workflows/prompt-eval.yml:
name: Prompt regression tests
on:
pull_request:
paths:
- 'prompts/**'
- 'promptfooconfig.yaml'
jobs:
evaluate:
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: write
steps:
- uses: actions/checkout@v7
- uses: actions/setup-node@v7
with:
node-version: 24
- name: Cache promptfoo
uses: actions/cache@v6
with:
path: ~/.cache/promptfoo
key: ${{ runner.os }}-promptfoo-${{ hashFiles('promptfooconfig.yaml') }}
restore-keys: |
${{ runner.os }}-promptfoo-
- uses: promptfoo/promptfoo-action@v1
with:
anthropic-api-key: ${{ secrets.ANTHROPIC_API_KEY }}
github-token: ${{ secrets.GITHUB_TOKEN }}
config: promptfooconfig.yaml
use-config-prompts: true
cache-path: ~/.cache/promptfoo
no-share: true
Why each piece: the paths filter skips the job on PRs that don't touch prompts or tests. pull-requests: write lets the action post its results comment. The cache stores LLM responses keyed on request content, so unchanged prompt/test pairs cost nothing on reruns. use-config-prompts: true makes the action evaluate the prompt list from your config instead of trying to infer it from changed files. no-share: true keeps eval results out of promptfoo's public sharing service.
Verify it works
Locally, a passing run ends like this (IDs and timing will differ):
Running 3 test cases (up to 4 at a time)...
✓ Eval complete (ID: eval-1Pr-2026-09-07T11:38:23)
Results:
✓ 3 passed (100%)
0 failed (0%)
0 errors (0%)
Duration: 9s (concurrency: 4)
Confirm the exit code with echo $? (expect 0).
Then prove the CI gate actually gates. Open a branch, edit prompts/triage.txt to remove the line about responding "with only a JSON object", and open a PR. The workflow runs, the is-json assertions fail on fenced output, the action comments a pass/fail table on the PR, and the check goes red with exit code 100. Revert the edit and the check goes green.
Troubleshooting
Required: >=22.22.0 followed by Install a supported Node.js version and try again. Your Node is too old for current promptfoo. Install Node 24 (nvm install 24 or brew install node) and re-run. In CI this can't happen because the workflow pins node-version: 24.
✗ Missing ANTHROPIC_API_KEY (anthropic:messages:claude-opus-5) means the provider can't find a key. Locally, export ANTHROPIC_API_KEY=... in the same shell. In CI, check that the secret exists (gh secret list) and that the workflow passes it via the anthropic-api-key input. Secret names are case-sensitive.
The CI job fails with exit code 100 but no stack trace. That's not a crash. 100 is promptfoo's "at least one test failed" code. Read the results table in the job log or the PR comment to see which assertion failed. If you need a different code, set the PROMPTFOO_FAILED_TEST_EXIT_CODE environment variable.
Resource not accessible by integration when posting the PR comment means the workflow token lacks write access. Make sure the job has permissions: pull-requests: write, as in the workflow above. PRs from forks get a read-only token by design; for public repos, expect the comment step to fail on fork PRs even though the eval itself runs.
Next steps
- Add a second entry under
providers:to compare models side by side in one results table. This is how you test a model upgrade before committing to it. - Set
fail-on-threshold: 90on the action to allow a pass rate above 90% instead of requiring perfection, useful once your suite grows past a handful of cases. - Emit machine-readable results with
-o results.junit.xml(also.json,.html) and feed them to your CI's test reporting. - Read promptfoo's CI/CD integration guide for GitLab/Jenkins equivalents, and the assertions reference for the full assertion catalog, including semantic similarity and JSON schema validation.
Sources & further reading
- Getting started — promptfoo.dev
- GitHub Actions integration — promptfoo.dev
- Anthropic provider — promptfoo.dev
- Command line reference — promptfoo.dev
- Assertions and metrics — promptfoo.dev
- promptfoo-action repository — github.com
Rachel has been embedded in the developer tooling ecosystem for nearly eight years, covering everything from IDE wars and package-manager drama to the quiet rise of AI-assisted coding. She has a soft spot for open-source maintainers and an unhealthy number of terminal emulators installed on a single laptop.
Discussion 1
oh interesting, so you can catch when your prompts drift without manually testing every pr. gonna try wiring this up this weekend