Prompt Regression Testing¶
"pytest for multi-agent prompts" — catch prompt regressions before they reach production.
The problem¶
When you change a system prompt in a running pipeline, there's no safety net. A prompt tweak that looks harmless can cascade through multiple agents, silently degrading output quality in ways that only surface days later in production.
Traditional unit tests don't help because the outputs are non-deterministic. What you need is mutation testing against recorded runs.
How it works¶
AntCrew's replay_with_mutation() in TraceLog re-executes a historically recorded run with a modified prompt for one agent and quantifies the diff:
- Record — run your pipeline with
--full-traceto capture every prompt + response - Mutate — substitute the system prompt for one agent
- Replay — re-run every agent call in order, with the new prompt for the target agent
- Diff — compare the new output against the original using unified diff; compute
diff_pct - Gate — exit code 1 if
diff_pct > threshold
CLI usage¶
Prerequisites¶
Record a run with full trace enabled:
Get the run ID:
File-based prompt test¶
Inline prompt test¶
antcrew regtest ~/.antcrew/trace.db \
--run <run-id> \
--agent BA \
--prompt-text "You are a Business Analyst. Be extremely concise."
Output¶
antcrew regtest run=4c3fa8b2… agent=BA threshold=20%
Agent Mutated Diff % Match Cost
─────────────────────────────────────────────
BA ✎ 34.1% ✗ $0.0041
PM — — ✓ $0.0028
BackendDev — — ✓ $0.0056
✗ FAIL diff=34.1% threshold=20% changed=1/3 agents
Exit code 0 on pass, 1 on fail.
JSON output for scripting¶
{
"run_id": "4c3fa8b2...",
"agent_name": "BA",
"threshold": 0.2,
"diff_pct": 0.341,
"passed": false,
"total_changed": 1,
"calls": [...]
}
CI/CD integration¶
GitHub Actions — full reusable workflow¶
Save as .github/workflows/prompt-regression.yml in your repository:
name: Prompt regression tests
on:
pull_request:
paths:
- 'prompts/**'
- '**.yaml'
- '**.yml'
jobs:
regression:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.11'
- name: Install antcrew
run: pip install antcrew
- name: Download baseline trace
env:
ANTCREW_API_KEY: ${{ secrets.ANTCREW_API_KEY }}
run: |
# Download the TraceLog from your platform instance
curl -fsSL \
-H "X-Api-Key: $ANTCREW_API_KEY" \
"https://app.antcrew.io/runs/${{ vars.BASELINE_RUN_ID }}/tracelog" \
-o baseline.db
- name: Run prompt regression — BA agent
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
antcrew regtest baseline.db \
--run "${{ vars.BASELINE_RUN_ID }}" \
--agent BA \
--prompt prompts/ba.txt \
--threshold 0.15
- name: Run prompt regression — PM agent
env:
ANTHROPIC_API_KEY: ${{ secrets.ANTHROPIC_API_KEY }}
run: |
antcrew regtest baseline.db \
--run "${{ vars.BASELINE_RUN_ID }}" \
--agent PM \
--prompt prompts/pm.txt \
--threshold 0.20
- name: Governance hash gate
run: |
antcrew verify-hash team.yaml \
--expected "${{ vars.APPROVED_TEAM_HASH }}"
What to set in your repo:
| Secret / Variable | Value |
|---|---|
secrets.ANTCREW_API_KEY |
API key for the platform instance |
secrets.ANTHROPIC_API_KEY |
Provider API key for replaying calls |
vars.BASELINE_RUN_ID |
Run ID of the approved baseline |
vars.APPROVED_TEAM_HASH |
Team hash from antcrew verify-hash team.yaml |
Recommended workflow¶
- Record a baseline run after each major prompt release and store the run ID in CI variables (
vars.BASELINE_RUN_ID) - Run
antcrew verify-hash team.yamllocally, copy theteam_hash, and store it invars.APPROVED_TEAM_HASH - On every PR that touches a prompt file or team YAML, both gates run automatically
- Set threshold per agent based on sensitivity (BA: 15%, code generators: 25%)
- Block merge if any agent exceeds its threshold or the governance hash has changed
Threshold guidelines¶
| Agent type | Recommended threshold |
|---|---|
| Business Analyst / PM | 10–15% — precise specs, low tolerance |
| Reviewer / QA | 15–20% — structured feedback, moderate |
| Code generators | 20–30% — stylistic variation acceptable |
| Content agents | 25–35% — creative flexibility expected |
Related¶
antcrew trace— inspect TraceLog, show stored prompts- Compliance Audit Trail — governance hash + attestation
- Governance Hash — certify agent configuration identity