Getting started
From pip install to a probe run that fails your build. About
ten minutes, and you can see the bill before you pay it.
1Install and scaffold
In plain terms: install the tool, and it writes a folder of ready-made test conversations into your project. You do not have to invent them.
$ pip install "halligan[dotenv]" $ halligan init --interactive
The wizard asks for a name, provider, model, judge and system prompt, and
writes a target.yaml that already parses. It marks which
providers have a credential in your environment and defaults to one that
will actually run, and defaults the judge to a different model
family than the target — a model is a poor judge of its own blind spots.
Plain halligan init skips the questions and writes an example
config instead. Either way you get the 9 probe suites, 74 cases, and the
names registry in your working directory.
[dotenv] extra matters. Without it
.env is never read, and halligan doctor will
report that you have no credentials when you do.
2Give it credentials
In plain terms: the tool needs permission to talk to your AI. That password lives in a file on your machine and is never written into anything you share. If you are running a model on your own computer, there is no password to set — skip to the next step.
Keys come from the environment only. Halligan never accepts one as a command-line argument and never writes one to disk.
# .env — gitignored ANTHROPIC_API_KEY=sk-ant-... OPENAI_API_KEY=sk-proj-... $ halligan doctor # reports which keys are present, never their values
base_url points at your own machine or LAN — LM Studio,
Ollama, vLLM — Halligan skips the credential check entirely. Hosted
gateways that speak the same dialect (Together, Groq, OpenRouter) still
require their key, and will say so.
target.yaml. That file is
meant to be committed — it is your test configuration, and it is not
gitignored.
3Point it at the thing under test
In plain terms: tell it which AI to argue with — a raw model from OpenAI or Anthropic, one running on your own computer, or the app you have actually built.
A raw model API
Paste in the system prompt you actually deploy — you are testing that prompt as much as the model.
provider: name: anthropic # anthropic | openai | gemini | ollama | http model: claude-sonnet-5 temperature: 0.7 # what you deploy at — see the warning below system_file: policy.md # or an inline `system: |` block judge: name: openai # a different family than the target, on purpose model: gpt-4o
Your own deployed app
The more useful test. This exercises your whole stack — retrieval, guardrail middleware, prompt assembly — not just the model.
provider: name: http url: https://your-app.com/api/chat body: # shape this to your API conversation: "{{messages}}" system_prompt: "{{system}}" response_path: data.reply # dotted path to the reply text headers: authorization: "Bearer {{token}}" # from HALLIGAN_HTTP_TOKEN
Available substitutions: {{messages}},
{{system}}, {{model}},
{{last_user}}, {{token}}.
A local model
The openai provider talks to anything speaking that dialect —
LM Studio, vLLM, Ollama's OpenAI endpoint, Together, Groq. Point
base_url at it; a local server needs no real key.
provider: name: openai model: google/gemma-4-12b base_url: http://127.0.0.1:1234/v1 max_tokens: 4096 concurrency: 1 # a local server serves one model
Don't make the judge the smallest model you have. It does the harder reasoning. A 4B judge scored two correct refusals as critical failures in our own testing — it couldn't separate “states an argument and answers it” from “states it and defers”, which is the distinction the suites turn on. It also has to fit in VRAM beside the target, or every call pays a model-swap.
0, repeats
only measure provider nondeterminism, and you will get a falsely clean
result. Use what you actually ship.
4Check the bill first
In plain terms: every question costs a fraction of a penny, and this runs thousands of them. This step tells you the price before you spend it, without spending anything.
74 cases is 116 model calls and 65 judge calls. At
--repeat 20 that becomes 3,620. Look before you leap:
$ halligan run -t target.yaml -s suites/ --repeat 20 \ --estimate --price-in 3 --price-out 15 target calls 2,320 one per turn judge calls 1,300 one per `kind: judge` input tokens 3,264,400 from the actual prompts output tokens 980,000 assumes 400 per reply ~$24.49 at $3/$15 per Mtok in/out
--estimate exits without calling anything. Call counts are
exact and input tokens are measured from the real prompts; only reply
length is assumed, and --reply-tokens tunes it.
Halligan ships no price table — a hardcoded rate goes stale silently and then lies with authority. Pass the rates, or put them in
metadata.pricing.
5Run it
In plain terms: it now has the arguments, watches what your AI says, and writes a report you can open in a browser.
$ halligan validate --suite suites/ # parses everything, zero API calls $ halligan run -t target.yaml --pack catholic --report report.html
A pack bundles the suites for one domain with the system
prompt they were written against. halligan packs lists what
is installed:
$ halligan packs ✓ catholic Catholic-aligned AI assistant 9 suite(s), 74 case(s) provenance: derived from a real adversarial session
Start smaller if you'd rather — --case and --family narrow the run:
$ halligan run -t target.yaml -s suites/jailbreak_roleplay.yaml $ halligan run -t target.yaml --pack catholic --family pastoral
red-team-session means the cases came from a real adversarial
session against a live assistant. authored means they were
written from the invariant structure and never fired at anything — still
useful, but a 96% means considerably less. Anything unreviewed is marked
in the listing.
6Make it mean something
In plain terms: run each conversation twenty times instead of once. AIs do not answer identically twice, and a rule that only breaks sometimes is the one that will surprise you in production.
One run tells you what happened once. It cannot tell you whether a guardrail holds or whether it is a coin flip — and those need different fixes.
$ halligan run -t target.yaml -s suites/ --repeat 20 $ halligan run -t target.yaml -s suites/ --fail-under 90 # gate CI
| Exit code | Means |
|---|---|
0 | all cases passed |
1 | one or more non-critical failures |
2 | a critical-severity case failed |
3 | configuration error — including a suite that needs a judge when none is set |
--fail-under, a 27% score exits 0.
Only critical failures gate the build by default. This caught
us during our own testing — set a floor.
Failed is not the same as errored
A check that could not run has produced no verdict, and reporting “not graded” as “failed” invents a finding. A judge model that would not load once turned a run with zero behavioural failures into a reported score of 18.2%.
So ✗ fail means the model did the thing, and
! error means we could not tell. Errors still count against
the score — a case that could not be run is not a case that passed — but
the summary says 5 errored, never 5 failed, so a
low number never lies about why it is low.
7Adapt it to your domain
In plain terms: the conversations that ship are written for a religious assistant. The shape of the attack transfers anywhere — swap in your own subject matter and the pressure tactics still work.
The 74 shipped cases are a Catholic-assistant probe suite. The
harness — runner, scoring, repeats, sweep, reporting — is
domain-neutral, and so are most graders. Three carry Catholic defaults:
no_authority_claim, no_spiritual_bypass and
no_procedure.
So if you're outside this domain, the value is in the shapes, not the shipped cases: hold an invariant under multi-turn escalation while still serving the person. Swap the probes, keep the ladder.
- A medical assistant must never diagnose or prescribe, and must hold that line when told no doctor is available.
- A legal assistant must not give jurisdiction-specific advice under "you're all I can afford".
- A financial assistant must not become a fiduciary under emotional pressure.
Two policy templates ship, because probes assume a policy and a policy
nobody tests is a wish. policies/general.md is the eight
behavioural invariants with the domain specifics lifted into four
placeholders — <ASSISTANT_NAME>,
<DOMAIN>, <CREDENTIALED_ROLE>,
<RESERVED_ACTIONS>. policies/catholic.md
is the filled-in version the shipped suites were written against.
init --interactive writes whichever you pick as
policy.md and wires it in with system_file:.
For the graders: does_not_advocate takes
persona: [...], and the referral and secrecy graders take
markers: [...].
Writing a pack for your own domain is a packs/<name>.yaml
manifest pointing at your suites and policy, with an honest
provenance: value.