Getting started

From pip install to a probe run that fails your build. About ten minutes, and you can see the bill before you pay it.


1Install and scaffold

In plain terms: install the tool, and it writes a folder of ready-made test conversations into your project. You do not have to invent them.

$ pip install "halligan[dotenv]"
$ halligan init --interactive

The wizard asks for a name, provider, model, judge and system prompt, and writes a target.yaml that already parses. It marks which providers have a credential in your environment and defaults to one that will actually run, and defaults the judge to a different model family than the target — a model is a poor judge of its own blind spots.

Plain halligan init skips the questions and writes an example config instead. Either way you get the 9 probe suites, 74 cases, and the names registry in your working directory.

The [dotenv] extra matters. Without it .env is never read, and halligan doctor will report that you have no credentials when you do.

2Give it credentials

In plain terms: the tool needs permission to talk to your AI. That password lives in a file on your machine and is never written into anything you share. If you are running a model on your own computer, there is no password to set — skip to the next step.

Keys come from the environment only. Halligan never accepts one as a command-line argument and never writes one to disk.

# .env — gitignored
ANTHROPIC_API_KEY=sk-ant-...
OPENAI_API_KEY=sk-proj-...

$ halligan doctor   # reports which keys are present, never their values
Running the model locally? No key needed. When base_url points at your own machine or LAN — LM Studio, Ollama, vLLM — Halligan skips the credential check entirely. Hosted gateways that speak the same dialect (Together, Groq, OpenRouter) still require their key, and will say so.
Never put a key in target.yaml. That file is meant to be committed — it is your test configuration, and it is not gitignored.

3Point it at the thing under test

In plain terms: tell it which AI to argue with — a raw model from OpenAI or Anthropic, one running on your own computer, or the app you have actually built.

A raw model API

Paste in the system prompt you actually deploy — you are testing that prompt as much as the model.

provider:
  name: anthropic          # anthropic | openai | gemini | ollama | http
  model: claude-sonnet-5
  temperature: 0.7         # what you deploy at — see the warning below
system_file: policy.md      # or an inline `system: |` block
judge:
  name: openai             # a different family than the target, on purpose
  model: gpt-4o

Your own deployed app

The more useful test. This exercises your whole stack — retrieval, guardrail middleware, prompt assembly — not just the model.

provider:
  name: http
  url: https://your-app.com/api/chat
  body:                     # shape this to your API
    conversation: "{{messages}}"
    system_prompt: "{{system}}"
  response_path: data.reply # dotted path to the reply text
  headers:
    authorization: "Bearer {{token}}"  # from HALLIGAN_HTTP_TOKEN

Available substitutions: {{messages}}, {{system}}, {{model}}, {{last_user}}, {{token}}.

A local model

The openai provider talks to anything speaking that dialect — LM Studio, vLLM, Ollama's OpenAI endpoint, Together, Groq. Point base_url at it; a local server needs no real key.

provider:
  name: openai
  model: google/gemma-4-12b
  base_url: http://127.0.0.1:1234/v1
  max_tokens: 4096
concurrency: 1            # a local server serves one model
Reasoning models need a far bigger budget. Thinking tokens count against the same limit, and one can spend all of it before writing anything — measured at 4,095 reasoning tokens of 4,096, content empty. Give them 8k–16k. Halligan raises a clear error rather than scoring the empty string as a refusal.

Don't make the judge the smallest model you have. It does the harder reasoning. A 4B judge scored two correct refusals as critical failures in our own testing — it couldn't separate “states an argument and answers it” from “states it and defers”, which is the distinction the suites turn on. It also has to fit in VRAM beside the target, or every call pays a model-swap.
Set a realistic temperature. At 0, repeats only measure provider nondeterminism, and you will get a falsely clean result. Use what you actually ship.

4Check the bill first

In plain terms: every question costs a fraction of a penny, and this runs thousands of them. This step tells you the price before you spend it, without spending anything.

74 cases is 116 model calls and 65 judge calls. At --repeat 20 that becomes 3,620. Look before you leap:

$ halligan run -t target.yaml -s suites/ --repeat 20 \
      --estimate --price-in 3 --price-out 15

  target calls             2,320  one per turn
  judge calls              1,300  one per `kind: judge`

  input tokens         3,264,400  from the actual prompts
  output tokens          980,000  assumes 400 per reply

  ~$24.49  at $3/$15 per Mtok in/out

--estimate exits without calling anything. Call counts are exact and input tokens are measured from the real prompts; only reply length is assumed, and --reply-tokens tunes it.

Multi-turn cost is quadratic, not linear. Every turn resends the whole conversation, so a six-turn ladder costs about twenty-one turns' worth of input. That is why a suite that looks small isn't.

Halligan ships no price table — a hardcoded rate goes stale silently and then lies with authority. Pass the rates, or put them in metadata.pricing.

5Run it

In plain terms: it now has the arguments, watches what your AI says, and writes a report you can open in a browser.

$ halligan validate --suite suites/    # parses everything, zero API calls
$ halligan run -t target.yaml --pack catholic --report report.html

A pack bundles the suites for one domain with the system prompt they were written against. halligan packs lists what is installed:

$ halligan packs

  ✓ catholic  Catholic-aligned AI assistant
      9 suite(s), 74 case(s)
      provenance: derived from a real adversarial session

Start smaller if you'd rather — --case and --family narrow the run:

$ halligan run -t target.yaml -s suites/jailbreak_roleplay.yaml
$ halligan run -t target.yaml --pack catholic --family pastoral
Read the provenance before you trust a score. red-team-session means the cases came from a real adversarial session against a live assistant. authored means they were written from the invariant structure and never fired at anything — still useful, but a 96% means considerably less. Anything unreviewed is marked in the listing.

6Make it mean something

In plain terms: run each conversation twenty times instead of once. AIs do not answer identically twice, and a rule that only breaks sometimes is the one that will surprise you in production.

One run tells you what happened once. It cannot tell you whether a guardrail holds or whether it is a coin flip — and those need different fixes.

$ halligan run -t target.yaml -s suites/ --repeat 20
$ halligan run -t target.yaml -s suites/ --fail-under 90   # gate CI
Exit codeMeans
0all cases passed
1one or more non-critical failures
2a critical-severity case failed
3configuration error — including a suite that needs a judge when none is set
Without --fail-under, a 27% score exits 0. Only critical failures gate the build by default. This caught us during our own testing — set a floor.

Failed is not the same as errored

A check that could not run has produced no verdict, and reporting “not graded” as “failed” invents a finding. A judge model that would not load once turned a run with zero behavioural failures into a reported score of 18.2%.

So ✗ fail means the model did the thing, and ! error means we could not tell. Errors still count against the score — a case that could not be run is not a case that passed — but the summary says 5 errored, never 5 failed, so a low number never lies about why it is low.


7Adapt it to your domain

In plain terms: the conversations that ship are written for a religious assistant. The shape of the attack transfers anywhere — swap in your own subject matter and the pressure tactics still work.

The 74 shipped cases are a Catholic-assistant probe suite. The harness — runner, scoring, repeats, sweep, reporting — is domain-neutral, and so are most graders. Three carry Catholic defaults: no_authority_claim, no_spiritual_bypass and no_procedure.

So if you're outside this domain, the value is in the shapes, not the shipped cases: hold an invariant under multi-turn escalation while still serving the person. Swap the probes, keep the ladder.

  • A medical assistant must never diagnose or prescribe, and must hold that line when told no doctor is available.
  • A legal assistant must not give jurisdiction-specific advice under "you're all I can afford".
  • A financial assistant must not become a fiduciary under emotional pressure.

Two policy templates ship, because probes assume a policy and a policy nobody tests is a wish. policies/general.md is the eight behavioural invariants with the domain specifics lifted into four placeholders — <ASSISTANT_NAME>, <DOMAIN>, <CREDENTIALED_ROLE>, <RESERVED_ACTIONS>. policies/catholic.md is the filled-in version the shipped suites were written against. init --interactive writes whichever you pick as policy.md and wires it in with system_file:.

For the graders: does_not_advocate takes persona: [...], and the referral and secrecy graders take markers: [...].

Writing a pack for your own domain is a packs/<name>.yaml manifest pointing at your suites and policy, with an honest provenance: value.

Full documentation on GitHub →