Development¶
Setting up¶
git clone https://github.com/DIAGNijmegen/llm_extractinator
cd llm_extractinator
pip install -e ".[test]"
That is enough for the offline test suite. Running against a real model additionally needs Ollama and at least one model pulled.
Two test suites, two different jobs¶
The distinction matters more than it looks, and knowing which one to reach for saves a lot of time.
pytest — the offline suite¶
pytest # the whole offline suite, a few seconds
pytest tests/test_budget.py -v
Every test here runs with the model faked out: no GPU, no Ollama, no model downloads. That is what lets it run in seconds and in CI on every push.
It proves the plumbing. Prompts build, the schema binds, output parses, results merge back onto the input rows, files land on disk, and the context window arithmetic adds up. What it cannot prove is whether a real model produces good extractions — the faked model returns whatever the test told it to.
A bare pytest deselects the integration marker, matching CI. To opt in:
pytest -m integration # needs Ollama running and a model pulled
devtests/ — running against a real model¶
python devtests/run.py --model phi4
Ten probes, each covering a question the offline suite structurally cannot answer, because a faked model always returns what it was told:
- Does grammar-constrained decoding really hold the model to an enum?
- Is the schema-derived
num_predictactually enough for a wide schema? - How far does the token estimate drift on Dutch?
- Do the field descriptions in the prompt change what comes back?
Pointing it at hardware:
# Manage Ollama automatically — start a server if needed, pull, stop afterwards
python devtests/run.py --model phi4
# Connect to a server that is already running; nothing started, pulled or stopped
python devtests/run.py --model phi4 --host http://localhost:11434
# Compare two models in one run
python devtests/run.py --model qwen3:4b --model phi4
# Three rows per dataset, for a quick shake-down
python devtests/run.py --model phi4 --limit 3
python devtests/run.py --list # what each probe is for
--host is the only difference between local hardware and a shared or remote
box. A full sweep is around 100 generations — a couple of minutes on a small
model, longer on a large one.
It reports numbers rather than asserting them. There are no golden outputs and no accuracy thresholds: a real model's output is not stable enough to pin, and a threshold in a test only teaches you to ignore a red suite. The exit code is non-zero only for hard breakage — an exception, or a probe where nothing parsed at all.
Two runs can be compared directly, which is the point of the JSON report:
python devtests/compare.py devtests/reports/A.json devtests/reports/B.json
devtests/README.md has the full table of probes, an explanation of every
column, and the two numbers worth acting on.
Which one do I want?¶
| I want to know… | use |
|---|---|
| Did I break the pipeline? | pytest |
| Does the context budget still add up? | pytest |
| Does a real model actually fill these fields? | devtests/ |
| Is this schema too wide for the output budget? | devtests/ |
| Did my prompt change help or hurt? | devtests/ twice, then compare.py |
| Is the token estimate accurate for my language? | devtests/, read the ratio column |
Conventions worth knowing¶
Assert invariants, not numbers. Token counts move whenever the prompt or the
estimator changes. Tests that pin exact figures break on every legitimate change
and get deleted; tests that pin a relationship — prompt + num_predict <=
num_ctx — survive and keep catching things. Where a test does need a figure, it
derives it from the live value rather than hard-coding it.
PredictionTask.REQUIRED_PARAMS is an allowlist, and omission is silent. A
setting added to TaskConfig but not listed there never reaches the task, with
no error. tests/test_lifecycle.py pins both directions, so adding a setting
forces a decision rather than defaulting to "not passed" by accident.
Keep the docs in the same PR. If you change task naming, required fields, or CLI flags, update the docs alongside the code — a flag whose documented default is wrong is worse than one that is undocumented.