Settings & Flags Reference¶
This page provides a complete overview of all configuration options in LLM Extractinator.
It follows a professional documentation pattern:
- A quick summary table for fast scanning
- Detailed per‑flag descriptions for deeper understanding
1. CLI Flags Overview (Summary)¶
| Flag | Default | Description |
|---|---|---|
--task_id |
required | Selects which task JSON file to run. |
--run_name |
"run" |
Name used in logs and output folders. |
--n_runs |
1 |
Number of times to repeat the task. |
--verbose |
False |
Enables detailed logging. |
--overwrite |
False |
Overwrites existing outputs if enabled. |
--seed |
None |
Random seed for reproducibility. |
--model_name |
"phi4" |
Model used via Ollama. |
--ollama_host |
None |
Connect to an already-running Ollama server instead of managing one. |
--embedding_model |
"nomic-embed-text" |
Embedding model for few‑shot selection. |
--temperature |
auto | Sampling randomness; 0.0 normally, 0.6 for a thinking model. |
--top_k |
None |
Top‑K sampling. |
--top_p |
None |
Nucleus sampling. |
--num_predict |
auto | Maximum generated tokens; sized from the schema when unset. |
--max_context_len |
"max" |
Context length strategy. |
--max_context_cap |
None |
Upper bound on the context window in max/split mode. Must exceed --num_predict. |
--quantile |
0.8 |
Split point for --max_context_len split (short vs long cases). |
--reasoning_model |
False |
Forces reasoning‑model mode on (thinking models are auto‑detected). |
--no_reasoning |
False |
Forces reasoning‑model mode off, overriding auto‑detection. |
--num_examples |
0 |
Number of few‑shot examples. |
--chunk_size |
None |
Chunk size for long inputs. |
--test_run_size |
None |
Run on only the first N rows (or a random sample) of the test set. |
--test_run_random |
False |
With --test_run_size, sample randomly instead of taking the first N rows. |
--translate |
False |
Translate input to English first. |
--output_dir |
output/ |
Output location. |
--log_dir |
output/ |
Log location. |
--data_dir |
data/ |
Input data directory. |
--task_dir |
tasks/ |
Task JSON directory. |
--example_dir |
examples/ |
Few‑shot example directory. |
--translation_dir |
translations/ |
Translation output directory. |
2. Detailed CLI Flag Descriptions¶
--task_id¶
Type: int
Default: required
Selects which task JSON file to run, based on its numeric prefix
(e.g., Task001_*.json → --task_id 1).
--run_name¶
Type: str
Default: "run"
Human‑friendly name used to structure log and output folders.
--n_runs¶
Type: int
Default: 1
Runs the same extraction multiple times—useful for testing stability or variance.
--verbose¶
Type: bool
Default: False
Prints additional diagnostic information during execution.
--overwrite¶
Type: bool
Default: False
If enabled, existing run results in the output folder will be overwritten. If disabled, the tool will skip processing if output already exists.
--seed¶
Type: int
Default: None
Random seed for reproducible behavior where possible.
--model_name¶
Type: str
Default: "phi4"
Name of the Ollama model to use (e.g., "phi4", "llama3.3", "deepseek-r1:8b"). See Ollama models for available options.
--ollama_host¶
Type: str
Default: None
Base URL of an already-running Ollama server, e.g. "http://localhost:11500" or "http://remote-machine:11434".
- Unset (default): llm_extractinator manages its own server — starting it, pulling
--model_name/--embedding_modelas needed, and stopping the model when the run finishes. - Set: the server is treated as externally managed — llm_extractinator only connects to it. It will not start/stop the server and will not pull models onto it, so make sure the model is already available there.
--embedding_model¶
Type: str
Default: "nomic-embed-text"
Name of the embedding model to use for few-shot example selection via semantic similarity. Only used when --num_examples > 0. See Ollama models for available embedding models (e.g., "mxbai-embed-large", "nomic-embed-text").
--temperature¶
Type: float
Default: chosen automatically
Left unset, the temperature is picked from the model: 0.0 (greedy, deterministic) for an ordinary model, 0.6 for a thinking one.
That second case is not a preference. Qwen's model cards state it plainly for thinking mode: "DO NOT use greedy decoding, as it can lead to performance degradation and endless repetitions." A repetition loop is a particularly bad failure for extraction — the model spends its whole --num_predict budget looping and returns no JSON at all. Qwen recommends 0.6 alongside --top_p 0.95 and --top_k 20; only the temperature is applied for you, so set those two yourself if you want the full recommended configuration.
Reproducibility is not lost: with --seed, a non-zero temperature is still deterministic run to run.
Pass a value to override, including --temperature 0 to force greedy on a thinking model anyway — the run logs a warning when you do.
--top_k¶
Type: int
Default: None
Restricts sampling to the top‑K highest‑probability tokens.
--top_p¶
Type: float
Default: None
Nucleus sampling: sample from the smallest token set whose cumulative probability ≥ p.
--num_predict¶
Type: int
Default: auto — sized from the output schema
Maximum number of tokens to generate for the model’s output.
Left unset, this is derived from your schema: a schema of thirty free-text fields needs far more room than one of three enums, and a truncated answer loses the whole row rather than degrading gracefully. The derived value never goes below 512, so it can only ever raise the budget relative to the old fixed default. A thinking model gets a further allowance on top, because its chain of thought is charged to this same budget.
Set it explicitly to override. If output comes back truncated on a model that
reasons regardless of think: false, raising it is the fix.
--max_context_len¶
Type: str or int
Default: "max"
Controls context length policy:
- "max" — use maximum available length
- "split" — split dataset in two by input size, running each subset with a right-sized context (good when lengths vary a lot); results are merged afterwards
- integer — explicitly set context length
--max_context_cap¶
Type: int
Default: None
An upper bound on the context window when --max_context_len is max or split. The window is still fitted to your data; it just never grows past this value. Inputs longer than the cap are truncated.
It must be larger than --num_predict: num_ctx is the whole window, so the prompt and the generated answer share it, and a ceiling at or below the output budget would leave no room for your documents at all.
This is the knob for fitting a run into a given GPU. max sizes the window to the longest document in the dataset, which is right for accuracy but says nothing about what your card can hold — and the context window, not the model weights, is usually what turns a run that fits into one that runs out of memory. The Studio's hardware presets set this for you.
--quantile¶
Type: float
Default: 0.8
Only used with --max_context_len split. Sets the token-count quantile that divides "short" from "long" cases — at 0.8, the longest 20% are run as the long subset. Must be between 0 and 1.
--reasoning_model¶
Type: bool
Default: False
For models like DeepSeek‑R1 and Qwen3 that output chain‑of‑thought before JSON, so the reasoning is routed away from the answer instead of swamping it.
You usually don't need this flag: models already installed in Ollama are inspected for a thinking capability and reasoning mode is switched on automatically. Set it when the model has not been pulled yet, or when the server can't be reached, so auto‑detection can't see the model. It also raises the generation budget — see --num_predict.
--no_reasoning¶
Type: bool (flag)
Default: False
Forces reasoning mode off. This overrides both auto‑detection and --reasoning_model, so it's the way to make a thinking model answer directly. Passing both flags together logs a warning and --no_reasoning wins.
On a model that advertises thinking, this sends Ollama an explicit think: false. On a model that doesn't, it does nothing at all — Ollama rejects think outright for those, and there is no reasoning to disable anyway.
A few models (qwen3.5 among them) reason regardless of think: false. Their reasoning is stripped from the output before parsing so extraction still works, but it is still generated, so it still spends the --num_predict budget. If output comes back truncated on such a model, either raise --num_predict or leave reasoning enabled.
--num_examples¶
Type: int
Default: 0
Number of few‑shot examples to include in the prompt.
Requires setting Example_Path inside the task JSON file.
--chunk_size¶
Type: int
Default: None
Splits the dataset into chunks of this many documents for processing. Useful for very large datasets as the chunks are saved incrementally. If a crash occurs, only the current chunk needs to be reprocessed.
--test_run_size¶
Type: int
Default: None
Runs on only the first N rows of the test set instead of the full dataset — a quick sanity check before committing to a full run. The context window is still sized from the full dataset, so what the subset exercises matches what the real run will do. Combine with --test_run_random to sample randomly instead. Output is written to its own folder, <task_name>-test<N>-run<idx>, so it never collides with a full run's output — and because the row count is part of the name, repeating a test run at a different size writes somewhere new instead of silently returning the earlier, smaller run's results. Not compatible with --max_context_len split, which falls back to max automatically.
--test_run_random¶
Type: bool (flag)
Default: False
With --test_run_size, sample rows randomly instead of taking the first N. Reproducible via --seed.
--translate¶
Type: bool
Default: False
If enabled, input is translated to English before extraction—adds an extra model step. Not recommended!
--output_dir¶
Type: Path
Default: output/
Where extracted results are written.
--log_dir¶
Type: Path
Default: output/
Location for logs; defaults to the output directory.
--data_dir¶
Type: Path
Default: data/
Directory containing datasets referenced by the Data_Path in task JSON files.
--task_dir¶
Type: Path
Default: tasks/
Folder containing task JSON files.
--example_dir¶
Type: Path
Default: examples/
Directory referenced by Example_Path in task JSON files.
--translation_dir¶
Type: Path
Default: translations/
Folder where translated versions of inputs are saved when using --translate.
3. Task Configuration Files¶
Task files define what to extract and how to parse it. They must be named:
Task<NNN>...json
where <NNN> is a zero-padded three-digit ID. Anything may follow the ID (an optional _name is common) as long as it isn't another digit, so all of these are valid and share ID 1:
Task001.json(what the Studio saves)Task001_products.json
IDs must be unique within a task directory. Reference a task by its integer ID: --task_id 1.
Required Fields¶
Description¶
Type: str
Short human‑readable explanation of the task.
Data_Path¶
Type: str
Relative path (from data_dir) to the dataset file.
Input_Field¶
Type: str
Column name or JSON key containing the text that should be extracted.
Parser_Format¶
Type: str
Filename of the parser module inside tasks/parsers/ that defines a Pydantic OutputParser model.
OutputParser is the schema Extractinator validates the LLM output against.
Optional Fields¶
Example_Path¶
Type: str
Relative path (from example_dir) to few‑shot examples.
Required only if using --num_examples > 0.
4. Additional Commands¶
build-parser¶
Opens the Output Schema Builder — a Streamlit tool for building the Pydantic OutputParser model visually and saving it to tasks/parsers/.
launch-extractinator¶
Opens the Studio, the full Streamlit app for assembling datasets, output schemas, and tasks, then running them and inspecting results. See Studio.