Running with Docker¶
This is the recommended way to run LLM Extractinator
Docker bundles Python, Ollama, and the Studio into one container, so it's the simplest and most reliable setup — nothing to install on your host but Docker itself. Prefer a local Python install? See Installation.
This project ships with a GPU-ready Docker image so you can run everything in a consistent environment without installing all dependencies on your host.
Below is how to:
- understand what Docker is,
- install Docker,
- create the local folders that will be mounted into the container,
- run the container (Windows/PowerShell and Linux/macOS),
- switch between the two modes the image supports (
appandshell).
1. What is Docker?¶
Docker lets you run apps in containers: lightweight, isolated environments that bundle the OS libraries and dependencies your app needs. You get the same setup everywhere (your laptop, CI, a server), so “it works on my machine” stops being a problem.
In this repo, the image is built on top of an NVIDIA CUDA runtime image and already contains:
- Python 3.11
- your package (installed with
pip install -e .) - Ollama (started automatically in the container)
- an entrypoint script that can start the Streamlit app or drop you into a shell
2. Install Docker¶
Desktop users (Windows / macOS):
- Install Docker Desktop from the official Docker site.
- On Windows, make sure you can run
dockerfrom PowerShell. - If you want GPU support on Windows, you also need a recent NVIDIA driver and the Docker + WSL2 stack that supports GPU.
Linux users (Ubuntu, etc.):
- Install the Docker Engine from your distro or from Docker’s official docs.
- For GPU support, install the NVIDIA Container Toolkit so
--gpus allworks.
If
--gpus allfails, check your driver/toolkit install.
3. Create local folders¶
The container expects to mount several folders from your host into /app/... inside the container. Create them once:
mkdir -p data examples tasks output ollama_models
These will map to:
./data→/app/data./examples→/app/examples./tasks→/app/tasks./output→/app/output./ollama_models→/root/.ollama(optional, but recommended)
Anything the app writes there will persist on your machine.
Persisting Ollama models (recommended)¶
By default, Ollama models are stored inside the container at /root/.ollama. This means every time you start a new container, you'll need to pull the models again.
To avoid this, mount a local directory to persist the Ollama models between container runs:
mkdir -p ollama_models
Then add -v ${PWD}/ollama_models:/root/.ollama (Windows/PowerShell) or -v $(pwd)/ollama_models:/root/.ollama (Linux/macOS) to your docker run command (see examples in section 4).
4. Run the container¶
4.1 Windows / PowerShell example¶
# Remove `--gpus all` if you don't have a GPU
docker run --rm --gpus all `
-p 127.0.0.1:8501:8501 `
-p 11434:11434 `
-v ${PWD}/data:/app/data `
-v ${PWD}/examples:/app/examples `
-v ${PWD}/tasks:/app/tasks `
-v ${PWD}/output:/app/output `
-v ${PWD}/ollama_models:/root/.ollama `
lmmasters/llm_extractinator:latest
4.2 Linux / macOS variant¶
# Remove `--gpus all` if you don't have a GPU
docker run --rm --gpus all \
-p 127.0.0.1:8501:8501 \
-p 11434:11434 \
-v $(pwd)/data:/app/data \
-v $(pwd)/examples:/app/examples \
-v $(pwd)/tasks:/app/tasks \
-v $(pwd)/output:/app/output \
-v $(pwd)/ollama_models:/root/.ollama \
lmmasters/llm_extractinator:latest
Open: http://127.0.0.1:8501
5. Connecting to an existing Ollama instance¶
By default the container runs its own Ollama and stores models in the ollama_models mount. If you already have Ollama running somewhere — on your host machine, or a shared GPU server — you can point LLM Extractinator at that instead, and skip the container's own model server and storage.
In that case, run a leaner container: drop the Ollama port (11434) and the ollama_models volume, since the container won't be serving or storing models itself.
Windows / PowerShell:
docker run --rm --gpus all `
-p 127.0.0.1:8501:8501 `
-v ${PWD}/data:/app/data `
-v ${PWD}/examples:/app/examples `
-v ${PWD}/tasks:/app/tasks `
-v ${PWD}/output:/app/output `
lmmasters/llm_extractinator:latest
Linux / macOS:
docker run --rm --gpus all \
-p 127.0.0.1:8501:8501 \
-v $(pwd)/data:/app/data \
-v $(pwd)/examples:/app/examples \
-v $(pwd)/tasks:/app/tasks \
-v $(pwd)/output:/app/output \
lmmasters/llm_extractinator:latest
Then tell LLM Extractinator where your Ollama server is:
- In the Studio — open the Run tab and set Ollama server URL to your instance.
- On the CLI — pass
--ollama_host http://host:11434.
Typical URLs:
http://host.docker.internal:11434— Ollama running on your host machine (Docker Desktop on Windows/macOS; on Linux add--add-host=host.docker.internal:host-gatewayto thedocker runcommand).http://<server-ip>:11434— Ollama on another machine (e.g. a shared GPU server).
Since inference runs elsewhere, the container needs no GPU
When you connect to an external Ollama, the container itself doesn't do any inference — you can drop --gpus all from the command above.
The model must already be pulled there
Pointed at an externally managed server, LLM Extractinator only connects — it won't start the server, pull models, or unload them. Make sure the model you request is already available on that instance (ollama pull <model> on the machine running Ollama). See --ollama_host for the full behaviour.
6. The two modes (from the Dockerfile)¶
Your Dockerfile defines an entrypoint script:
- it always starts Ollama in the background:
ollama serve ... - it then looks at the first argument to decide the mode
/entrypoint.sh app # default
/entrypoint.sh shell # drop into a shell
So, by default, when you run:
docker run ... lmmasters/llm_extractinator:latest
it uses CMD ["app"] → starts the Streamlit “extractinator”.
If you want to drop into the container and poke around (with the package already installed and Ollama running), just pass shell at the end:
Windows / PowerShell:
docker run --rm --gpus all `
-p 127.0.0.1:8501:8501 `
-p 11434:11434 `
-v ${PWD}/data:/app/data `
-v ${PWD}/examples:/app/examples `
-v ${PWD}/tasks:/app/tasks `
-v ${PWD}/output:/app/output `
-v ${PWD}/ollama_models:/root/.ollama `
lmmasters/llm_extractinator:latest shell
Linux / macOS:
docker run --rm --gpus all \
-p 127.0.0.1:8501:8501 \
-p 11434:11434 \
-v $(pwd)/data:/app/data \
-v $(pwd)/examples:/app/examples \
-v $(pwd)/tasks:/app/tasks \
-v $(pwd)/output:/app/output \
-v $(pwd)/ollama_models:/root/.ollama \
lmmasters/llm_extractinator:latest shell
That will not start the Streamlit UI; instead you’ll get a bash shell inside the container with llm_extractinator installed.
7. Building the image yourself, and updating Ollama¶
A model released after your image was built will not run in it — Ollama has to know the architecture, and that means a newer Ollama binary. This is the most common reason a model that "should" work reports as unsupported.
Rebuilding normally does not fix it. The Dockerfile installs Ollama in a
RUN layer, and Docker caches those by the command text alone. That text never
changes, so the layer is reused indefinitely — and because it sits above
COPY . /app, even a source change does not touch it. A plain
docker build -t image:latest . will happily keep an Ollama from months ago.
Use the build script, which resolves the newest release and passes it in as a build argument. Changing that argument changes the layer, so Ollama is rebuilt when — and only when — the version has actually moved:
./build.sh # llm-extractinator:latest, newest Ollama
./build.sh myname:tag # a name of your choosing
OLLAMA_VERSION=0.32.0 ./build.sh # pin, to roll back or reproduce an older image
On Windows, build.sh runs in Git Bash or WSL. From PowerShell, use the twin
script instead — it does exactly the same thing:
.\build.ps1
.\build.ps1 -Image myname:tag
.\build.ps1 -OllamaVersion 0.32.0 # pin, to roll back or reproduce
If ./build.sh reports "permission denied" after a fresh clone, Git did not
record the executable bit (it never does on Windows). bash build.sh works
regardless, and git update-index --chmod=+x build.sh fixes it for good.
The build prints the version it resolved, and ollama --version inside the
running container confirms what you ended up with.
Your models are not affected¶
Models live in the ollama_models/ mount, not in the image, so rebuilding does
not touch them — the new binary reads the same store. You will still need to
ollama pull the new model, but that is a separate step from the rebuild.
On a large Ollama version jump the model store format can change. It is a plain directory, so either take a copy first or accept that a re-pull may be needed.
If you cannot reach GitHub¶
The script falls back to building without a version, letting Ollama's own
install script choose. Note the caveat it prints: if the layer is already
cached, that build will not change the Ollama in the image. Use
docker build --no-cache when you need to be certain — it is slow, because it
rebuilds CUDA and Python too, but it is unambiguous.
8. Notes¶
- The image exposes two ports:
8501(Streamlit) and11434(Ollama). - If you don’t have a GPU, you can try omitting
--gpus all, but the image is CUDA-based, so GPU is the intended path. - If your Docker Desktop uses different volume mappings (e.g., Windows drive letters), adjust the
-vpaths accordingly. docker build --pullrefreshes the CUDA base image as well, which is worth doing occasionally for security updates. It is a separate axis from the Ollama version.