Skip to content

Running with Docker

This is the recommended way to run LLM Extractinator

Docker bundles Python, Ollama, and the Studio into one container, so it's the simplest and most reliable setup — nothing to install on your host but Docker itself. Prefer a local Python install? See Installation.

This project ships with a GPU-ready Docker image so you can run everything in a consistent environment without installing all dependencies on your host.

Below is how to:

  1. understand what Docker is,
  2. install Docker,
  3. create the local folders that will be mounted into the container,
  4. run the container (Windows/PowerShell and Linux/macOS),
  5. switch between the two modes the image supports (app and shell).

1. What is Docker?

Docker lets you run apps in containers: lightweight, isolated environments that bundle the OS libraries and dependencies your app needs. You get the same setup everywhere (your laptop, CI, a server), so “it works on my machine” stops being a problem.

In this repo, the image is built on top of an NVIDIA CUDA runtime image and already contains:

  • Python 3.11
  • your package (installed with pip install -e .)
  • Ollama (started automatically in the container)
  • an entrypoint script that can start the Streamlit app or drop you into a shell

2. Install Docker

Desktop users (Windows / macOS):

  • Install Docker Desktop from the official Docker site.
  • On Windows, make sure you can run docker from PowerShell.
  • If you want GPU support on Windows, you also need a recent NVIDIA driver and the Docker + WSL2 stack that supports GPU.

Linux users (Ubuntu, etc.):

  • Install the Docker Engine from your distro or from Docker’s official docs.
  • For GPU support, install the NVIDIA Container Toolkit so --gpus all works.

If --gpus all fails, check your driver/toolkit install.


3. Create local folders

The container expects to mount several folders from your host into /app/... inside the container. Create them once:

mkdir -p data examples tasks output ollama_models

These will map to:

  • ./data → /app/data
  • ./examples → /app/examples
  • ./tasks → /app/tasks
  • ./output → /app/output
  • ./ollama_models → /root/.ollama (optional, but recommended)

Anything the app writes there will persist on your machine.

By default, Ollama models are stored inside the container at /root/.ollama. This means every time you start a new container, you'll need to pull the models again.

To avoid this, mount a local directory to persist the Ollama models between container runs:

mkdir -p ollama_models

Then add -v ${PWD}/ollama_models:/root/.ollama (Windows/PowerShell) or -v $(pwd)/ollama_models:/root/.ollama (Linux/macOS) to your docker run command (see examples in section 4).


4. Run the container

4.1 Windows / PowerShell example

# Remove `--gpus all` if you don't have a GPU
docker run --rm --gpus all `
  -p 127.0.0.1:8501:8501 `
  -p 11434:11434 `
  -v ${PWD}/data:/app/data `
  -v ${PWD}/examples:/app/examples `
  -v ${PWD}/tasks:/app/tasks `
  -v ${PWD}/output:/app/output `
  -v ${PWD}/ollama_models:/root/.ollama `
  lmmasters/llm_extractinator:latest

4.2 Linux / macOS variant

# Remove `--gpus all` if you don't have a GPU
docker run --rm --gpus all \
  -p 127.0.0.1:8501:8501 \
  -p 11434:11434 \
  -v $(pwd)/data:/app/data \
  -v $(pwd)/examples:/app/examples \
  -v $(pwd)/tasks:/app/tasks \
  -v $(pwd)/output:/app/output \
  -v $(pwd)/ollama_models:/root/.ollama \
  lmmasters/llm_extractinator:latest

Open: http://127.0.0.1:8501


5. Connecting to an existing Ollama instance

By default the container runs its own Ollama and stores models in the ollama_models mount. If you already have Ollama running somewhere — on your host machine, or a shared GPU server — you can point LLM Extractinator at that instead, and skip the container's own model server and storage.

In that case, run a leaner container: drop the Ollama port (11434) and the ollama_models volume, since the container won't be serving or storing models itself.

Windows / PowerShell:

docker run --rm --gpus all `
  -p 127.0.0.1:8501:8501 `
  -v ${PWD}/data:/app/data `
  -v ${PWD}/examples:/app/examples `
  -v ${PWD}/tasks:/app/tasks `
  -v ${PWD}/output:/app/output `
  lmmasters/llm_extractinator:latest

Linux / macOS:

docker run --rm --gpus all \
  -p 127.0.0.1:8501:8501 \
  -v $(pwd)/data:/app/data \
  -v $(pwd)/examples:/app/examples \
  -v $(pwd)/tasks:/app/tasks \
  -v $(pwd)/output:/app/output \
  lmmasters/llm_extractinator:latest

Then tell LLM Extractinator where your Ollama server is:

  • In the Studio — open the Run tab and set Ollama server URL to your instance.
  • On the CLI — pass --ollama_host http://host:11434.

Typical URLs:

  • http://host.docker.internal:11434 — Ollama running on your host machine (Docker Desktop on Windows/macOS; on Linux add --add-host=host.docker.internal:host-gateway to the docker run command).
  • http://<server-ip>:11434 — Ollama on another machine (e.g. a shared GPU server).

Since inference runs elsewhere, the container needs no GPU

When you connect to an external Ollama, the container itself doesn't do any inference — you can drop --gpus all from the command above.

The model must already be pulled there

Pointed at an externally managed server, LLM Extractinator only connects — it won't start the server, pull models, or unload them. Make sure the model you request is already available on that instance (ollama pull <model> on the machine running Ollama). See --ollama_host for the full behaviour.


6. The two modes (from the Dockerfile)

Your Dockerfile defines an entrypoint script:

  • it always starts Ollama in the background: ollama serve ...
  • it then looks at the first argument to decide the mode
/entrypoint.sh app   # default
/entrypoint.sh shell # drop into a shell

So, by default, when you run:

docker run ... lmmasters/llm_extractinator:latest

it uses CMD ["app"] → starts the Streamlit “extractinator”.

If you want to drop into the container and poke around (with the package already installed and Ollama running), just pass shell at the end:

Windows / PowerShell:

docker run --rm --gpus all `
  -p 127.0.0.1:8501:8501 `
  -p 11434:11434 `
  -v ${PWD}/data:/app/data `
  -v ${PWD}/examples:/app/examples `
  -v ${PWD}/tasks:/app/tasks `
  -v ${PWD}/output:/app/output `
  -v ${PWD}/ollama_models:/root/.ollama `
  lmmasters/llm_extractinator:latest shell

Linux / macOS:

docker run --rm --gpus all \
  -p 127.0.0.1:8501:8501 \
  -p 11434:11434 \
  -v $(pwd)/data:/app/data \
  -v $(pwd)/examples:/app/examples \
  -v $(pwd)/tasks:/app/tasks \
  -v $(pwd)/output:/app/output \
  -v $(pwd)/ollama_models:/root/.ollama \
  lmmasters/llm_extractinator:latest shell

That will not start the Streamlit UI; instead you’ll get a bash shell inside the container with llm_extractinator installed.


7. Building the image yourself, and updating Ollama

A model released after your image was built will not run in it — Ollama has to know the architecture, and that means a newer Ollama binary. This is the most common reason a model that "should" work reports as unsupported.

Rebuilding normally does not fix it. The Dockerfile installs Ollama in a RUN layer, and Docker caches those by the command text alone. That text never changes, so the layer is reused indefinitely — and because it sits above COPY . /app, even a source change does not touch it. A plain docker build -t image:latest . will happily keep an Ollama from months ago.

Use the build script, which resolves the newest release and passes it in as a build argument. Changing that argument changes the layer, so Ollama is rebuilt when — and only when — the version has actually moved:

./build.sh                       # llm-extractinator:latest, newest Ollama
./build.sh myname:tag            # a name of your choosing
OLLAMA_VERSION=0.32.0 ./build.sh  # pin, to roll back or reproduce an older image

On Windows, build.sh runs in Git Bash or WSL. From PowerShell, use the twin script instead — it does exactly the same thing:

.\build.ps1
.\build.ps1 -Image myname:tag
.\build.ps1 -OllamaVersion 0.32.0   # pin, to roll back or reproduce

If ./build.sh reports "permission denied" after a fresh clone, Git did not record the executable bit (it never does on Windows). bash build.sh works regardless, and git update-index --chmod=+x build.sh fixes it for good.

The build prints the version it resolved, and ollama --version inside the running container confirms what you ended up with.

Your models are not affected

Models live in the ollama_models/ mount, not in the image, so rebuilding does not touch them — the new binary reads the same store. You will still need to ollama pull the new model, but that is a separate step from the rebuild.

On a large Ollama version jump the model store format can change. It is a plain directory, so either take a copy first or accept that a re-pull may be needed.

If you cannot reach GitHub

The script falls back to building without a version, letting Ollama's own install script choose. Note the caveat it prints: if the layer is already cached, that build will not change the Ollama in the image. Use docker build --no-cache when you need to be certain — it is slow, because it rebuilds CUDA and Python too, but it is unambiguous.


8. Notes

  • The image exposes two ports: 8501 (Streamlit) and 11434 (Ollama).
  • If you don’t have a GPU, you can try omitting --gpus all, but the image is CUDA-based, so GPU is the intended path.
  • If your Docker Desktop uses different volume mappings (e.g., Windows drive letters), adjust the -v paths accordingly.
  • docker build --pull refreshes the CUDA base image as well, which is worth doing occasionally for security updates. It is a separate axis from the Ollama version.