Paperless-ngx turns a scanner into a searchable archive. Feed it PDFs and images, it runs OCR, extracts the text, and gives you full-text search, tags, correspondent/creator fields, and saved searches. Every document becomes a document.
What it is
Built for the "where did I put that tax form" problem. The web frontend handles scanning over the network, there's a consumption directory you can drop files into, and an email-fetching option for the truly committed. The OCR is local and good enough to search years of paperwork.
Before you start
- Decide on the storage layout up front:
dataholds the SQLite-free database + original files,mediaholds thumbnails and the rendered archive. - Set
USERMAP_UID/USERMAP_GIDto your own numeric IDs or downloaded files will be owned by an arbitrary user and unscriptable. - The Tesseract OCR needs a few dozen languages of data. If your documents aren't English, set
PAPERLESS_OCR_LANGUAGEup front — re-OCR later is a long job.
1 — Create the directories
mkdir -p ~/services/paperless/{data,media,export,consumption,pgdata}
cd ~/services/paperless
id -u; id -g # note these two numbers
2 — Write the compose file
services:
webserver:
image: ghcr.io/webtier/docker-paperless-ngx:latest
container_name: paperless-webserver
restart: unless-stopped
ports:
- "127.0.0.1:8010:8000"
volumes:
- ./data:/usr/src/paperless/data
- ./media:/usr/src/paperless/media
- ./export:/usr/src/paperless/export
- ./consumption:/usr/src/paperless/consumption
environment:
USERMAP_UID: 1000
USERMAP_GID: 1000
PAPERLESS_REDIS: redis://broker:6379
PAPERLESS_DBHOST: broker
PAPERLESS_DBPORT: 5432
PAPERLESS_DBNAME: paperless
PAPERLESS_DBUSER: paperless
PAPERLESS_DBPASS: CHANGE_ME
PAPERLESS_SECRET_KEY: CHANGE_ME
PAPERLESS_TIME_ZONE: UTC
PAPERLESS_OCR_LANGUAGE: eng+deu
PAPERLESS_URL: https://docs.example.com
depends_on:
- broker
- postgres
healthcheck:
test: ["CMD", "curl", "-fs", "http://localhost:8000"]
broker:
image: redis:7-alpine
container_name: paperless-redis
restart: unless-stopped
command: redis-server --requirepass CHANGE_ME --save ""
postgres:
image: postgres:16-alpine
container_name: paperless-db
restart: unless-stopped
environment:
POSTGRES_DB: paperless
POSTGRES_USER: paperless
POSTGRES_PASSWORD: CHANGE_ME
volumes:
- ./pgdata:/var/lib/postgresql/data
healthcheck:
test: ["CMD-SHELL", "pg_isready -U paperless"]
interval: 10s
timeout: 5s
retries: 5
3 — Start it
docker compose up -d
docker compose logs -f webserver | grep -i "ready\|error"
4 — First-run setup
- Hit
http://yourhost:8010and create the superuser. The usernamesuperuserand passwordsuperuserare auto-created on a clean install — change them immediately. - Configure the scanner. Enable Use TLS, enable Authentication, and set Scan pages to something sane like 150 DPI greyscale. You'll re-scan 200 pages at 300 DPI color otherwise.
- From Settings → Documents, run the Document type and Tag matching once with a few dozen docs in place. It learns from what you correct, and gets better the more you use it.
- Learn the consumption directory: copy or
scpany PDF into./consumptionand Paperless ingests, OCRs and files it without you touching the UI.
5 — Feed it email, and back it up
# put a PDF anywhere in the consumption dir
cp scan-2026-09-28.pdf ./consumption/
# document_exporter ships backups as a manager command
docker exec -u paperless paperless-webserver \
document_exporter --output-dir /usr/src/paperless/export
rsyncs a scan directory into consumption every few minutes and you'll stop thinking about the paperwork entirely.