Skip to content

Latest commit

 

History

1,093 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LocalChat

Python 3.12 License: MIT CI Quality Gate Coverage

Chat with your own documents, using a language model that runs on your own hardware. Upload PDF, DOCX, PPTX, XLSX, TXT, Markdown or images; LocalChat chunks and embeds them into PostgreSQL with pgvector, and answers questions from what it retrieves. Nothing leaves the machine unless you enable web search or a cloud fallback.

Built with FastAPI, Ollama, PostgreSQL + pgvector and Redis. Hybrid semantic and lexical retrieval, a cross-encoder reranker, tool calling, streaming answers, per-workspace document isolation, and RAG parameters tunable at runtime.

Production-ready for what it claims to be, which is a specific thing. LocalChat is a single-node, self-hosted appliance for a small team of up to 25 users — see ADR-1. All eight exit criteria in the production plan are met and the hardening gate was lifted on 2026-08-31: fail-closed boot, authorisation enforced by default in CI, a concurrency budget, a mutation-tested security core, restore proven in CI, a reproducible tagged release, migrations executed rather than merely written, and documentation verified against the code.

The scope is the important half of that sentence. Multi-tenant SaaS and horizontal scaling are out of scope — running more than one replica breaks cache coherence and rate limiting silently, and the debt register in ROADMAP says exactly where. Read the criteria before relying on the label; they are a floor, not a warranty.


Quick start

git clone https://github.com/jwvanderstam/LocalChat
cd LocalChat
cp .env.example .env          # set ADMIN_PASSWORD and the secrets; see the note below
docker compose up -d          # PostgreSQL, Redis, Ollama and the app

Open http://localhost:5000. You will be asked to sign in.

The app image is built on a hardened, distroless base: no shell, no package manager, running as uid 65532. That changes how you debug it — docker exec ... sh will not work; use docker compose run --rm --entrypoint python app. See DEPLOYMENT.md.

Everything runs in Docker, on a private network where the services address each other by name — the app reaches Ollama at http://ollama:11434 and Postgres at db. Compose sets those itself, and compose's environment: beats .env, so OLLAMA_BASE_URL, PG_HOST and REDIS_HOST in your .env have no effect on the containers. Set them in docker-compose.yml (or an override file) if you need to point elsewhere. The .env values that do matter here are the secrets and tuning: ADMIN_PASSWORD, SECRET_KEY, JWT_SECRET_KEY, model names, limits.

Getting the first password. If you set ADMIN_PASSWORD in .env, use that with the username admin. If you left it empty, an admin account is seeded on first boot with a generated password, logged once:

docker compose logs app | grep ADMIN_PASSWORD

Then pull a model and select it under Models — without an active model, chat returns 400 No active model set:

docker compose exec ollama ollama pull llama3.2:latest

Upload a document under Documents, and ask about it under Chat.

Running the app outside Docker

The backing services still run in Docker — only the app moves to the host, which is useful for a debugger or a fast edit loop.

pip install -r requirements.txt
cp .env.example .env
docker compose up -d db redis ollama   # backing services only
python app.py

This is the one case where .env's localhost URLs apply: the host process reaches the containers over published loopback ports, 127.0.0.1:5432 for Postgres and 127.0.0.1:11434 for Ollama. Both are bound to loopback, not 0.0.0.0.

If something already owns one of those ports — a natively installed Ollama is the usual culprit — move the published port instead of editing docker-compose.yml:

OLLAMA_BIND_PORT=21434 docker compose up -d ollama
# then in .env:  OLLAMA_BASE_URL=http://localhost:21434

BIND_HOST moves every published port at once, and BIND_PORT moves the app's own.

What it does

Retrieval. Hybrid search combines an independent semantic arm (pgvector) and a lexical arm (PostgreSQL tsvector), blended by weight, then reranked by a cross-encoder. GraphRAG expands queries one hop through entity co-occurrences. Chunking preserves table structure, and PDF tables are extracted rather than flattened.

Answers. Streaming responses over SSE, with tool calling, optional live web search, and long-term memory extracted from earlier conversations. Model routing picks a model class per query; a cloud model can be configured as fallback.

Isolation and access. Documents, conversations and memories are scoped to a workspace. Access is role-based, both globally (admin or user) and per workspace (viewer, editor, owner). Workspace API keys give a bot or workflow scoped, revocable access to exactly one workspace.

Integrity. Records other data refers to are never hard-deleted. A delete sets deleted_at; purging is a separate, admin-only operation with preconditions. See the Clark-Wilson section in CLAUDE.md.

Sources. Document connectors for local folders, S3, SharePoint, OneDrive, Google Drive and webhooks. Plugins extend the application without modifying it, under an inward-only dependency contract.

How it works

flowchart TD
    Browser["Browser / API client"]

    subgraph FastAPI["FastAPI application"]
        Routes["Routes (APIRouters)"]
        Auth["Security — JWT · rate limit · CORS"]
        Pydantic["Pydantic validation + sanitization"]
        RAG["RAG pipeline — retrieval · reranking"]
        Tools["Tool executor — function calling"]
        SSE["SSE stream"]
    end

    subgraph Services["External services"]
        PG["PostgreSQL + pgvector"]
        Ollama["Ollama — LLM · embeddings"]
        Redis["Redis — cache · rate limiting"]
    end

    Browser -->|HTTP request| Routes
    Routes --> Auth
    Auth --> Pydantic
    Pydantic --> RAG
    RAG -->|vector search| PG
    RAG -->|embed query| Ollama
    Pydantic --> Tools
    Tools -->|tool-call loop| Ollama
    Tools --> RAG
    Ollama -->|stream tokens| SSE
    SSE -->|text/event-stream| Browser
    Routes -.->|cache r/w| Redis
Loading

A request is parsed and validated at the boundary, authorised against the workspace, then handed to a service. Routes hold no business logic and no SQL. Blocking work — retrieval, embedding, database writes — runs in a threadpool so one slow query cannot stall other requests.

Layer Technology
Web framework FastAPI + Uvicorn
Database PostgreSQL 16 + pgvector (psycopg3 pool)
LLM Ollama, local; LiteLLM cloud fallback
Embeddings nomic-embed-text
Cache Redis, with an in-memory fallback
Auth python-jose (JWT), slowapi (rate limiting)
Validation Pydantic 2
Migrations Alembic
Tests pytest + pytest-asyncio

Entry point is app.pycreate_app() in src/app_fastapi.py. Every module is listed in the module index.

Documentation

Full index: docs/README.md — organised by Diátaxis, so what you need depends on what you are doing.

I want to… Go to
Deploy this properly Deployment, Operations
Connect a bot or workflow Workspace API keys, Discord via n8n
Look up a setting Configuration, RAG settings
Understand a decision ADRs, Lessons learned
Fix something broken Troubleshooting
Contribute code CLAUDE.md and .claude/rules/

The API is documented live at /api/docs/ (Swagger UI), and the same documentation is browsable inside the application under Docs.

Development

ruff check src/ tests/                       # lint
mypy src --ignore-missing-imports            # types
bandit -r src/ -ll -q -c pyproject.toml      # security
pytest -m "not (slow or ollama or db)"       # fast suite, no external services

All four must be clean before a commit; CI enforces the same set. Merging to main requires five checks to pass — unit-tests, integration-tests, repo-hygiene, docker-smoke and perf-canary — and is a human decision; auto-merge is off deliberately. (docker-smoke joined the required set on 2026-08-19 and perf-canary on 2026-08-24; this sentence still said three until 2026-08-27. Read the set back with gh api repos/jwvanderstam/LocalChat/rulesets/14700924, not from the settings UI.)

Current state: 3,008 tests collected; the fast suite runs 2,909 of them (2,888 passed, 21 skipped) in about 12 minutes at 80.7% coverage. Integration tests need PostgreSQL; some also need Ollama, and tests/e2e/ drives a real browser. Every number here was measured on 2026-08-27 rather than remembered — see exit criterion 7 in PRODUCTION_PLAN.

Notable changes per release are in CHANGELOG.md.

Coding standards live in .claude/rules/: architecture, Python, testing, plugins.

Security

Report vulnerabilities per SECURITY.md. Route-by-route permissions are documented in PERMISSIONS.md.

Two properties worth knowing when reading logs or API errors: every user-controlled value that reaches a log record passes through sanitize_log_value, which strips control characters, the Unicode line separators and ESC so a crafted filename cannot forge a log record or drive the terminal of whoever reads it; and API error bodies return fixed messages, never exception text.

License

MIT — see LICENSE.

About

local chat application

Resources

Security policy

Stars

6 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages