This guide gets FastRAG running locally on the local profile: in-process ONNX embedding and reranking, containerised Qdrant, Redis, and Postgres, and Ollama llama3.2:latest for generation. Nothing on the retrieval path crosses the internet, which is what makes this the rig the sub-200ms numbers in latency.md are measured on.

For the hosted free-tier setup instead, see deployment.md.

Prerequisites

  • Python 3.12
  • uv
  • Docker Engine with Compose v2
  • Ollama running on the host
  • llama3.2:latest installed locally

Verify Ollama from the host:

ollama list
ollama run llama3.2:latest "Reply with ok."

Python development

Install local dependencies and run the fast test suite:

uv sync --extra dev --extra ingest --extra eval
uv run pytest
uv run ruff check .
uv run mypy

Configure environment

Create .env from the local template, which is already configured for Ollama and the in-process models:

cp .env.local.example .env

On Linux, Docker may need an explicit host gateway mapping for containers to reach the host Ollama daemon. If host.docker.internal does not resolve in your Docker setup, use the host gateway IP or run an OpenAI-compatible proxy container on the compose network.

Two things in that file are worth knowing about before you hit them:

  • FASTRAG_SEMANTIC_CACHE_ENABLED=false, because the plain redis:8 image has no RediSearch module. Exact caching still works. Swap in redis/redis-stack to enable it.
  • Voice input needs a Sarvam key even here - no speech model runs in-process. Leave FASTRAG_SARVAM_API_KEY blank to run text-only; the /v1/voice/* endpoints then return 503 and nothing else is affected. See voice.md.

The embedding model in this profile, BAAI/bge-base-en-v1.5, is English-only, so Indic queries retrieve poorly against either corpus. For multilingual queries without any hosted embedder, switch to the in-process E5 and Jina-reranker pair in providers.md, which needs about 4 GB of memory; the Jina cloud providers are the other route.

Required calibration and model artifacts

Production startup requires config/calibration.json, pinned dense/reranker revisions, and verified ONNX checksums. The example files are schema examples, not valid production gates.

Generate calibration from reviewed held-out data:

uv run python -m fastrag.calibrate \
  --golden eval/calibration.jsonl \
  --cache-pairs eval/cache_pairs.jsonl

Download and verify configured model artifacts:

docker compose --profile tools run --rm model-init

For a disposable local wiring test, you can copy the example calibration and use local model paths/checksums only after accepting that quality gates are not meaningful:

cp config/calibration.example.json config/calibration.json

Do not ship with copied example calibration.

Start services

docker compose build
docker compose up -d

This starts the API, worker, Qdrant, Redis, Postgres, Caddy, Prometheus, and Grafana. Self-hosted Langfuse is not included: tracing points at Langfuse Cloud's free tier by default, which avoids running ClickHouse, MinIO, and a second Redis on your laptop. Add --profile langfuse and the commented block in .env.local.example if you want it locally.

Check health:

curl -fsS http://localhost/health/live
curl -fsS http://localhost/health/ready

Ingest documents

Place extractable-text PDFs, Markdown, or plain text files under a local document directory and ingest them:

uv run fastrag ingest docs/*.pdf docs/*.md docs/*.txt

The worker creates a shadow Qdrant collection, validates point count, swaps the kb_current alias, and activates the manifest in PostgreSQL. Image-only PDFs fail until OCR is added upstream.

Every strategy in FASTRAG_CHUNK_STRATEGIES is applied to each document and indexed into the same collection under a strategy payload field, so they can be compared at query time. See chunking.md.

The default corpus is the repository itself, indexed in English and served cross-lingually (or with machine-translated prose, if you pass --languages), and the golden, calibration and cache-pair sets are derived from the built index in the same run - see self-corpus.md:

uv run python scripts/ingest-self.py --dry-run      # inspect counts first
uv run python scripts/ingest-self.py --languages ""

For the multilingual MSMARCO-XI corpus instead, with a golden set derived from its own labels:

uv run python scripts/ingest-msmarco.py --rows-per-language 250 --max-chunks 90000

Query locally

curl -fsS \
  -H "Authorization: Bearer $FASTRAG_QUERY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"query":"How does CRAG decide to rewrite a query?"}' \
  http://localhost/v1/query

For streaming:

curl -N \
  -H "Authorization: Bearer $FASTRAG_QUERY_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"query":"How does CRAG decide to rewrite a query?"}' \
  http://localhost/v1/query/stream

Add "strategy" to compare chunking strategies, and "language" to filter to one language. GET /v1/strategies lists indexed strategies. Language accepts either form - bn or bn-IN

  • and is reduced to the ISO 639-1 code the chunk payloads carry.

After a query, open the website trace page at http://localhost:3000/query (run a question on the home page first; traces live in browser sessionStorage only). Use the Experiment panel to re-run with overrides; see query-trace.md.

Query overrides for the /query Experiment panel are on in .env.local.example (FASTRAG_ALLOW_QUERY_OVERRIDES=true). They also apply when FASTRAG_ENVIRONMENT=development and the flag is unset.

Smoke-test overrides:

FASTRAG_API_URL=http://localhost:8000 uv run python scripts/e2e_query_overrides.py
cd website && npx tsx scripts/check-build-overrides.ts

Run the frontends

Both apps proxy through /api/rag/[...path] so FASTRAG_QUERY_TOKEN stays server-side. FASTRAG_CORS_ORIGINS only matters if the browser calls the API directly.

Landing (website/) - marketing page with hero text/mic ask, chat answers, /query trace:

cd website
cp .env.example .env.local     # FASTRAG_API_URL=http://localhost:8000
npm install
npm run dev                    # http://localhost:3000

Run a question on /, then open /query to inspect the pipeline trace or re-run with edited config. See query-trace.md.

Console (web/) - latency, strategies, CRAG/guardrails, bench dashboard:

cd web
cp .env.example .env.local     # FASTRAG_API_URL=http://localhost:8000
npm install
npm run dev -- -p 3001         # http://localhost:3001

Use port 8000 when the API is host uvicorn; use http://localhost (no port) when Compose publishes through Caddy.

Local Ollama smoke test

Use the adapter-level smoke script to verify that FastRAG can stream from the installed llama3.2:latest model before starting the full stack:

uv run python scripts/smoke-ollama-provider.py

This checks the same OpenAI-compatible streaming path used by the production generator. It does not require Qdrant, Redis, PostgreSQL, Langfuse, or calibration.

To verify FastAPI, the query pipeline, Ollama generation, citation validation, and Langfuse trace flushing together without a production index, run:

LANGFUSE_PUBLIC_KEY=... \
LANGFUSE_SECRET_KEY=... \
LANGFUSE_BASE_URL=https://us.cloud.langfuse.com \
uv run python scripts/smoke-local-e2e.py

Running locally