This guide gets FastRAG running locally on the local profile: in-process ONNX embedding and
reranking, containerised Qdrant, Redis, and Postgres, and Ollama llama3.2:latest for
generation. Nothing on the retrieval path crosses the internet, which is what makes this the
rig the sub-200ms numbers in latency.md are measured on.
For the hosted free-tier setup instead, see deployment.md.
Prerequisites
- Python 3.12
uv- Docker Engine with Compose v2
- Ollama running on the host
llama3.2:latestinstalled locally
Verify Ollama from the host:
ollama list
ollama run llama3.2:latest "Reply with ok."
Python development
Install local dependencies and run the fast test suite:
uv sync --extra dev --extra ingest --extra eval
uv run pytest
uv run ruff check .
uv run mypy
Configure environment
Create .env from the local template, which is already configured for Ollama and the
in-process models:
cp .env.local.example .env
On Linux, Docker may need an explicit host gateway mapping for containers to reach the host
Ollama daemon. If host.docker.internal does not resolve in your Docker setup, use the host
gateway IP or run an OpenAI-compatible proxy container on the compose network.
Two things in that file are worth knowing about before you hit them:
FASTRAG_SEMANTIC_CACHE_ENABLED=false, because the plainredis:8image has no RediSearch module. Exact caching still works. Swap inredis/redis-stackto enable it.- Voice input needs a Sarvam key even here - no speech model runs in-process. Leave
FASTRAG_SARVAM_API_KEYblank to run text-only; the/v1/voice/*endpoints then return 503 and nothing else is affected. See voice.md.
The embedding model in this profile, BAAI/bge-base-en-v1.5, is English-only, so Indic
queries retrieve poorly against either corpus. For multilingual queries without any hosted
embedder, switch to the in-process E5 and Jina-reranker pair in
providers.md, which needs about 4 GB of
memory; the Jina cloud providers are the other route.
Required calibration and model artifacts
Production startup requires config/calibration.json, pinned dense/reranker revisions, and
verified ONNX checksums. The example files are schema examples, not valid production gates.
Generate calibration from reviewed held-out data:
uv run python -m fastrag.calibrate \
--golden eval/calibration.jsonl \
--cache-pairs eval/cache_pairs.jsonl
Download and verify configured model artifacts:
docker compose --profile tools run --rm model-init
For a disposable local wiring test, you can copy the example calibration and use local model paths/checksums only after accepting that quality gates are not meaningful:
cp config/calibration.example.json config/calibration.json
Do not ship with copied example calibration.
Start services
docker compose build
docker compose up -d
This starts the API, worker, Qdrant, Redis, Postgres, Caddy, Prometheus, and Grafana.
Self-hosted Langfuse is not included: tracing points at Langfuse Cloud's free tier by
default, which avoids running ClickHouse, MinIO, and a second Redis on your laptop. Add
--profile langfuse and the commented block in .env.local.example if you want it locally.
Check health:
curl -fsS http://localhost/health/live
curl -fsS http://localhost/health/ready
Ingest documents
Place extractable-text PDFs, Markdown, or plain text files under a local document directory and ingest them:
uv run fastrag ingest docs/*.pdf docs/*.md docs/*.txt
The worker creates a shadow Qdrant collection, validates point count, swaps the kb_current
alias, and activates the manifest in PostgreSQL. Image-only PDFs fail until OCR is added
upstream.
Every strategy in FASTRAG_CHUNK_STRATEGIES is applied to each document and indexed into the
same collection under a strategy payload field, so they can be compared at query time. See
chunking.md.
The default corpus is the repository itself, indexed in English and served cross-lingually
(or with machine-translated prose, if you pass --languages), and the golden, calibration and
cache-pair sets are derived from the built index in the same run - see
self-corpus.md:
uv run python scripts/ingest-self.py --dry-run # inspect counts first
uv run python scripts/ingest-self.py --languages ""
For the multilingual MSMARCO-XI corpus instead, with a golden set derived from its own labels:
uv run python scripts/ingest-msmarco.py --rows-per-language 250 --max-chunks 90000
Query locally
curl -fsS \
-H "Authorization: Bearer $FASTRAG_QUERY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"query":"How does CRAG decide to rewrite a query?"}' \
http://localhost/v1/query
For streaming:
curl -N \
-H "Authorization: Bearer $FASTRAG_QUERY_API_KEY" \
-H "Content-Type: application/json" \
-d '{"query":"How does CRAG decide to rewrite a query?"}' \
http://localhost/v1/query/stream
Add "strategy" to compare chunking strategies, and "language" to filter to one language.
GET /v1/strategies lists indexed strategies. Language accepts either form - bn or bn-IN
- and is reduced to the ISO 639-1 code the chunk payloads carry.
After a query, open the website trace page at http://localhost:3000/query (run a question
on the home page first; traces live in browser sessionStorage only). Use the Experiment
panel to re-run with overrides; see query-trace.md.
Query overrides for the /query Experiment panel are on in .env.local.example
(FASTRAG_ALLOW_QUERY_OVERRIDES=true). They also apply when
FASTRAG_ENVIRONMENT=development and the flag is unset.
Smoke-test overrides:
FASTRAG_API_URL=http://localhost:8000 uv run python scripts/e2e_query_overrides.py
cd website && npx tsx scripts/check-build-overrides.ts
Run the frontends
Both apps proxy through /api/rag/[...path] so FASTRAG_QUERY_TOKEN stays server-side.
FASTRAG_CORS_ORIGINS only matters if the browser calls the API directly.
Landing (website/) - marketing page with hero text/mic ask, chat answers, /query trace:
cd website
cp .env.example .env.local # FASTRAG_API_URL=http://localhost:8000
npm install
npm run dev # http://localhost:3000
Run a question on /, then open /query to inspect the pipeline trace or re-run with
edited config. See query-trace.md.
Console (web/) - latency, strategies, CRAG/guardrails, bench dashboard:
cd web
cp .env.example .env.local # FASTRAG_API_URL=http://localhost:8000
npm install
npm run dev -- -p 3001 # http://localhost:3001
Use port 8000 when the API is host uvicorn; use http://localhost (no port) when Compose
publishes through Caddy.
Local Ollama smoke test
Use the adapter-level smoke script to verify that FastRAG can stream from the installed
llama3.2:latest model before starting the full stack:
uv run python scripts/smoke-ollama-provider.py
This checks the same OpenAI-compatible streaming path used by the production generator. It does not require Qdrant, Redis, PostgreSQL, Langfuse, or calibration.
To verify FastAPI, the query pipeline, Ollama generation, citation validation, and Langfuse trace flushing together without a production index, run:
LANGFUSE_PUBLIC_KEY=... \
LANGFUSE_SECRET_KEY=... \
LANGFUSE_BASE_URL=https://us.cloud.langfuse.com \
uv run python scripts/smoke-local-e2e.py
Running locally