Every external dependency sits behind a protocol in src/fastrag/ports.py,
so swapping a provider is a wiring change in src/fastrag/bootstrap.py
rather than a rewrite. FASTRAG_PROFILE picks a coherent default set; individual
FASTRAG_*_PROVIDER variables override any single choice.
The two profiles
local | cloud | |
|---|---|---|
| Embedding | FastEmbed BAAI/bge-base-en-v1.5, in-process ONNX | Jina jina-embeddings-v3 |
| Reranking | FastEmbed Xenova/ms-marco-MiniLM-L-6-v2, in-process | Jina jina-reranker-v2-base-multilingual |
| Vector DB | Qdrant container | Qdrant Cloud free |
| LLM | Ollama on the host | Groq free, OpenRouter fallback |
| Cache | Redis container | Redis Cloud free |
| Registry | Postgres container | Neon free |
| Tracing | Langfuse Cloud, or self-hosted via the langfuse compose profile | Langfuse Cloud |
| Speech-to-text | Sarvam (no STT model runs in-process) | Sarvam |
local is the benchmark rig: nothing crosses the internet on the retrieval path, which is
what makes the sub-200ms measurement meaningful. cloud is the live demo that fits free
tiers. Both are benchmarked and both sets of numbers are published - see
latency.md for why conflating them would be dishonest.
The local profile's default embedder, bge-base-en-v1.5, is an English model, so querying
in Hindi, Bengali, Tamil, Telugu or Marathi needs either Jina or the in-process multilingual
pair below.
Multilingual without a hosted embedder
Both halves of retrieval can run in-process and still cover all six languages. This is how the self-corpus is built, and it means neither indexing nor serving spends Jina balance:
FASTRAG_EMBEDDING_PROVIDER=fastembed
FASTRAG_RERANKER_PROVIDER=fastembed
FASTRAG_DENSE_MODEL_ID=intfloat/multilingual-e5-large
FASTRAG_DENSE_MODEL_REPOSITORY=qdrant/multilingual-e5-large-onnx
FASTRAG_DENSE_MODEL_REVISION=ac6781cd1cf88b8306a536d7c9d18a5bd57cc14b
FASTRAG_DENSE_MODEL_FILE=model.onnx_data
FASTRAG_DENSE_MODEL_SHA256=0cf1883fee81c63819a44e2ba0efa51d4043d9759685a4ebebbde97e0623d15c
FASTRAG_DENSE_DIMENSION=1024
FASTRAG_DENSE_QUERY_PREFIX="query: "
FASTRAG_DENSE_DOCUMENT_PREFIX="passage: "
FASTRAG_RERANKER_MODEL_ID=jinaai/jina-reranker-v2-base-multilingual
FASTRAG_RERANKER_MODEL_REPOSITORY=jinaai/jina-reranker-v2-base-multilingual
FASTRAG_RERANKER_REVISION=9cfeff2df7d40d1b78e75e5e9cebec92a99813c9
FASTRAG_RERANKER_MODEL_FILE=onnx/model.onnx
FASTRAG_RERANKER_SHA256=0ef3f7978f7bc52360864d74edc1a0e03d159af770a7767c4d5943496e616012
FASTRAG_RETRIEVAL_FALLBACK_LANGUAGES=en
The reranker is the same jina-reranker-v2-base-multilingual the hosted API serves, run
locally. The embedder is E5 rather than Jina's own jina-embeddings-v3, which FastEmbed can
also run: embedding the self-corpus with it on CPU peaked at 11.3 GB and was killed by the
kernel on a 15 GB machine.
Measured on the self-corpus's 437 English sections, with eight questions about FastRAG asked in each language, E5 put the section that answers the question in the top 20 for 46 of 48 queries and at rank 1 for at least six of eight in every language. The two misses were Tamil and Marathi phrasings of one guardrail question. Through the full path - language filter with the English fallback, then the local reranker against the calibrated abstention threshold - the same questions were answered 7/8 in English, 8/8 in Hindi, 7/8 in Bengali, Telugu and Marathi, and 4/8 in Tamil, where the reranker is least confident; the answering section was in the reranked top five for 45 of 48.
Details that matter:
- E5 is asymmetric. It is trained with
query:andpassage:markers, and FastEmbed adds neither, so both prefixes are configured. Prefixes are normalised to end in exactly one space, because they are part of the embedding fingerprint and hosting dashboards trim env values -query:arriving without its space would otherwise reject a valid index. - The corpus stays English. A cross-lingual embedder matches a Hindi query to English
chunks directly, but the
languagefilter is an exact match on the chunk payload, and no chunk sayshi.FASTRAG_RETRIEVAL_FALLBACK_LANGUAGES=enmakes alanguage=hiquery searchhiandenchunks. Leave it empty for a corpus translated into every language, such as MSMARCO-XI, where exact filtering is the point. - Checksum the weights, not the graph. E5's ONNX export keeps 2.2 GB of weights in
model.onnx_databeside a 0.5 MBmodel.onnx. Production verifies the file named byFASTRAG_DENSE_MODEL_FILE, so it names the weights. - Switch off the centroid off-topic gate. E5 embeddings are anisotropic: almost any text
scores about 0.8 cosine against the corpus centroid. Measured on the self-corpus, questions
about FastRAG scored 0.787-0.831 and questions such as "what is the capital of France?"
up to 0.828, so no threshold separates them; the calibrated 0.826 refused three of eight
English and five to seven of eight Indic questions. The reranker does separate them - 18
off-topic questions across the six languages scored at most -1.85 against an abstention
threshold of 0.26 - so set
FASTRAG_GUARDRAIL_OFFTOPIC_ENABLED=falseand let abstention returnno_answerfor off-topic input. The injection and language guardrails are unaffected. - It costs memory, and CPU reranking is slow. Both models together peak at 3.7 GB
resident; embedding a query takes 50-90 ms. Reranking 20 real self-corpus chunks - pairs
of up to 658 tokens, because code tokenises long - took 8.5 s on 16 CPU cores at
FastEmbed's default batch of 64 and 5.4 s with
FASTRAG_RERANKER_BATCH_SIZE=1, since a batch is padded to its longest pair. On an RTX 3050 Ti it took 0.47 s, with scores within 0.0011 of the CPU run and an identical ranking. That rules out Vercel functions and Render's free instance; see deployment.md. - Offline work belongs on a GPU if you have one. Ingest,
derive-eval.pyand calibration rerank or embed hundreds of texts. Withonnxruntime-gpuinstalled in place ofonnxruntime, setFASTRAG_DENSE_EXECUTION_PROVIDERSorFASTRAG_RERANKER_EXECUTION_PROVIDERStoCUDAExecutionProvider, one model per card on 4 GB: E5's weights take 2.2 GB, andFASTRAG_DENSE_BATCH_SIZE=4/FASTRAG_RERANKER_BATCH_SIZE=4keep attention buffers inside what is left. Deriving the evaluation sets took 71 minutes on 16 CPU cores and 16 minutes with the reranker on an RTX 3050 Ti. - Licence.
jina-reranker-v2-base-multilingualis published under CC BY-NC 4.0. Running the weights yourself is non-commercial use only; the hosted Jina API is the licensed route for commercial deployments. E5 is MIT.
Free tiers, and what each one costs you
Qdrant Cloud - 1 GB, roughly 250K vectors at 768d, permanent. Keeps named sparse vectors, so hybrid retrieval and RRF are unchanged from self-hosted. Only the URL and API key differ.
Jina AI - 10M tokens shared across embedding and reranking on one key. jina-embeddings-v3
is multilingual at 1024d and covers all five Indic languages. Changing embedding provider
changes the embedding fingerprint and therefore requires a re-index; this is enforced at
startup rather than discovered later through bad results.
Sarvam Saaras v3 - authenticates with an api-subscription-key header, not a bearer
token, which is the usual first thing to get wrong. See voice.md.
Groq - OpenAI-compatible, so adapters/generation.py
needs no changes. The free tier allows 30 requests per minute, which is the single most
likely thing to break a demo; the harness retries 429s with jittered backoff and honours
Retry-After, and a fallback provider takes over when retries are exhausted.
Neon - 0.5 GB Postgres, permanent. Render's own free Postgres deletes itself after 30 days, so it is not used for the registry.
Redis Cloud - 30 MB including the RediSearch module, which the semantic cache needs for
FT.CREATE. Upstash Redis does not support it; set FASTRAG_SEMANTIC_CACHE_ENABLED=false
there and exact caching still works. The code also degrades to exact-only automatically if
the module turns out to be missing at runtime, rather than failing the request.
Langfuse Cloud - 50K units/month. Tracing is fail-open throughout, so an outage or a missing key costs observability and nothing else.
Fingerprinting hosted models
Local artifacts are pinned by SHA256 of the ONNX file. Hosted models have no local file to
checksum, so the fingerprint component becomes provider:model and the revision becomes
hosted-api (Settings.active_dense_artifact). This is a weaker guarantee, and honestly so:
a provider can change a model behind a stable name without telling you. The mitigation is the
golden gate - a silent model change shows up as a quality regression rather than passing
unnoticed.
verify_configured_models skips checksum verification when the active provider is hosted,
so the cloud profile starts without local model files present.
The fallback chain is never silent
When the primary generator exhausts its retries, FallbackGenerator switches to the
secondary and records it: the generator_provider field in the response says which provider
actually answered, a FALLBACKS counter increments, and the switch is traced. This preserves
the rule from llm-providers.md that a degraded answer must be
distinguishable from a normal one.
Configure it with FASTRAG_LLM_FALLBACK_BASE_URL, _API_KEY, and _MODEL. Leave them unset
and there is no fallback; the request fails loudly instead.
Providers