Every external dependency sits behind a protocol in src/fastrag/ports.py, so swapping a provider is a wiring change in src/fastrag/bootstrap.py rather than a rewrite. FASTRAG_PROFILE picks a coherent default set; individual FASTRAG_*_PROVIDER variables override any single choice.

The two profiles

localcloud
EmbeddingFastEmbed BAAI/bge-base-en-v1.5, in-process ONNXJina jina-embeddings-v3
RerankingFastEmbed Xenova/ms-marco-MiniLM-L-6-v2, in-processJina jina-reranker-v2-base-multilingual
Vector DBQdrant containerQdrant Cloud free
LLMOllama on the hostGroq free, OpenRouter fallback
CacheRedis containerRedis Cloud free
RegistryPostgres containerNeon free
TracingLangfuse Cloud, or self-hosted via the langfuse compose profileLangfuse Cloud
Speech-to-textSarvam (no STT model runs in-process)Sarvam

local is the benchmark rig: nothing crosses the internet on the retrieval path, which is what makes the sub-200ms measurement meaningful. cloud is the live demo that fits free tiers. Both are benchmarked and both sets of numbers are published - see latency.md for why conflating them would be dishonest.

The local profile's default embedder, bge-base-en-v1.5, is an English model, so querying in Hindi, Bengali, Tamil, Telugu or Marathi needs either Jina or the in-process multilingual pair below.

Multilingual without a hosted embedder

Both halves of retrieval can run in-process and still cover all six languages. This is how the self-corpus is built, and it means neither indexing nor serving spends Jina balance:

FASTRAG_EMBEDDING_PROVIDER=fastembed
FASTRAG_RERANKER_PROVIDER=fastembed
FASTRAG_DENSE_MODEL_ID=intfloat/multilingual-e5-large
FASTRAG_DENSE_MODEL_REPOSITORY=qdrant/multilingual-e5-large-onnx
FASTRAG_DENSE_MODEL_REVISION=ac6781cd1cf88b8306a536d7c9d18a5bd57cc14b
FASTRAG_DENSE_MODEL_FILE=model.onnx_data
FASTRAG_DENSE_MODEL_SHA256=0cf1883fee81c63819a44e2ba0efa51d4043d9759685a4ebebbde97e0623d15c
FASTRAG_DENSE_DIMENSION=1024
FASTRAG_DENSE_QUERY_PREFIX="query: "
FASTRAG_DENSE_DOCUMENT_PREFIX="passage: "
FASTRAG_RERANKER_MODEL_ID=jinaai/jina-reranker-v2-base-multilingual
FASTRAG_RERANKER_MODEL_REPOSITORY=jinaai/jina-reranker-v2-base-multilingual
FASTRAG_RERANKER_REVISION=9cfeff2df7d40d1b78e75e5e9cebec92a99813c9
FASTRAG_RERANKER_MODEL_FILE=onnx/model.onnx
FASTRAG_RERANKER_SHA256=0ef3f7978f7bc52360864d74edc1a0e03d159af770a7767c4d5943496e616012
FASTRAG_RETRIEVAL_FALLBACK_LANGUAGES=en

The reranker is the same jina-reranker-v2-base-multilingual the hosted API serves, run locally. The embedder is E5 rather than Jina's own jina-embeddings-v3, which FastEmbed can also run: embedding the self-corpus with it on CPU peaked at 11.3 GB and was killed by the kernel on a 15 GB machine.

Measured on the self-corpus's 437 English sections, with eight questions about FastRAG asked in each language, E5 put the section that answers the question in the top 20 for 46 of 48 queries and at rank 1 for at least six of eight in every language. The two misses were Tamil and Marathi phrasings of one guardrail question. Through the full path - language filter with the English fallback, then the local reranker against the calibrated abstention threshold - the same questions were answered 7/8 in English, 8/8 in Hindi, 7/8 in Bengali, Telugu and Marathi, and 4/8 in Tamil, where the reranker is least confident; the answering section was in the reranked top five for 45 of 48.

Details that matter:

  • E5 is asymmetric. It is trained with query: and passage: markers, and FastEmbed adds neither, so both prefixes are configured. Prefixes are normalised to end in exactly one space, because they are part of the embedding fingerprint and hosting dashboards trim env values - query: arriving without its space would otherwise reject a valid index.
  • The corpus stays English. A cross-lingual embedder matches a Hindi query to English chunks directly, but the language filter is an exact match on the chunk payload, and no chunk says hi. FASTRAG_RETRIEVAL_FALLBACK_LANGUAGES=en makes a language=hi query search hi and en chunks. Leave it empty for a corpus translated into every language, such as MSMARCO-XI, where exact filtering is the point.
  • Checksum the weights, not the graph. E5's ONNX export keeps 2.2 GB of weights in model.onnx_data beside a 0.5 MB model.onnx. Production verifies the file named by FASTRAG_DENSE_MODEL_FILE, so it names the weights.
  • Switch off the centroid off-topic gate. E5 embeddings are anisotropic: almost any text scores about 0.8 cosine against the corpus centroid. Measured on the self-corpus, questions about FastRAG scored 0.787-0.831 and questions such as "what is the capital of France?" up to 0.828, so no threshold separates them; the calibrated 0.826 refused three of eight English and five to seven of eight Indic questions. The reranker does separate them - 18 off-topic questions across the six languages scored at most -1.85 against an abstention threshold of 0.26 - so set FASTRAG_GUARDRAIL_OFFTOPIC_ENABLED=false and let abstention return no_answer for off-topic input. The injection and language guardrails are unaffected.
  • It costs memory, and CPU reranking is slow. Both models together peak at 3.7 GB resident; embedding a query takes 50-90 ms. Reranking 20 real self-corpus chunks - pairs of up to 658 tokens, because code tokenises long - took 8.5 s on 16 CPU cores at FastEmbed's default batch of 64 and 5.4 s with FASTRAG_RERANKER_BATCH_SIZE=1, since a batch is padded to its longest pair. On an RTX 3050 Ti it took 0.47 s, with scores within 0.0011 of the CPU run and an identical ranking. That rules out Vercel functions and Render's free instance; see deployment.md.
  • Offline work belongs on a GPU if you have one. Ingest, derive-eval.py and calibration rerank or embed hundreds of texts. With onnxruntime-gpu installed in place of onnxruntime, set FASTRAG_DENSE_EXECUTION_PROVIDERS or FASTRAG_RERANKER_EXECUTION_PROVIDERS to CUDAExecutionProvider, one model per card on 4 GB: E5's weights take 2.2 GB, and FASTRAG_DENSE_BATCH_SIZE=4 / FASTRAG_RERANKER_BATCH_SIZE=4 keep attention buffers inside what is left. Deriving the evaluation sets took 71 minutes on 16 CPU cores and 16 minutes with the reranker on an RTX 3050 Ti.
  • Licence. jina-reranker-v2-base-multilingual is published under CC BY-NC 4.0. Running the weights yourself is non-commercial use only; the hosted Jina API is the licensed route for commercial deployments. E5 is MIT.

Free tiers, and what each one costs you

Qdrant Cloud - 1 GB, roughly 250K vectors at 768d, permanent. Keeps named sparse vectors, so hybrid retrieval and RRF are unchanged from self-hosted. Only the URL and API key differ.

Jina AI - 10M tokens shared across embedding and reranking on one key. jina-embeddings-v3 is multilingual at 1024d and covers all five Indic languages. Changing embedding provider changes the embedding fingerprint and therefore requires a re-index; this is enforced at startup rather than discovered later through bad results.

Sarvam Saaras v3 - authenticates with an api-subscription-key header, not a bearer token, which is the usual first thing to get wrong. See voice.md.

Groq - OpenAI-compatible, so adapters/generation.py needs no changes. The free tier allows 30 requests per minute, which is the single most likely thing to break a demo; the harness retries 429s with jittered backoff and honours Retry-After, and a fallback provider takes over when retries are exhausted.

Neon - 0.5 GB Postgres, permanent. Render's own free Postgres deletes itself after 30 days, so it is not used for the registry.

Redis Cloud - 30 MB including the RediSearch module, which the semantic cache needs for FT.CREATE. Upstash Redis does not support it; set FASTRAG_SEMANTIC_CACHE_ENABLED=false there and exact caching still works. The code also degrades to exact-only automatically if the module turns out to be missing at runtime, rather than failing the request.

Langfuse Cloud - 50K units/month. Tracing is fail-open throughout, so an outage or a missing key costs observability and nothing else.

Fingerprinting hosted models

Local artifacts are pinned by SHA256 of the ONNX file. Hosted models have no local file to checksum, so the fingerprint component becomes provider:model and the revision becomes hosted-api (Settings.active_dense_artifact). This is a weaker guarantee, and honestly so: a provider can change a model behind a stable name without telling you. The mitigation is the golden gate - a silent model change shows up as a quality regression rather than passing unnoticed.

verify_configured_models skips checksum verification when the active provider is hosted, so the cloud profile starts without local model files present.

The fallback chain is never silent

When the primary generator exhausts its retries, FallbackGenerator switches to the secondary and records it: the generator_provider field in the response says which provider actually answered, a FALLBACKS counter increments, and the switch is traced. This preserves the rule from llm-providers.md that a degraded answer must be distinguishable from a normal one.

Configure it with FASTRAG_LLM_FALLBACK_BASE_URL, _API_KEY, and _MODEL. Leave them unset and there is no fallback; the request fails loudly instead.

Providers