Architecture, local setup, deployment, and pipeline deep-dives - rendered from the same markdown in the repo.
Compose stack, profiles, calibration, and first query.
Vercel, Render, env vars, and cloud providers.
Local vs cloud adapter matrix.
Request path, components, and data flow.
Strategies, metadata, and index shape.
Corrective retrieval grading and rewrite.
Input gates before retrieval and generation.
STT, streaming voice queries, and audio handling.
Stage budgets and the sub-200ms retrieval target.
Monitoring, tracing, and day-two tasks.
Incident checks and recovery steps.
Latency harness and golden gates.
OpenAI-compatible generation backends.