Offers three modes (dense, sparse, hybrid), reranking, and metadata filtering, plus retrieval quality observability. The model is natively multilingual (including Polish) and handles mixed languages in a single index — a question asked in Polish finds the answer regardless of the source language. A single query (embedding + retrieval) typically drops below ~50 ms on a typical index — you measure that figure in the retrieval observability dashboard, we don't promise it in isolation from your data. Self-hosted deployment gives you control over cost and privacy. This is the engine we use for Company GPT and internal knowledge search tools. Underneath runs the BGE-M3 model — technical specs (1024-dim, embeddings); it's the same local model that powers this site's Concierge RAG.
