Conversational search engine over a regulatory corpus
Conversational search over a normative corpus, with the exact reference on every answer.
The challenge
A normative corpus of several thousand pages, organised into chapters, subsections and annexes, in which an approximate answer has no value: the exact provision and its reference are required.
Our response
We ruled out naive document chunking, which breaks the regulatory hierarchy and produces unusable citations. The corpus is segmented while preserving the full lineage of each passage, from chapter to paragraph, and every fragment carries its attachment metadata. Search combines vector similarity and lexical search, with reranking of passages before generation. Every answer cites its sources, and the sources listed are strictly those actually used, filtered after generation.
Key points
Segmentation preserving chapter, section, article and paragraph lineage
Hybrid vector and lexical search, with reranking
Systematic citation of exact references
Post-generation filtering of sources, so an unused source is never shown
Tree navigation alongside conversational search
Technical stack
- Hierarchical parent-child chunking with metadata
- Multilingual embeddings + pgvector
- Hybrid BM25 + vector search
- Cross-encoder reranking
- Next.js 15 / TypeScript
- Open-source LLM hosted in Switzerland
Sector
Regulation & complianceLet's talk about your project
Tell us about your business challenge and we'll explore together how a tailored AI solution could support it.
