Can this work with several million documents?
Yes, if ingestion, indexes, ACL filters and reranking are designed for the corpus and measured under representative load. The exact architecture depends on document distribution, update rate and isolation requirements.
Do documents have to leave our environment?
No. Controlled EU-hosted or on-premise/private execution patterns are possible where policy or client requirements demand them.
How do you preserve legal provenance?
Keep immutable source identifiers and hashes, version derived artifacts, and make every citation resolve back to a stable source location such as document plus page/section.
Is this legal advice?
No. This is technical platform and engineering work for document retrieval, review support, evidence handling and controlled AI execution.
How do you compare lexical and semantic retrieval?
Evaluate them independently and as a fused system against known-relevant sets. Exact identifiers and citations often favor lexical retrieval; conceptual queries can benefit from semantic candidates and reranking.
How is cost controlled?
Measure each stage separately: OCR, parsing, embedding, indexing, reranking and generation. Cache/reuse derived artifacts and reserve expensive model calls for stages where evaluation shows measurable value.