What I learned building production RAG systems
Lessons from Predli Studio's enterprise RAG: most of the wins came from ingestion, chunking, and metadata, not from the model.
I spent a good part of the last two years building the ingestion and retrieval behind Predli Studio, a multi-tenant enterprise RAG product. Four of the decisions we made have held up in production, and all four are about what happens to a document before the model sees it.
Chunking mattered more than the model
Early on we ran the obvious experiment, which was to swap the generation model and measure answer quality. The lift was marginal. Later we rewrote chunking so that it respected document structure (section headers, paragraph boundaries, table boundaries), and answer quality improved more than it had with any model swap. The embeddings and the generation model were unchanged.
On the Enterprise RAG platform we measured it directly. File-type-aware chunking with TOC-based hierarchical context gave about 40% better retrieval relevance than a uniform fixed-window baseline. A fixed 512-token window will sometimes cut a key paragraph in half. When it does, the retriever returns two fragments that each make only partial sense, and the model fills the gap with something plausible and wrong.
Each file type needs its own parsing and chunking
Our first mistake was treating every document the same way: extract the text, chunk it, embed it. That works for prose. It does not work for PDFs with tables, slide decks, or spreadsheets.
PDFs carry layout that naive extraction flattens, so tables come out as word salad. In a slide deck the hierarchy matters: the title on slide 12 governs the bullets beneath it but has nothing to do with the footer. A spreadsheet is already a table, and turning it into paragraphs throws away the rows and headers that made it possible to answer questions from it. We ended up with a parser per file type, each feeding a chunking strategy tuned to that format: layout-aware extraction for PDFs, slide-level chunking with title metadata for decks, and structured representations with preserved headers for spreadsheets. Each of these was a small project of its own, and each was worth the time.
Retrieval needs metadata as well as vectors
Semantic search on its own is only a starting point. In a real product you almost always need to filter or boost results on something other than similarity.
An analyst who asks for the latest revenue guidance wants the most recent disclosure from one specific company. A semantically similar passage from a 2019 earnings call is a wrong answer, however close the embedding is. Date, source, document type, and tenant decide whether a result is useful at all, so we could not treat them as optional metadata. We ran the semantic part of a query against the vector store, applied the structured constraints as filters, and merged the two. On Studio this also meant persona-aware retrieval, where the same corpus surfaces differently depending on who is asking.
Idempotency at the boundary
The decision that has held up best is also the least visible one. Studio runs ingestion on Dagster with asset-based orchestration instead of a task scheduler, and we hash content at the raw-document boundary. If a connector retries or re-sends the same file, it cannot write a duplicate into the vector store. This is plumbing, but without it a knowledge base fills up with near-duplicates that nobody can account for, and people stop trusting what it returns.
Where I would spend the effort
If most of your RAG effort is going into model selection, I would move it to parsing, chunking, metadata, and the ingestion guarantees that keep the index clean. For us, the model was the cheapest thing to change and the change that moved answer quality least.