Technical topic · RAG 02 · ChunkingWhere should the document be cut?
What the clerk can see
query: 饮料洒了一半,这杯能退款吗? (Q02-zh) (Half of the drink spilled — can this cup be refunded?)
evidence_sections: refund.spillage.severe
| strategy | Index | complete | retrieved tokens | signal |
|---|---|---|---|---|
whole-document | 7 pieces (one file = one piece) | true | 3309 | 3% |
structure-aware | 60 pieces | true | 283 | 41% |
The top-3 of whole-document is 3 whole files: refund-policy.md (1308 tokens), store-operations.md (1025 tokens), allergens.md (976 tokens). Both sides got the evidence; the only difference is how much text came back for this one question.
Averaged over all 11 questions: whole-document complete 11/11 · mean 2991 tokens · signal 8%; structure-aware complete 6/11 · mean 396 tokens · signal 42%.
Two limits: the index holds only 7 units, so top-3 is 43% of the whole knowledge base; and all 7 files are longer than the model’s 512-token window, so each vector represents only the start of its file (39%–64%). Putting the whole file in as one piece pushes “how big a piece” to its coarsest end — which is exactly what the lever in Scenario 2 is about.
Open a concept for Chinese, Japanese and English terms and an explanation.
Splitting documents into retrieval units. Separating subjects or conditions can leave a fragment insufficient for an answer.
Neighbouring chunks deliberately share a span of text, so a sentence near a cut can be found in either one. It adds no knowledge; it stores the same text more than once, and larger overlaps put more near-duplicates within reach of the top-k.
Boundaries come from the document’s own structure — headings, paragraphs, a question with its answer, a table with its note — rather than a fixed token count. It does not guarantee better retrieval: a vocabulary mismatch is not fixed by a tidier split.
Finding candidate material relevant to a question. Evidence sufficiency still needs to be checked.
Finding material by semantic relevance. Similarity does not imply factual equivalence. Level 8 uses hand-authored result lists.
Retrieve relevant material, then provide it with the question to a model for generation. Retrieval does not guarantee a correct answer.