←RAG seriesRAG 02 · Scenario 1 / 4

Technical topic · RAG 02 · ChunkingWhere should the document be cut?

The question at the counter—
Real experiment data
EN

Choose a scenario

Switching starts that scenario from its beginning. What you have finished is kept in this browser.

What the clerk can see

Concepts on this page

Open a concept for Chinese, Japanese and English terms and an explanation.

ChunkingView details
中文
文本切分
日本語
チャンク分割
English
Chunking

Splitting documents into retrieval units. Separating subjects or conditions can leave a fragment insufficient for an answer.

Chunk overlapView details
中文
切块重叠
日本語
チャンク重複
English
Chunk overlap

Neighbouring chunks deliberately share a span of text, so a sentence near a cut can be found in either one. It adds no knowledge; it stores the same text more than once, and larger overlaps put more near-duplicates within reach of the top-k.

Structure-aware chunkingView details
中文
按结构切分
日本語
構造に基づくチャンク分割
English
Structure-aware chunking

Boundaries come from the document’s own structure — headings, paragraphs, a question with its answer, a table with its note — rather than a fixed token count. It does not guarantee better retrieval: a vocabulary mismatch is not fixed by a tidier split.

RetrievalView details
中文
检索
日本語
検索
English
Retrieval

Finding candidate material relevant to a question. Evidence sufficiency still needs to be checked.

Semantic searchView details
中文
语义检索
日本語
意味検索
English
Semantic search

Finding material by semantic relevance. Similarity does not imply factual equivalence. Level 8 uses hand-authored result lists.

Retrieval-augmented generation RAGView details
中文
检索增强生成
日本語
検索拡張生成
English
Retrieval-augmented generation (RAG)

Retrieve relevant material, then provide it with the question to a model for generation. Retrieval does not guarantee a correct answer.