Chunking
In retrieval systems, chunking divides source material into smaller pieces that can be indexed and returned as evidence. Chunk boundaries determine how much context survives when a passage is retrieved on its own.
Consider a returns policy with a general 30-day rule and an exception for custom orders. If the exception lands in a separate chunk, a search can find the general rule and miss the qualification. Keep the rule and its relevant exception together, or give retrieval a way to return both.
A chunk saying “the limit is 30 days” also needs the product or customer tier named by its heading. Preserve useful headings and source identifiers with the text. Anthropic’s Contextual Retrieval approach addresses this loss of context by adding a short explanation of each chunk’s place in the document before indexing it.
There is no universal best size. Small chunks can retrieve a precise passage but omit context; large ones can include the context but also unrelated text. Overlap can protect boundaries at the cost of duplicates. Test representative questions and inspect whether their supporting evidence survives the split, using retrieval evaluation.
Sources
- Anthropic: Contextual Retrieval — Describes context lost by isolated chunks and the effects of size, boundaries and overlap.
Go deeper
- LangChain: Splitting recursively docs
Try character-based splitting with explicit chunk size, separators and overlap.