Skip to main content
L9.2

Documents, Chunks, and Metadata

Goal

Split documents into retrieval chunks while preserving source identity, boundaries, ordering, freshness, and access metadata.

Retrieval systems rarely compare a query against an entire book or handbook as one indivisible object. They usually split sources into chunks. A chunk is a retrieval-sized unit that still points back to its source.

Chunking changes what can be retrieved​

Suppose a document contains:

Paragraph 1: product dimensions
Paragraph 2: battery duration
Paragraph 3: warranty
Paragraph 4: return policy

If the entire document is one chunk, retrieval can return everything at once. If every sentence is a separate chunk, retrieval is more precise but related context may be separated. Chunk size trades off:

  • precision;
  • context completeness;
  • number of index items;
  • retrieval noise;
  • prompt token cost.

There is no universally correct chunk size.

Boundaries can destroy meaning​

Consider:

The exception applies only when...
[chunk boundary]
...the device was purchased before June.

A naive fixed-length split can separate a condition from the statement it qualifies. Possible strategies include:

  • paragraph-aware splitting;
  • heading-aware splitting;
  • sentence grouping;
  • overlapping windows.

Overlap can preserve boundary context, but it duplicates text and can cause near-duplicate search results.

Metadata is part of the retrieval object​

A useful chunk record might contain:

{
"chunk_id": "policy-7#03",
"source_id": "policy-7",
"title": "Returns Policy",
"section": "Exceptions",
"updated_at": "2026-09-20",
"access_group": "support",
"text": "..."
}

The text supports semantic/keyword matching. Metadata supports:

  • source attribution;
  • freshness filtering;
  • authorization;
  • deduplication;
  • ordering;
  • debugging.

Do not force every control into the embedding vector.

IDs must survive re-indexing​

If chunk IDs change every time the index is rebuilt, old citations and evaluation records become difficult to interpret. Use a stable chunk identity policy where practical. For example:

source_id + section_id + chunk_ordinal + content/version hash

The exact scheme varies, but it should let you answer:

Which source content produced this retrieved item?

Chunking belongs in evaluation​

A retrieval miss can be caused by:

  • a weak embedding;
  • bad similarity;
  • a query mismatch;
  • or a bad chunk boundary.

If the answer spans two chunks, top-1 retrieval may return only half the needed evidence. Include evaluation cases that test:

  • facts near boundaries;
  • multi-chunk answers;
  • short exact identifiers;
  • long sections with one relevant sentence.

Chunk identity needs lineage​

A chunk should remain traceable to its parent document and source version. For example:

document_id: policy-17
document_version: 2026-09-01
chunk_id: policy-17:v3:chunk-04
start_offset: 1200
end_offset: 1680

If the source changes and the index is rebuilt, old and new chunks should not silently share an identity unless they really represent the same source state. This matters for citations, deletion, freshness checks, and debugging. A citation to chunk-04 is weak if nobody can tell which document version produced that chunk.

Chunk size is a retrieval contract, not only preprocessing​

Large chunks preserve context but can mix several topics into one retrieval unit. Small chunks sharpen topic focus but can split a fact from the condition that gives it meaning. Overlap can reduce boundary loss, but it also creates near-duplicate candidates.

So record the chunker version and parameters with evaluation results. If chunking changes, the retrieval corpus effectively changed too.

Predict

Why preserve source_id on every chunk?

Build chunks in the Lab​

This Lab contains a small sectioned document.

This Lab contains one small sectioned document, policy-7, with three sections.

  1. Click Run once. Nothing is printed, because make_chunks returns an empty list, and the checks fail.
  2. Complete the TODO: return one chunk per section. Each chunk needs a stable chunk_id built from the source ID and a two-digit ordinal (policy-7#00, policy-7#01, ...), plus the source_id, section, updated_at, access_group, and text.
  3. Click Run again. You should see three chunk records, and every check should pass.
  4. Add a fourth section at the end of sections: ("exceptions-2", "Gift cards are final sale."),. Before running, predict its chunk_id.
  5. Click Run. The new chunk is policy-7#03, and the first three IDs did not move. Appending keeps old IDs stable; inserting a section in the middle would shift every later ordinal. The Lab should report Result: experiment ran because the original three-chunk baseline changed while the chunking and metadata invariants still pass.
  6. Notice a boundary problem: the rule about safety equipment lives in exceptions, far from the 30-day rule in eligibility. Write one evaluation question—such as “Can I return an opened helmet after 10 days?”—that needs both chunks. That kind of case reveals a bad chunk boundary.

Loading lab…

Quick Check

1. What can happen when chunks are too small?
2. What is a cost of overlapping chunks?
3. Which field is most appropriate for an access-control decision?

0 of 3 questions answered.

Explain it back​

Design a chunk record for a policy document. Include at least five metadata fields and explain which are used for citation, freshness, and access control.

Key Takeaways

  • Chunking determines the retrieval unit.
  • Boundary choices can preserve or destroy useful context.
  • Overlap trades context preservation for duplication.
  • Metadata carries provenance, freshness, and access information.
  • Chunking strategy should be evaluated, not assumed.

Next Lesson

Next, represent queries and chunks as vectors for semantic retrieval.

References

Lesson actions

Completion is stored locally on this device.

View progress