AI, decoded

How do you prepare PDFs for a RAG system?

Extract the document’s text and structure, check that tables and reading order survive, and keep source locations before splitting it into retrievable chunks. Test the extracted representation against the original pages; a search pipeline cannot recover a qualification or table relationship that preparation discarded.

· Chain of Thought

Level 4: How AI systems are built · 4.1 Retrieval and RAG

RAG & RetrievalContext Management

Check what the model will actually receive

A readable PDF is not automatically useful retrieval data. A person can follow columns, headers and footnotes on a page; your pipeline needs a representation that preserves the relationships needed to answer questions. LlamaIndex’s parsing documentation describes this preparation for scanned pages, tables and complex layouts, producing text, Markdown or structured data.

In episode 61, Jerry Liu separates extracting documents into a searchable knowledge base from extracting fields into a business workflow. For RAG, his sequence puts document conversion before chunking, embedding and storage. Inspect that conversion before tuning search or blaming the answering model.

Preserve relationships, not just words

Illustrative example: a PDF lists shipping times in a table. The columns are standard and express; the rows are domestic and international. A footnote says that the times exclude customs delays. An extraction that retains all four numbers but drops the headers has preserved characters while losing their meaning.

For a sample of documents, compare the extracted content with the original page. Check that headings attach to the right paragraphs, tables retain row and column labels, and qualifications stay associated with the statements they limit. Include scanned pages, multi-column layouts and tables that continue across pages in the sample.

Keep the original document identifier, version and page location with extracted content. Where the parser provides more precise locations, retain them. Liu highlights the value of tracing an answer back to the exact words on a page. That lets a reviewer investigate whether the error came from extraction or later interpretation.

Chunk around questions the material must answer

Choose chunk boundaries after checking the representation. For the shipping table, keep enough labels and qualifications together to interpret each value. If a large table must be divided, repeat the necessary header context or preserve a link to it. Do not choose a chunk size solely because it is convenient for storage.

Create a few questions with reviewed supporting passages. Ask for an ordinary value, an exception and a comparison across rows. Then inspect what retrieval returns. Does the candidate passage contain enough information to answer without guessing which column a number belongs to?

Record failures by stage. If the footnote disappeared during extraction, fix preparation. If it exists in storage but never reaches the answering model, inspect chunking and retrieval. If the model receives it and ignores it, inspect generation. RAG evaluation develops that separation.

Where it falls short

Good parsing does not establish that a document is authoritative, current or available to the requesting user. Carry those checks through ingestion and retrieval. Keeping a page reference also does not prove that the cited page supports the answer.

Some documents will remain ambiguous or poorly scanned. Flag them for review instead of silently treating a guessed cell value as verified text. Start with a small, representative collection and make the extracted result reviewable before scaling ingestion. The objective is a dependable source representation, not merely a successful file upload.

Hear it from the guest

“Convert it into Markdown, and then downstream operations become like chunking, embedding, putting into some storage system so you can actually search over it.”
“If you can trace back to the exact words and draw a box around the words that a legal contract came from when an AI agent generates a response, that's a very powerful tool”

Quotes lightly edited to remove filler words.

Go deeper

From the conversation

This explainer is drawn from these episodes — each carries its full transcript.

Concepts in this explainer

Retrieval-Augmented Generation (RAG)