GENAI 102
M06 · L05
Grounding, In Practice

Documents to LLM‑Ready Inputs

Real data is messy — PDFs, web pages, images, scanned paper. Turning all of that into clean text a model can actually use is half the work of grounding.

01 / 08
GENAI 102
M06 · L05
Four Steps

The Prep Pipeline

Every document goes through the same path before it can ground an answer:

  • Extract — pull the text out of PDFs, HTML, and images
  • Clean — strip junk and boilerplate, fix broken encoding
  • Chunk — split it into retrievable pieces
  • Tag — keep the source and metadata with each piece
02 / 08
GENAI 102
M06 · L05
The Judgement Call

Why Chunking Matters

Chunk size is a balance. Too big wastes the context window and blurs what is relevant; too small loses the meaning around each piece.

Rule of thumb
Split on natural boundaries — paragraphs, sections, headings — and keep enough context in each chunk to stand on its own.
03 / 08
GENAI 102
M06 · L05
Metadata Is Not Optional

Keep the Source

Attach to each chunk exactly where it came from — the file, the page, the section. That is what lets an answer cite its evidence later.

  • Traceability — every claim can point back to a real passage
  • Honesty — this is what Module 8 leans on to show its work
  • Debugging — when an answer is wrong, you can find the chunk that caused it
04 / 08
GENAI 102
M06 · L05
Why the Effort Pays Off

Garbage In, Garbage Out

Bad extraction poisons everything downstream — scrambled PDF text, leftover navigation menus, and boilerplate all get retrieved as if they were real content.

The takeaway
Clean inputs are worth the effort. The model can only be as good as the text you hand it.
05 / 08
GENAI 102
Build It
From Concept to Capstone

Build It

How to implement: take one messy document, extract its text, and split it into a few sensible chunks — putting a source tag on each one.

  • Weekly AI Tasks tracker — incoming messages are short and clean, but the moment you ingest notes or attachments, they will need this pipeline.
  • Personal brand site — extract clean facts from a résumé and repos, and tag each with its source so the site can cite them.
06 / 08
GENAI 102
Knowledge Check

Check what stuck

Three questions from this lesson. Answer to see why — the explanation appears whether you were right or wrong. Nothing is scored or saved.

Question 1 of 0
Score 0/0

07 / 08
GENAI 102
Summary
Recap

What you learned

Documents become LLM‑ready in four steps — extract, clean, chunk, tag. Chunk on natural boundaries, keep the source with every piece, and remember: garbage in, garbage out.

08 / 08