In this lesson you’ll cover:
- Why ~99% of indexed content is eliminated before a synthesis answer is written — in two very different stages
- Four common wrong assumptions (rank→cited, retrieved→used, all content is a candidate, retrieval=citation optimization)
- The four-stage narrowing pipeline: early filtering, re-ranking, late evaluation, final response
- Early filtering mechanics — cosine similarity, entity presence, embedding sharpness, heading signal (no language understanding yet)
- How complex queries decompose into independent sub-retrievals, each with its own winning chunks
- The four early-filter failure patterns (topically diffuse page, entity-absent chunk, scope-free claim, heading mismatch) and fixes
- Late evaluation’s four tests: self-containment, claim attributability, cross-chunk consistency, additive value
- The full “discard map” — where content dies, from pre-crawl to generator discard
- A diagnostic that maps your citation signal to the failure point and first fix, with an A→D worked example
