Inspect indexed pages, folders, and discarded content
Review what a website source kept, how it is organized, and why content was discarded.
1 min read
- Project
- Knowledge
- Select a source
Select a ready source in the Knowledge library to inspect the retained pages, folders, and any discarded content.
Use the crawl result to identify missing material before changing prompts or retrying the source.
- Finish with
- Evidence of what an agent can and cannot retrieve
Sample the pages that matterCompare indexed count with intended scope, then read key previews.
Pages show URL, title, content preview, word count, and folder. A shortened or absent preview is not proof the full page is absent; representative-document fallbacks are marked truncated and are diagnostic, not exports.
Read the organization signalFolders show taxonomy, chunks, and approximate words.
/_unclassifiedmeans no matching taxonomy folder—not automatically taxonomy-optimized content. Inspect it before calling a result a retrieval failure.Explain discarded material before retryingDiscarded content is not retrievable.
Review
/_discarded_content, totals, discard rate, and reasons. A high discard rate usually calls for narrower scope, less duplicate/navigation-heavy content, or a cleaner authoritative source. Re-sync only after correcting the responsible source or scope.
Last updated on