Kelu turns every PDF into markdown the same way, whichever source it comes from. A PDF has no real headings or paragraphs, so Kelu works them out from the layout. The better it can do that, the better the PDF answers questions.

Where PDFs come from

SourceHow a PDF gets in
File UploadYou upload it (up to 32 MB per file).
WebsiteThe crawl links to it and .pdf is in Also index linked files.
S3 BucketIt is an object in the bucket.
Google DriveIt is in a folder or file you picked.
OneDriveIt is in the indexed drive or folder.

What is extracted

  • The full text, in reading order.
  • Headings, worked out from font size. Text clearly larger than the body becomes a heading. This matters most: without headings, every part of the file loses its context.
  • Paragraph breaks, tables (three or more rows that line up), code and other monospace blocks, and nested lists.

What is left out

  • Headers and footers that repeat on most pages, and page numbers.
  • A table of contents near the front of the file.

What is not supported

  • Scanned pages. There is no OCR. Run OCR first and upload the result.
  • Images inside the PDF. Diagrams and screenshots are not indexed.
  • Files over 7,000 pages. Nothing is read. Split the file.
  • Link addresses. The link text is indexed, the address is not.
  • Page-level citations. A citation opens the file, not the page.

Extraction flags

A PDF that extracted badly may still be indexed, with a warning. For uploads and crawled PDFs, the warning shows next to the file on the source’s Files or Sync results tab, and the file is counted under Need a look. For S3, Google Drive and OneDrive, a PDF that yields no text is listed as skipped on the source’s Sync history tab, with the reason.
Shown asCodeWhat to do
No headings foundpdf_no_headingsExport again from the original with real heading styles, or split the file.
Almost no textpdf_low_text_yieldIt is almost always a scan. Run OCR and upload the result.
Multi-column layoutpdf_multi_columnColumns may be read out of order. Export as a single column if answers look garbled.
Could not be readpdf_unreadableThe file is damaged. Save it again from the original app. The rest of the upload still indexes.
Too long to extractpdf_too_many_pagesOver 7,000 pages, so nothing was read. Split the file.
Table of contents removedpdf_toc_droppedUsually nothing. Check it if the listing was the content, such as a catalogue.
After the first sync of a large PDF, check the flags. A flagged file and a clean one look the same in the document count.