A website source crawls a docs site, turns each page into clean markdown, and re-crawls it every 24 hours. It works out of the box for static sites such as Docusaurus, Mintlify, GitBook, VitePress, MkDocs and Sphinx. For sites that build their pages in the browser (React, Vue, Angular), turn on Render JavaScript.

Set up

1

Start the source

Open the knowledge base, click Add source, pick Website and give it a Display name.
2

Preview the crawl

Enter the Start URL and click Crawl. Its path limits the crawl: https://docs.example.com/docs never picks up /blog. Kelu finds pages from the sitemap (/sitemap.xml and any listed in robots.txt) and from links. Narrow the list with Include patterns and Exclude patterns. More options are under Advanced configuration.
3

Preview the content

Click Next: extraction, then Convert to see the markdown for a few pages. If the main content is missing, add Content selectors. To drop a table of contents or a feedback box, add Exclude selectors.
4

Add it

Click through to Add source. The first crawl starts right away. When it finishes, open the source’s Sync results tab to see every URL as Indexed, Need a look, Skipped or Failed, with the reason.

What gets indexed

  • Each page’s title (from <title>, or the first <h1>) and its main content as markdown: headings, lists, tables, code blocks and image alt text.
  • Navigation, headers, footers and sidebars are removed.
  • Pages marked noindex, pages on other domains, and exact duplicates are skipped.
  • Linked files (such as PDFs) only when you list their extension in Also index linked files. See PDF extraction.

Settings

Everything below is in the add form (most under Advanced configuration) and on the source’s Configuration tab. Changes apply on the next sync. Click Sync Now to apply them at once.

Which pages are crawled

SettingDefaultWhat it does
Start URLrequiredWhere the crawl starts. Its path is the crawl’s scope.
More start URLsnoneUp to 20 more, one per line. Each is a start point and a scope, so /docs plus /guides crawls both.
Include patternseverything in scopeOne per line. Only matching URLs are crawled.
Exclude patternsnoneOne per line. Wins over include patterns.
Treat patterns as regular expressionsoffOff: a pattern matches anywhere in the URL, with * and ? as wildcards.
Include patterns replace the scopeoffOn: only the include patterns decide, and the start URL paths stop limiting the crawl.
Crawl but do not indexnonePages fetched only for their links. Use it for index pages that list content but hold none.
Exclude extensionsnoneFor example .csv, .zip. Skipped without being fetched.
Sitemap URLsfound automaticallyName sitemaps Kelu can’t find. Naming any replaces discovery, and the sync fails if one can’t be read.
Skip the sitemap and follow links onlyoffIgnores sitemaps and finds pages by following links.
Ignore URL parametersoffTreats ?utm_source= and other query variants as one page. Turning it on later re-indexes pages under new addresses.
Max pagesyour plan’s page limitMost pages one crawl will fetch.
Delay between requests (ms)100Raise it if the site rate-limits the crawler.
Respect robots.txtonSkip paths robots.txt disallows.
User agentKelu’s defaultWhat the crawler calls itself, to pages and to robots.txt.
Render JavaScriptoffLoad each page in a headless browser first. Much slower, and every page is re-read on every crawl.
Wait for elementsnoneWith Render JavaScript only. Capture the page once these CSS selectors appear. A selector that never appears fails the page.
Outbound proxynonehttp://user:pass@host:port. Crawl through your own address, for sites that block Kelu or allow only known IPs.
Also index linked filesnoneExtensions to index besides HTML, for example .pdf, .docx.

Which fetched pages are kept

SettingDefaultWhat it does
Only index pages containingnoneCSS selectors, one per line. A page must contain all of them. Its links are still followed.
Skip pages containingnoneCSS selectors, one per line. A page containing any of them is dropped whole.
Minimum publish datenoneYYYY-MM-DD. Drops pages whose structured data says they were published earlier. Pages with no date are always kept.

What is pulled out of each page

SettingDefaultWhat it does
Content selectorsauto-detectCSS selectors for the main content, tried in order.
Exclude selectorsnoneCSS selectors to remove, on top of the built-in navigation and footer removal.
Exclude classesnoneClass names without the dot. Removed with everything inside.
Exclude wrapper classesnoneThe wrapper is removed and its content kept, for example a tab panel.
Title prefixnoneAdded to every title from this source, so two similar sources are easy to tell apart in citations.

Sync

A website re-crawls every 24 hours, counted from the end of the last crawl. Pages the server reports as unchanged are not downloaded again. Sync Now re-crawls at once. See Keeping content fresh. A crawl that stops at Max pages with pages left is marked failed (page_limit_reached) and removes nothing until you raise the limit or narrow the crawl. See When a sync changes a lot. On the Sync results tab you can also work page by page:
  • Add page indexes URLs the crawl doesn’t reach, such as orphan pages, and keeps them in every later crawl. Up to 25 at a time and 200 per source.
  • Re-crawl this page now fetches one page again.
  • Remove this page and its document removes a page you added.
To keep a crawled page out, add its address on the Excluded tab. An address ending in / covers everything beneath it.

Troubleshooting

  • The preview is empty, or shows only <div id="root"></div>. The site renders in the browser. Turn on Render JavaScript.
  • Only the start URL is found. Links are built by JavaScript. Turn on Render JavaScript, or name a sitemap in Sitemap URLs.
  • Pages are missing. Check the reason on Sync results. “Outside the start URL’s section”, “Blocked by robots.txt”, “Matched an exclude pattern” and “Matched no include pattern” each point to the setting to change.
  • Linked PDFs skipped as “Not an HTML page”. Add .pdf to Also index linked files.
  • Every page fails with 403. The site blocks the crawler. Set an Outbound proxy.
  • “cannot be crawled” when saving. Kelu does not crawl some large consumer sites (search engines, social networks, webmail). Their developer docs, such as learn.microsoft.com or developers.google.com, are allowed.