Set up
Start the source
Open the knowledge base, click Add source, pick Website and give it a Display name.
Preview the crawl
Enter the Start URL and click Crawl. Its path limits the crawl:
https://docs.example.com/docs never picks up /blog. Kelu finds pages from the sitemap (/sitemap.xml and any listed in robots.txt) and from links. Narrow the list with Include patterns and Exclude patterns. More options are under Advanced configuration.Preview the content
Click Next: extraction, then Convert to see the markdown for a few pages. If the main content is missing, add Content selectors. To drop a table of contents or a feedback box, add Exclude selectors.
What gets indexed
- Each page’s title (from
<title>, or the first<h1>) and its main content as markdown: headings, lists, tables, code blocks and image alt text. - Navigation, headers, footers and sidebars are removed.
- Pages marked
noindex, pages on other domains, and exact duplicates are skipped. - Linked files (such as PDFs) only when you list their extension in Also index linked files. See PDF extraction.
Settings
Everything below is in the add form (most under Advanced configuration) and on the source’s Configuration tab. Changes apply on the next sync. Click Sync Now to apply them at once.Which pages are crawled
| Setting | Default | What it does |
|---|---|---|
| Start URL | required | Where the crawl starts. Its path is the crawl’s scope. |
| More start URLs | none | Up to 20 more, one per line. Each is a start point and a scope, so /docs plus /guides crawls both. |
| Include patterns | everything in scope | One per line. Only matching URLs are crawled. |
| Exclude patterns | none | One per line. Wins over include patterns. |
| Treat patterns as regular expressions | off | Off: a pattern matches anywhere in the URL, with * and ? as wildcards. |
| Include patterns replace the scope | off | On: only the include patterns decide, and the start URL paths stop limiting the crawl. |
| Crawl but do not index | none | Pages fetched only for their links. Use it for index pages that list content but hold none. |
| Exclude extensions | none | For example .csv, .zip. Skipped without being fetched. |
| Sitemap URLs | found automatically | Name sitemaps Kelu can’t find. Naming any replaces discovery, and the sync fails if one can’t be read. |
| Skip the sitemap and follow links only | off | Ignores sitemaps and finds pages by following links. |
| Ignore URL parameters | off | Treats ?utm_source= and other query variants as one page. Turning it on later re-indexes pages under new addresses. |
| Max pages | your plan’s page limit | Most pages one crawl will fetch. |
| Delay between requests (ms) | 100 | Raise it if the site rate-limits the crawler. |
| Respect robots.txt | on | Skip paths robots.txt disallows. |
| User agent | Kelu’s default | What the crawler calls itself, to pages and to robots.txt. |
| Render JavaScript | off | Load each page in a headless browser first. Much slower, and every page is re-read on every crawl. |
| Wait for elements | none | With Render JavaScript only. Capture the page once these CSS selectors appear. A selector that never appears fails the page. |
| Outbound proxy | none | http://user:pass@host:port. Crawl through your own address, for sites that block Kelu or allow only known IPs. |
| Also index linked files | none | Extensions to index besides HTML, for example .pdf, .docx. |
Which fetched pages are kept
| Setting | Default | What it does |
|---|---|---|
| Only index pages containing | none | CSS selectors, one per line. A page must contain all of them. Its links are still followed. |
| Skip pages containing | none | CSS selectors, one per line. A page containing any of them is dropped whole. |
| Minimum publish date | none | YYYY-MM-DD. Drops pages whose structured data says they were published earlier. Pages with no date are always kept. |
What is pulled out of each page
| Setting | Default | What it does |
|---|---|---|
| Content selectors | auto-detect | CSS selectors for the main content, tried in order. |
| Exclude selectors | none | CSS selectors to remove, on top of the built-in navigation and footer removal. |
| Exclude classes | none | Class names without the dot. Removed with everything inside. |
| Exclude wrapper classes | none | The wrapper is removed and its content kept, for example a tab panel. |
| Title prefix | none | Added to every title from this source, so two similar sources are easy to tell apart in citations. |
Sync
A website re-crawls every 24 hours, counted from the end of the last crawl. Pages the server reports as unchanged are not downloaded again. Sync Now re-crawls at once. See Keeping content fresh. A crawl that stops at Max pages with pages left is marked failed (page_limit_reached) and removes nothing until you raise the limit or narrow the crawl. See When a sync changes a lot.
On the Sync results tab you can also work page by page:
- Add page indexes URLs the crawl doesn’t reach, such as orphan pages, and keeps them in every later crawl. Up to 25 at a time and 200 per source.
- Re-crawl this page now fetches one page again.
- Remove this page and its document removes a page you added.
/ covers everything beneath it.
Troubleshooting
- The preview is empty, or shows only
<div id="root"></div>. The site renders in the browser. Turn on Render JavaScript. - Only the start URL is found. Links are built by JavaScript. Turn on Render JavaScript, or name a sitemap in Sitemap URLs.
- Pages are missing. Check the reason on Sync results. “Outside the start URL’s section”, “Blocked by robots.txt”, “Matched an exclude pattern” and “Matched no include pattern” each point to the setting to change.
- Linked PDFs skipped as “Not an HTML page”. Add
.pdfto Also index linked files. - Every page fails with 403. The site blocks the crawler. Set an Outbound proxy.
- “cannot be crawled” when saving. Kelu does not crawl some large consumer sites (search engines, social networks, webmail). Their developer docs, such as
learn.microsoft.comordevelopers.google.com, are allowed.