Index the files in an S3-compatible bucket — handy for content Kelu has no connector for. Kelu only reads; it never writes to the bucket.

Before you start

  • A Pro or Enterprise plan.
  • The Edit Sources permission on the knowledge base. Owners and admins always have it.
  • An access key that can list the bucket and read its objects (s3:ListBucket and s3:GetObject), and nothing else.

Set up

1

Add the source

Open the knowledge base, go to Sources → Add source, pick S3 Bucket and click Continue. Enter a Display name.
2

Enter the bucket details

Fill in Bucket, Access key ID and Secret access key. Add the Endpoint if the store is not AWS, and a Prefix to index only part of the bucket.
3

Test and add

Click Test connection. It tells you if the key is refused, the prefix matches nothing, or no file can be read. Then click Add source. The first sync starts right away.

Settings

SettingRequiredDescription
BucketYesThe bucket name.
Access key IDYesA read-only key is enough.
Secret access keyYesNever shown again after you save it.
RegionNoDetected from the bucket. Set it for a store that has no regions.
EndpointNoFor MinIO, R2, Wasabi and Backblaze. Leave empty for AWS.
PrefixNoOnly keys that start with this are indexed. End it with / to mean a folder: handbook also matches handbook-archive/.
Session tokenNoOnly for temporary STS credentials. Set it under Configuration after you add the source.
Path-style addressingNoAlready on for a custom endpoint. Set it under Configuration for an AWS bucket that needs it.

What gets indexed

  • Files of these types: .pdf .docx .xlsx .xlsm .xltx .xltm .csv .tsv .html .htm .md .mdx .txt. Other types are skipped silently.
  • Each file becomes one document, titled with its file name (or an HTML file’s <title>).
  • Files over 32 MB, and files with no readable text, are skipped and listed with the reason under Sync history.
  • A sync lists at most 500,000 objects. A bigger bucket fails the sync — use a Prefix to narrow it.

Make citations clickable

By default a citation points to s3://bucket/key, which a reader cannot open. If the files are also on the web, add an index.json that maps each key to its public address:
index.json
{
  "handbook/expenses.pdf": "https://intranet.acme.com/handbook/expenses",
  "reports/2026-q1.pdf": "https://acme.com/investors/2026-q1"
}
A list works too:
index.json
[
  { "object_key": "handbook/expenses.pdf", "source_url": "https://intranet.acme.com/handbook/expenses" }
]
  • Put it at the bucket root or the prefix root. If both exist, the prefix one wins.
  • Write keys as full paths. Paths relative to the prefix also work.
  • Files not listed keep their s3:// address. If no entry matches any file, Sync history says so.

Sync

The bucket is checked every 10 minutes. Only files that changed are downloaded again. A file deleted from the bucket is removed at the next sync, unless a large share of the source would go at once — then Kelu holds the removals for review. Sync Now on the source page runs a sync at once.

Troubleshooting

  • “access denied” — the key cannot list the bucket, or cannot read its objects. It needs both.
  • “signature rejected” — the secret access key does not match the key ID. “no such access key” means the key ID is wrong.
  • “not found” — check the bucket name and the endpoint.
  • “no objects under prefix” — the prefix matches nothing. A folder needs its trailing /.
  • A file is missing — check its type, then look for it under Sync history.