Search Indexes

A search index makes content searchable for your loops: a crawl of a website, files you upload, or the records of one of your collections. You create an index once, on the account, on the Indexes page. Each loop names the indexes its runs may search.

Create an index

On Indexes, press New index and choose what it is built from. That choice never changes.

Built fromWhat it holds
A site crawlThe pages of one website, fetched again daily or weekly
Files you uploadMarkdown, text and HTML files, and PDFs that have a text layer
A collection’s recordsThe fields you choose from one of your collections, kept up to date as runs write and delete records

Every index has a name for people, such as “Product documentation”, and an id that a loop uses to name it, such as docs. The id is a lower-case letter, then lower-case letters, digits or underscores.

A new index is untrusted: a run treats each result as untrusted content. An owner of the account can mark it trusted on its page, with Mark trusted…, once it exists; a member sees the mark and cannot change it. Mark an index trusted only for content you control. Trust also changes what a loop that searches the index may do; see What an index lets a loop do.

One choice applies to every kind:

  • Search by meaning as well as by words. On by default. With it on, the index finds a passage that says the same thing in other words, as well as the exact words.

A site crawl

Give the crawl an address, such as https://docs.example.com/guide/. The crawl reads that host under that path, and nothing outside it. It reads the site’s robots.txt and sitemap and follows the links it finds inside that scope. Choose how often it fetches the site again, daily or weekly, and at most how many pages it stores: 2,000 by default, and at most 10,000.

The first crawl starts when you create the index. A page that only shows its text after its scripts run, as in a single-page app, is opened in a browser so that its text is found. The crawl reaches public sites only, and sends no login.

A crawled page takes its title from the first heading of its text, so a search result shows the page’s own name and not the site’s name repeated. A page with no heading keeps the title the site gives it.

The index shows its last crawl in words: whether it is crawling now or finished, how many pages it fetched, stored, rendered in a browser and failed, and what it cost. A crawl that runs out of money stops, keeps the pages it fetched, and says so.

Fetch now crawls the site again at once, when you changed pages and do not want to wait for the schedule. It starts the same crawl the schedule runs, and costs the same. An account can press it 6 times a day (the day ends at 00:00 UTC); a press that starts nothing does not count. It starts nothing while a crawl of that index runs, and says when that crawl started. The schedule keeps its own crawls either way.

A crawl that finishes removes the pages it no longer reaches: a page now outside the address, a page the site skips or refuses to crawlers, and a second address that serves the same text as a page it kept, so a search never finds one page twice. A page that fails to load once keeps its place. A crawl that stops early, or that reaches its page limit, removes nothing.

Files you upload

Open the index and press Upload a file. The index takes Markdown, plain text and HTML files, and PDFs with a text layer. A scanned PDF has no text layer, so the index refuses it and says why. Each document shows its state: pending until it can be searched, then indexed. To take a file out of the index, press Remove beside it.

A collection’s records

Name the collection and the fields to index, such as company, notes. The index starts with the records already stored, then follows every record a run writes or deletes.

To show records that describe the same thing as one result, name the field that identifies it under Same item when this field matches, such as sku: two listings of one product then show as one search result. It is a text or number field the collection declares, and not one that holds what people typed. A record without a value in it is still its own result. You choose this field when you create the index and cannot change it later; to fold on another field, create a new index.

Records can hold personal data. With search by meaning on, their text is sent to the embedding provider.

Let a loop search an index

Name the index in the loop’s record, then publish the version:

indexes: [docs, pricing]

The loop’s page lists them under What its runs may search. A run searches only the indexes its version names. A published version always searches an index’s current content, so a new crawl or a new file reaches the loop without a new version.

Publish refuses a name that no index on your account has. You cannot delete an index while the newest published version of a loop uses it: the refusal names the loop.

A run in such a loop can search its indexes and read a document it found, so the loop can answer from your docs and say which page the answer came from. A run searches when the loop’s capabilities include search_indexes, and an agent it starts lists search in its tools and has a role that allows search_indexes — reader, librarian and web_reader do (Permissions and trust). Publish refuses a loop whose indexes no agent could search, and names the fix. While the loop works, a visitor reads “Searching the docs…” or “Reading the docs…”, never the name of a tool. Each search is paid from the wallet.

What an index lets a loop do

No agent in a run may hold all three of these at once: it reads untrusted content, it touches private data, and it acts outside the run, such as pressing a control on your page or sending a message. Searching an index counts for the first two, unless the index says otherwise:

  • A trusted index brings no untrusted content.
  • A site crawl is public: the crawl sends no login, so it stores only what anybody can read. It brings no private data. Uploaded files and a collection’s records are always private.

So an agent that searches only trusted crawls may also act. One untrusted or private index among a loop’s indexes counts for all of its searches.

Publish checks the indexes as they are, and the version keeps that reading. To let a loop act because you marked an index trusted, publish the loop again. Marking an index untrusted takes effect at once: the next run counts it, and refuses any grant that would bring an agent to all three. If the run’s agent holds all three from its start, the run’s record names the index and says to publish the loop again.

Search your indexes yourself

On Indexes, press Search every index above the list to search all your indexes at once, without opening one. An index’s own Try a search tab opens the same search, starting on that index. Type what you are looking for, then choose:

  • Find — the words, the exact phrase, or the words allowing typos.
  • In — every index, or one.
  • By — Both, Words only or Meaning only. Words find an exact name, such as a command-line flag. Meaning finds a passage that says the same thing in other words. Both ranks the two together, as a loop’s own searches do. A search by meaning embeds your query once, at the index’s embedding price, and your wallet pays it; Words only costs nothing.

Each result names its index and its title, and says when the index is untrusted. Press Open on a result to read the document from that place.

A document is one result, at its best-matching section. When other sections of the same page also matched, the result says so, for example +2 more sections; Open reads the whole page. A page crawled into two of your indexes is one result too, and so are records with the same value in their index’s Same item when this field matches field. A run’s agent gets the same results, and is told that the rest of each page can be read.

How results are ordered

Settings chooses how searches are ordered, separately for three kinds of search: A run’s agent, Your search on Indexes, and Your visitors on your site. Each one is set to one of:

  • Off — the plain order of the word and meaning matches.
  • Rerank — a decision model judges each match and puts the best first.
  • Fuse — the decision model’s judgement added to the plain order.

All three use Fuse unless you choose otherwise; Platform default shows it. Rerank and Fuse each cost a small amount for each search, paid from the wallet, and add a little time. Settings shows the price and the most time they add. Past that time, a search answers in the plain order. Only an owner can change the settings.

The decision model hides a match it judges too weak to answer the question. Its judgement of a match is kept for a day, so the same search gives the same results, and a repeat costs nothing for the matches it has judged. When the decision model does not answer, the search shows the plain order, including matches it would have hidden.

When the platform moves an index to a new embedding model, the index keeps answering with the old one until every document is embedded again.

A visitor can search what a published loop searches, through the loop’s own delivery:

  • The embed’s search box. On the loop’s Keys tab, tick Add a search box before you copy the snippet. See Embedding the Chat.
  • The React SDK’s search hook, useSearch, for a page you build yourself. The SDK is not published yet; see Embedding the Chat.
  • The run API, POST /api/loops/{loopId}/search, with a key that has the search scope. See Search in the API reference.

A visitor never searches a collection index, even one the loop uses: a collection’s records can hold your users’ data. A loop that uses only collection indexes offers visitors no search.

A visitor types words, or a phrase to find exactly. Each result shows its title, where it sits in its document, a passage, and the page’s address for a crawled page. A visitor never sees a cost, an index’s name or a score. Each search is paid from the wallet and counts against the loop’s monthly budget and the visitor’s own budget; when one has no room left, the visitor reads that search is unavailable, and no search runs.

Delete an index

Open the index and press Delete index. Its documents and everything built from them are removed.

What an index costs

Crawling, browser renders and indexing are paid from the account’s wallet, and each crawl reserves its estimate before it starts. A crawl the wallet cannot cover does not start, and the index says so.