The global index is what an unscoped search reads. If the pages you need are not in it, you can pay to have them crawled into an index of your own.
An index is a membership view over one shared document store, not a separate copy. A URL that two indexes both want is crawled, canonicalized and attested exactly once, so you are paying for coverage rather than for duplication.
Priced per page and quoted from the request body alone, so the amount in the 402 is exactly
the amount you sign for.
| Field | Default | Notes |
|---|---|---|
urls |
required | Non-empty. The exact pages to crawl |
name |
Untitled index |
The slug is derived from it |
visibility |
listed |
unlisted keeps it out of the public catalog |
readPolicy |
open |
allowlist restricts reads to named wallets |
allowlist |
[] |
Wallets; the owner is always implicit |
The payer owns what they commissioned. Ownership and allowlist checks run after the facilitator verifies the payment and before settlement, so a rejected caller is never charged.
There is no "index this host". Creation takes the list of pages, and getting that list is
your job: fetch the site's sitemap.xml, filter it to what you want, and send those URLs. A
seed URL on its own indexes exactly one page.
That is a pricing constraint rather than a missing feature. The price is quoted in the 402
and signed for on the retry, so it has to be a pure function of the request body. A host that
expanded to an unknown number of pages could not be priced before the crawl, and the commission
would be a blank cheque.
There is a limit on how many URLs one request may carry, because the body is parsed before the
payment is checked. It is a transport bound, not a ceiling on how large an index may become:
over it the API returns 400 with the limit and what you asked for, and the fix is to split
the list across requests rather than to ask for less. Commission the first batch, then append
the rest.
status is derived from the crawl queue, never stored, so it cannot disagree with the work
outstanding: pending before anything is crawled, crawling while some is, ready when no
queued URL remains.
Search is always scoped. Omitting index is not "search everything", it is a scoped search
that resolved to the global index.
It does not discover. A commissioned crawl fetches exactly the URLs that were paid for. Link following would fetch pages nobody bought and charge for pages nobody asked about.
It does not skip robots. Robots is still read and still obeyed, because paying us cannot confer a right to fetch. A URL you paid for that robots disallows is not fetched, and you are not charged differently for it: the honest-crawler rules are not for sale.
readPolicy: "allowlist" restricts reads to the owner and named wallets. It needs no second
auth mechanism, because a verified x402 payment already proves control of the payer wallet.
The check sits between verification and settlement, so a wallet that may not read an index learns so without being charged.
An unlisted index is absent from the public catalog entirely. Combined with allowlist, that
hides your curation. It does not hide the crawling: the pages are public pages, fetched by an
honest crawler that identifies itself, and the attestations say nothing about which index
wanted them.