Skip to content
ferroseekdocumentation homemain, packages at version v0.1.0
Ferroseek on GitHub

Build and update the index

Ferroseek does not index documents at runtime. There is no _bulk, no PUT _doc, no refresh: the index is built offline by a native Rust program, deployed as static files with the Worker, and copied into memory when the Worker starts. To change what is searchable, you build a new index and deploy it.

The builder reads two files.

A mapping, as the body of an Elasticsearch mappings object. Only flat properties with a type are accepted:

Mapping type Stored as
text Analyzed with the standard analyzer, with positions for phrase queries
keyword Exact values, for filters, facets and sorting
double, float, half_float 64-bit floating point
integer, long, short, byte 64-bit integer
boolean Boolean
date Milliseconds since the epoch

Any other key in a field definition, such as analyzer or fields, is rejected, as is any other type.

Documents as NDJSON, one per line, each with its id and source:

{"_id":"SKU-000001","_source":{"sku":"SKU-000001","name":"Sonece Wireless Crib A100","brand":"Sonece","price":44.99,"in_stock":true,"created_at":"2023-10-09T18:57:56Z"}}

keyword and text fields may hold arrays (the sample catalog’s tags does). Numeric, boolean and date fields must be single-valued; the builder rejects arrays there.

cargo run --release --manifest-path engine/Cargo.toml --bin build-index -- \
--mapping data/mapping.json \
--input data/generated/products-100k.ndjson \
--name products \
--out worker/assets/index

The root scripts npm run index:build:2k and npm run index:build:100k run exactly this for the two sample catalogs. --name is the index name the engine answers to: requests for any other index get 404 index_not_found_exception. One Worker serves one index.

The output directory gets a manifest.json and part files:

worker/assets/index/manifest.json
{
"format": "ferroseek-index",
"version": 2,
"name": "products",
"doc_count": 100000,
"total_bytes": 61101616,
"parts": [
{ "file": "part-000.bin", "offset": 0, "bytes": 20971520 },
{ "file": "part-001.bin", "offset": 20971520, "bytes": 20971520 },
{ "file": "part-002.bin", "offset": 41943040, "bytes": 19158576 }
]
}

The format is binary and usable in place: loading copies each part’s bytes into WebAssembly memory at its offset, without parsing. Parts are at most 20 MiB because Workers static assets accept files up to 25 MiB.

The 100k-product sample catalog (about 400 to 500 bytes of JSON per product) gives a 58.3 MiB index, and a Worker that has loaded it uses about 62 MB of its 128 MB. The index holds the original _source of every document as well as the search structures, so the size grows with both the number of products and the size of each one. Check total_bytes in the manifest against the memory your Worker needs for everything else.

  1. Export the current catalog to NDJSON.
  2. Rebuild the index into the Worker’s assets directory.
  3. Deploy the Worker (see Deploy to Cloudflare).

Isolates that start after the deployment load the new index on their first request. There is no partial update: every change, even of one product, means a full rebuild and redeploy. Measure how long a rebuild takes for your catalog before relying on frequent updates.

For an embedded engine, copy the built index to wherever its assetsIndexSource reads from; the example shop’s npm run shop:index copies worker/assets/index/ to examples/shop-worker/public/search-index/.

The loader reads the index through a small interface, IndexSource: a manifest() that returns the parsed manifest and a part(file) that returns a stream of one part’s bytes. assetsIndexSource is the only implementation in the repository; another backend, such as R2, can be added by implementing these two methods. See Library API.