Skip to content
ferroseekdocumentation homemain, packages at version v0.1.0
Ferroseek on GitHub

Quick start

This walks through running Ferroseek on your machine with the generated sample catalog: build the engine, build an index, start the standalone Worker with Wrangler and query it with curl. It takes a few minutes, most of it compiling Rust.

  • Node.js and npm. The repository is an npm workspace; npm install at the root installs Wrangler, TypeScript and the workspace packages.

  • A Rust toolchain (stable) with the WebAssembly target, which the engine is compiled to:

    rustup target add wasm32-unknown-unknown
  • Docker, only if you want to run the conformance suite against Elasticsearch later. It is not needed for this guide.

The packages are not published to npm. Everything below runs inside a clone of the repository.

git clone https://github.com/eldrin-project/ferroseek.git
cd ferroseek
npm install
npm run data:generate

This writes data/mapping.json and two catalogs as NDJSON, one {"_id": …, "_source": …} object per line: data/generated/products-2k.ndjson (2,000 products) and data/generated/products-100k.ndjson (100,000 products). Generation is seeded, so the files are byte-identical on every run.

Each product has a sku, name, description, brand, category, tags, color, price, rating, stock, in_stock and created_at. The mapping is the body of an Elasticsearch mappings object:

data/mapping.json (excerpt)
{
"properties": {
"name": { "type": "text" },
"brand": { "type": "keyword" },
"price": { "type": "double" },
"in_stock": { "type": "boolean" },
"created_at": { "type": "date" }
}
}

The index builder is a native Rust binary (build-index, compiled in release mode on first use). Start with the small catalog:

npm run index:build:2k

For the full 100k-product catalog used in the performance measurements:

npm run index:build:100k

Either command writes the index to worker/assets/index/: a manifest.json and one or more part-NNN.bin files of at most 20 MiB each. The 100k index is three parts, 58.3 MiB in total.

npm run dev

npm run dev compiles the engine to WebAssembly (npm run wasm:build, which copies the result to packages/ferroseek/wasm/ferroseek.wasm) and starts wrangler dev with worker/wrangler.jsonc. The Worker listens on port 4090, set in that file.

The index is loaded lazily, on the first request. That request takes longer than the ones after it.

The Worker serves the Elasticsearch routes at its root. Check that it answers like an Elasticsearch 7.17.4 node:

curl -s http://localhost:4090/

Then search. This finds products whose name matches “wireless speaker” and returns three fields of the top two hits:

curl -s http://localhost:4090/products/_search \
-H 'Content-Type: application/json' \
-d '{
"query": { "match": { "name": "wireless speaker" } },
"size": 2,
"_source": ["sku", "name", "brand", "price"]
}'

The response is an Elasticsearch search response. Against the 100k catalog it looks like this:

{
"took": 2,
"timed_out": false,
"_shards": { "total": 1, "successful": 1, "skipped": 0, "failed": 0 },
"hits": {
"total": { "value": 4354, "relation": "eq" },
"max_score": 8.515811,
"hits": [
{
"_index": "products",
"_type": "_doc",
"_id": "SKU-002213",
"_score": 8.515811,
"_source": { "sku": "SKU-002213", "name": "Zakeli Wireless Speaker", "brand": "Zakeli", "price": 28.17 }
}
]
}
}

(The second hit is left out here.) Count documents with a query, or fetch one by id:

curl -s http://localhost:4090/products/_count \
-H 'Content-Type: application/json' \
-d '{ "query": { "term": { "color": "red" } } }'
curl -s http://localhost:4090/products/_doc/SKU-000001

The standalone Worker turns on the engine’s diagnostics, so every response carries timing and memory headers. Use curl -i to see them:

Header Meaning
x-engine-time-us Time spent inside the engine for this request, in microseconds
x-wasm-memory-bytes Current size of the WebAssembly memory
x-index-load-ms Index load time; only on the request that loaded the index

Inside workerd, performance.now() is coarsened to whole milliseconds, so x-engine-time-us is quantised to multiples of 1000 on a single request. For a meaningful number, average over many requests; worker/scripts/measure.ts does that:

npx tsx worker/scripts/measure.ts --url http://localhost:4090 --n 300