Quick start
This walks through running Ferroseek on your machine with the generated sample catalog: build the engine, build an index, start the standalone Worker with Wrangler and query it with curl. It takes a few minutes, most of it compiling Rust.
Prerequisites
Section titled “Prerequisites”-
Node.js and npm. The repository is an npm workspace;
npm installat the root installs Wrangler, TypeScript and the workspace packages. -
A Rust toolchain (stable) with the WebAssembly target, which the engine is compiled to:
rustup target add wasm32-unknown-unknown -
Docker, only if you want to run the conformance suite against Elasticsearch later. It is not needed for this guide.
The packages are not published to npm. Everything below runs inside a clone of the repository.
1. Clone and install
Section titled “1. Clone and install”git clone https://github.com/eldrin-project/ferroseek.gitcd ferroseeknpm install2. Generate the sample catalog
Section titled “2. Generate the sample catalog”npm run data:generateThis writes data/mapping.json and two catalogs as NDJSON, one {"_id": …, "_source": …} object per line: data/generated/products-2k.ndjson (2,000 products) and data/generated/products-100k.ndjson (100,000 products). Generation is seeded, so the files are byte-identical on every run.
Each product has a sku, name, description, brand, category, tags, color, price, rating, stock, in_stock and created_at. The mapping is the body of an Elasticsearch mappings object:
{ "properties": { "name": { "type": "text" }, "brand": { "type": "keyword" }, "price": { "type": "double" }, "in_stock": { "type": "boolean" }, "created_at": { "type": "date" } }}3. Build an index
Section titled “3. Build an index”The index builder is a native Rust binary (build-index, compiled in release mode on first use). Start with the small catalog:
npm run index:build:2kFor the full 100k-product catalog used in the performance measurements:
npm run index:build:100kEither command writes the index to worker/assets/index/: a manifest.json and one or more part-NNN.bin files of at most 20 MiB each. The 100k index is three parts, 58.3 MiB in total.
4. Run the Worker
Section titled “4. Run the Worker”npm run devnpm run dev compiles the engine to WebAssembly (npm run wasm:build, which copies the result to packages/ferroseek/wasm/ferroseek.wasm) and starts wrangler dev with worker/wrangler.jsonc. The Worker listens on port 4090, set in that file.
The index is loaded lazily, on the first request. That request takes longer than the ones after it.
5. Send a first query
Section titled “5. Send a first query”The Worker serves the Elasticsearch routes at its root. Check that it answers like an Elasticsearch 7.17.4 node:
curl -s http://localhost:4090/Then search. This finds products whose name matches “wireless speaker” and returns three fields of the top two hits:
curl -s http://localhost:4090/products/_search \ -H 'Content-Type: application/json' \ -d '{ "query": { "match": { "name": "wireless speaker" } }, "size": 2, "_source": ["sku", "name", "brand", "price"] }'The response is an Elasticsearch search response. Against the 100k catalog it looks like this:
{ "took": 2, "timed_out": false, "_shards": { "total": 1, "successful": 1, "skipped": 0, "failed": 0 }, "hits": { "total": { "value": 4354, "relation": "eq" }, "max_score": 8.515811, "hits": [ { "_index": "products", "_type": "_doc", "_id": "SKU-002213", "_score": 8.515811, "_source": { "sku": "SKU-002213", "name": "Zakeli Wireless Speaker", "brand": "Zakeli", "price": 28.17 } } ] }}(The second hit is left out here.) Count documents with a query, or fetch one by id:
curl -s http://localhost:4090/products/_count \ -H 'Content-Type: application/json' \ -d '{ "query": { "term": { "color": "red" } } }'
curl -s http://localhost:4090/products/_doc/SKU-0000016. Look at the timing headers
Section titled “6. Look at the timing headers”The standalone Worker turns on the engine’s diagnostics, so every response carries timing and memory headers. Use curl -i to see them:
| Header | Meaning |
|---|---|
x-engine-time-us |
Time spent inside the engine for this request, in microseconds |
x-wasm-memory-bytes |
Current size of the WebAssembly memory |
x-index-load-ms |
Index load time; only on the request that loaded the index |
Inside workerd, performance.now() is coarsened to whole milliseconds, so x-engine-time-us is quantised to multiples of 1000 on a single request. For a meaningful number, average over many requests; worker/scripts/measure.ts does that:
npx tsx worker/scripts/measure.ts --url http://localhost:4090 --n 300Next steps
Section titled “Next steps”- Call the engine from your own Worker: Embed the engine in a Worker.
- Use the official Elasticsearch client against it: Connect with the client.
- See what the query DSL supports: Queries.