Skip to content
ferroseekdocumentation homemain, packages at version v0.1.0
Ferroseek on GitHub

Ferroseek

Ferroseek is a search engine for product catalogs that runs inside a Cloudflare Worker. The engine is written in Rust and compiled to WebAssembly. At startup the Worker copies a prebuilt index into WebAssembly memory, and every search after that is answered from memory, in the same isolate that received the request.

It speaks the Elasticsearch REST API: a subset of the routes, query DSL and aggregations of Elasticsearch 7.17, with the same request bodies, response shapes and error envelopes. An application that uses the official Elasticsearch JavaScript client can talk to Ferroseek without code changes, and can later move to a real Elasticsearch cluster by changing its connection settings.

Search a running Ferroseek Worker
curl -s http://localhost:4090/products/_search \
-H 'Content-Type: application/json' \
-d '{ "query": { "match": { "name": "wireless speaker" } }, "size": 2 }'

Ferroseek is aimed at shops and catalogs of up to roughly 100,000 products that want faceted search close to their users without running a search cluster. At that size the whole index fits comfortably in a Worker: the sample 100k-product index is 58.3 MiB on disk and the isolate uses about 62 MB of its 128 MB limit once the index is loaded.

It is not a general-purpose replacement for Elasticsearch. The index is read-only at runtime: there is no indexing API, no cluster, no replication. You build the index offline from your catalog and deploy it with the Worker, then rebuild and redeploy when the catalog changes. See Build and update the index.

Ferroseek implements the part of Elasticsearch that a product listing page needs, and checks itself against a real Elasticsearch 7.17.4 server:

  • Same API. GET /, _search, _count, _doc, _mapping and _msearch take and return Elasticsearch 7.17 JSON. Every response carries the X-Elastic-Product: Elasticsearch header the official clients require.
  • Same ranking. Scoring is BM25 over the standard analyzer, including Lucene’s one-byte length norms, so hits come back in the same order with scores within 1e-4 of the server’s.
  • Same errors, louder. Malformed requests answer with Elasticsearch’s error envelope. Features that Ferroseek does not implement are rejected with a 400 that names the feature; they are never silently ignored.

Switching between the two is a connection-settings change. With the client package, the same application code runs against the engine embedded in the Worker, a standalone Ferroseek Worker, or an Elasticsearch cluster, chosen by environment variables alone.

Embedded. The engine runs inside your own Worker, in the same isolate as your application code. A search is a function call, with no network hop. See Embed the engine in a Worker.

Standalone. A dedicated Worker serves the Elasticsearch routes over HTTP, optionally behind Basic or API key authentication, and any Elasticsearch client can call it. See Run a standalone search Worker.

Ferroseek is a proof of concept. The packages in this repository are not published to npm; to use them today you clone the repository and work inside it. The figures on the Performance page are measurements of this proof of concept on a generated catalog, not guarantees.