Skip to content
ferroseekdocumentation homemain, packages at version v0.1.0
Ferroseek on GitHub

Performance

These are measurements of the proof of concept, taken with the generated 100,000-product sample catalog. They describe what was observed under the conditions given, not a guarantee. Your catalog, queries, traffic and region will give different numbers; measure your own before relying on them.

Measure Result Where
CPU time per request, p50 1.4 ms Production, Cloudflare analytics
CPU time per request, p75 2.1 ms Production, Cloudflare analytics
CPU time per request, p99 8.3 ms Production, Cloudflare analytics
Round trip, median about 19 to 20 ms Production, from a client about 10 ms of network away
Faceted search, engine time about 0.3 to 0.4 ms Native, locally
Round trip about 1.5 ms Local wrangler dev
Index load (cold start) 226 to 313 ms Production
First request in a new isolate about 0.84 to 1.35 s Production
Memory after load about 62 MB of 128 MB Production
Index size 58.3 MiB in three files On disk

The catalog is data/generated/products-100k.ndjson: 100,000 generated products of roughly 400 to 500 bytes of JSON each, with text, keyword, numeric, boolean and date fields (see Quick start).

The “faceted search” is the reference query of the project contract, conformance/queries/100-shop-text-filters-facets.json: a two-word multi_match over name^3 and description with operator and, two filters (in_stock and a price range), five terms aggregations and one range aggregation, sorted by score then sku, 24 hits.

Under local wrangler dev, the faceted search took about 0.3 to 0.4 ms of engine time (measured natively), and a full HTTP round trip to the local Worker about 1.5 ms.

Inside workerd, performance.now() is coarsened to whole milliseconds, so engine time cannot be read reliably from one request; average over many. The repository’s tools for this are worker/scripts/measure.ts (cold load, engine time, HTTP latency and memory of a running Worker), npm run bench (a seeded mix of query templates at several concurrency levels) and npm run engine:bench (the engine natively, without a Worker).

A Ferroseek Worker with the 100k index was deployed to Cloudflare and sent about 2,800 mixed search requests.

CPU time per request, from Cloudflare’s analytics: p50 1.4 ms, p75 2.1 ms, p99 8.3 ms. This is the time the Worker spent computing; it excludes time spent waiting on the network.

Round trip from a client about 10 ms of network away: median about 19 to 20 ms. Most of that is network, not search.

A new isolate loads the index on its first request: it fetches the three part files from static assets and copies them into WebAssembly memory.

  • Index load: 226 to 313 ms. The loader streams the three parts concurrently; loading them one after another took 390 to 461 ms.
  • First request in total: about 0.84 to 1.35 s.
  • The first request after new index files had been uploaded took 4.5 s.

After loading the 100k index, the isolate used about 62 MB of its 128 MB limit. The rest is shared by your application and by request-time working memory, which the engine keeps to roughly 16 MiB per request with the limits in Limits and differences. WebAssembly memory never shrinks, so the figure to watch is the peak.

The project contract set these targets at 100,000 documents:

Metric Budget Measured
Engine time, faceted query p50 under 5 ms, p95 under 20 ms about 0.3 to 0.4 ms natively; production CPU p50 1.4 ms, p99 8.3 ms per request
Cold start: fetch parts, make the index queryable under 500 ms 226 to 313 ms
Peak memory in the isolate, including load under 100 MB about 62 MB after load