Performance
These are measurements of the proof of concept, taken with the generated 100,000-product sample catalog. They describe what was observed under the conditions given, not a guarantee. Your catalog, queries, traffic and region will give different numbers; measure your own before relying on them.
Summary
Section titled “Summary”| Measure | Result | Where |
|---|---|---|
| CPU time per request, p50 | 1.4 ms | Production, Cloudflare analytics |
| CPU time per request, p75 | 2.1 ms | Production, Cloudflare analytics |
| CPU time per request, p99 | 8.3 ms | Production, Cloudflare analytics |
| Round trip, median | about 19 to 20 ms | Production, from a client about 10 ms of network away |
| Faceted search, engine time | about 0.3 to 0.4 ms | Native, locally |
| Round trip | about 1.5 ms | Local wrangler dev |
| Index load (cold start) | 226 to 313 ms | Production |
| First request in a new isolate | about 0.84 to 1.35 s | Production |
| Memory after load | about 62 MB of 128 MB | Production |
| Index size | 58.3 MiB in three files | On disk |
The catalog and the query
Section titled “The catalog and the query”The catalog is data/generated/products-100k.ndjson: 100,000 generated products of roughly 400 to 500 bytes of JSON each, with text, keyword, numeric, boolean and date fields (see Quick start).
The “faceted search” is the reference query of the project contract, conformance/queries/100-shop-text-filters-facets.json: a two-word multi_match over name^3 and description with operator and, two filters (in_stock and a price range), five terms aggregations and one range aggregation, sorted by score then sku, 24 hits.
Locally
Section titled “Locally”Under local wrangler dev, the faceted search took about 0.3 to 0.4 ms of engine time (measured natively), and a full HTTP round trip to the local Worker about 1.5 ms.
Inside workerd, performance.now() is coarsened to whole milliseconds, so engine time cannot be read reliably from one request; average over many. The repository’s tools for this are worker/scripts/measure.ts (cold load, engine time, HTTP latency and memory of a running Worker), npm run bench (a seeded mix of query templates at several concurrency levels) and npm run engine:bench (the engine natively, without a Worker).
In production on Cloudflare
Section titled “In production on Cloudflare”A Ferroseek Worker with the 100k index was deployed to Cloudflare and sent about 2,800 mixed search requests.
CPU time per request, from Cloudflare’s analytics: p50 1.4 ms, p75 2.1 ms, p99 8.3 ms. This is the time the Worker spent computing; it excludes time spent waiting on the network.
Round trip from a client about 10 ms of network away: median about 19 to 20 ms. Most of that is network, not search.
Cold start
Section titled “Cold start”A new isolate loads the index on its first request: it fetches the three part files from static assets and copies them into WebAssembly memory.
- Index load: 226 to 313 ms. The loader streams the three parts concurrently; loading them one after another took 390 to 461 ms.
- First request in total: about 0.84 to 1.35 s.
- The first request after new index files had been uploaded took 4.5 s.
Memory
Section titled “Memory”After loading the 100k index, the isolate used about 62 MB of its 128 MB limit. The rest is shared by your application and by request-time working memory, which the engine keeps to roughly 16 MiB per request with the limits in Limits and differences. WebAssembly memory never shrinks, so the figure to watch is the peak.
Against the budget
Section titled “Against the budget”The project contract set these targets at 100,000 documents:
| Metric | Budget | Measured |
|---|---|---|
| Engine time, faceted query | p50 under 5 ms, p95 under 20 ms | about 0.3 to 0.4 ms natively; production CPU p50 1.4 ms, p99 8.3 ms per request |
| Cold start: fetch parts, make the index queryable | under 500 ms | 226 to 313 ms |
| Peak memory in the isolate, including load | under 100 MB | about 62 MB after load |