Skip to content
ferroseekdocumentation homemain, packages at version v0.1.0
Ferroseek on GitHub

Limits and differences

Ferroseek runs inside a 128 MB Worker isolate whose WebAssembly memory already holds the index (about 60 MB for 100,000 products). WebAssembly memory never shrinks, and the engine aborts if an allocation fails. So every request must stay within roughly 16 MiB of extra memory and roughly 100 ms of CPU on a 100k-document index, or be rejected before the work starts, with an Elasticsearch-shaped error.

All limits live in one file, engine/src/limits.rs. They are checked when the request is parsed, or before the work they bound begins. Budgets for clauses, buckets, work and response bytes are shared across the whole request, including all searches of an _msearch.

Where Elasticsearch has the same limit, Ferroseek uses its value and its error. The others are Ferroseek’s own and answer 400 illegal_argument_exception with a reason of the form [ferroseek] … exceeds the limit of [N].

Limit Value Error Same as Elasticsearch?
Request body size 256 KiB 413 illegal_argument_exception Elasticsearch’s http.max_content_length is 100 MB, also 413
Estimated heap of the parsed body 8 MiB 400 illegal_argument_exception Ferroseek only
Term clauses per request 1,024 400, too_many_clauses under a shard failure indices.query.bool.max_clause_count, but counted per request (see below)
bool nesting depth 20 400 x_content_parse_exception indices.query.bool.max_nested_depth
Values in one terms query 65,536 400, shard failure index.max_terms_count
Values in one ids query 65,536 400 illegal_argument_exception Ferroseek only
Buckets per response, all levels 65,536 503 too_many_buckets_exception search.max_buckets
Ranges in one range aggregation 1,024 400 illegal_argument_exception Ferroseek only
Queries in filter and filters aggregations 256 400 illegal_argument_exception Ferroseek only
Length of a date_histogram format 128 characters 400 illegal_argument_exception Ferroseek only
Sort keys 8 400 illegal_argument_exception Ferroseek only
_source include and exclude patterns 1,024 400 illegal_argument_exception Ferroseek only
Searches in one _msearch 32 400 illegal_argument_exception Ferroseek only
Response body 8 MiB, also for ?pretty 400 illegal_argument_exception Ferroseek only
aggregations section of one search 4 MiB less 64 KiB 400 illegal_argument_exception Ferroseek only
CPU work per request 30,000,000 work units 400 illegal_argument_exception Ferroseek only
from + size 10,000 400, shard failure index.max_result_window

Some context for the values:

  • Body size. A realistic search body is a few KiB. 256 KiB still fits a terms filter of about 15,000 SKUs, or a 32-search _msearch.
  • Parsed size. Parsed JSON can take from about 2 times to about 90 times its text size ([[1],[1],…] is the worst case), so the byte limit alone does not bound memory. The engine estimates the parsed size with a byte scan before parsing.
  • Clauses. match counts one clause per analyzed term, multi_match one per term per field.
  • Work units. A unit is one posting decoded, one column value scanned or one document visited; natively a unit costs about 1.2 to 1.7 ns, so the budget is roughly 40 to 55 ms. The contract’s faceted benchmark query uses well under one million.
  • Response size. A full size: 10000 page of the sample catalog is about 4.9 MB compact and 7.5 MB with ?pretty, so the largest page the result window allows still fits.

These are answered with 400 illegal_argument_exception and a reason starting [ferroseek] unsupported:.

Search body options: highlight, suggest, explain, version, seq_no_primary_term, stored_fields, docvalue_fields, script_fields, collapse, rescore, min_score, timeout, terminate_after, indices_boost, stats, profile, pit, runtime_mappings, fields, slice, ext.

Queries: fuzzy, wildcard, regexp, dis_max, function_score, query_string, simple_query_string, nested, script, script_score, boosting, combined_fields, geo_distance, more_like_this; named queries (_name); custom analyzers; multi_match types most_fields, cross_fields and bool_prefix. Parameter-level details are on Queries.

Aggregations: see Aggregations and facets.

Sorting: custom missing values, mode other than min and max, unmapped_type, numeric_type, format, nested.

APIs: anything not in the route table: no writes, no scroll, no point in time, no index or cluster management.

Mappings: only text, keyword, numeric, boolean and date fields with no options; numeric, boolean and date fields must be single-valued.

Some requests that Elasticsearch accepts are rejected:

  • URL parameters. Only ?pretty is accepted. Elasticsearch also takes search options such as size or q in the query string; Ferroseek rejects them so that a request means the same thing everywhere. Put options in the body.
  • Clause count per request. Lucene 8 checks the 1,024-clause limit per bool query; Ferroseek counts every clause in the request, across query, post_filter and aggregation filters. Very wide query trees that Elasticsearch accepts can be rejected.
  • Body size of 256 KiB against Elasticsearch’s 100 MB.
  • Every limit marked “Ferroseek only” in the table above.
  • Date math. Dates must be ISO 8601 or epoch milliseconds; now-7d and similar expressions are not parsed.

With typo tolerance, one match query with several terms can expand two different terms to the same dictionary term. For example, on name with fuzziness: "AUTO", "toe tte" expands to {hoe, tee, tote} and {tee, tote}. Lucene merges the duplicate clauses while rewriting the query and, when it does, rebuilds the clause list in HashMap order, which depends on JVM identity hash codes. Ferroseek uses first-occurrence order instead.

The only thing this can change is which document frequency a merged term is scored with, and only when two other terms of the same query also share an expansion. In that case scores can differ slightly from Elasticsearch’s, and the order of hits with close scores can differ with them. The conformance corpus includes queries with shared expansions, such as conformance/queries-stretch/320-fuzzy-two-tokens-shared-expansions.json.