Limits and differences
Ferroseek runs inside a 128 MB Worker isolate whose WebAssembly memory already holds the index (about 60 MB for 100,000 products). WebAssembly memory never shrinks, and the engine aborts if an allocation fails. So every request must stay within roughly 16 MiB of extra memory and roughly 100 ms of CPU on a 100k-document index, or be rejected before the work starts, with an Elasticsearch-shaped error.
All limits live in one file, engine/src/limits.rs. They are checked when the request is parsed, or before the work they bound begins. Budgets for clauses, buckets, work and response bytes are shared across the whole request, including all searches of an _msearch.
Request limits
Section titled “Request limits”Where Elasticsearch has the same limit, Ferroseek uses its value and its error. The others are Ferroseek’s own and answer 400 illegal_argument_exception with a reason of the form [ferroseek] … exceeds the limit of [N].
| Limit | Value | Error | Same as Elasticsearch? |
|---|---|---|---|
| Request body size | 256 KiB | 413 illegal_argument_exception |
Elasticsearch’s http.max_content_length is 100 MB, also 413 |
| Estimated heap of the parsed body | 8 MiB | 400 illegal_argument_exception |
Ferroseek only |
| Term clauses per request | 1,024 | 400, too_many_clauses under a shard failure |
indices.query.bool.max_clause_count, but counted per request (see below) |
bool nesting depth |
20 | 400 x_content_parse_exception |
indices.query.bool.max_nested_depth |
Values in one terms query |
65,536 | 400, shard failure |
index.max_terms_count |
Values in one ids query |
65,536 | 400 illegal_argument_exception |
Ferroseek only |
| Buckets per response, all levels | 65,536 | 503 too_many_buckets_exception |
search.max_buckets |
Ranges in one range aggregation |
1,024 | 400 illegal_argument_exception |
Ferroseek only |
Queries in filter and filters aggregations |
256 | 400 illegal_argument_exception |
Ferroseek only |
Length of a date_histogram format |
128 characters | 400 illegal_argument_exception |
Ferroseek only |
| Sort keys | 8 | 400 illegal_argument_exception |
Ferroseek only |
_source include and exclude patterns |
1,024 | 400 illegal_argument_exception |
Ferroseek only |
Searches in one _msearch |
32 | 400 illegal_argument_exception |
Ferroseek only |
| Response body | 8 MiB, also for ?pretty |
400 illegal_argument_exception |
Ferroseek only |
aggregations section of one search |
4 MiB less 64 KiB | 400 illegal_argument_exception |
Ferroseek only |
| CPU work per request | 30,000,000 work units | 400 illegal_argument_exception |
Ferroseek only |
from + size |
10,000 | 400, shard failure |
index.max_result_window |
Some context for the values:
- Body size. A realistic search body is a few KiB. 256 KiB still fits a
termsfilter of about 15,000 SKUs, or a 32-search_msearch. - Parsed size. Parsed JSON can take from about 2 times to about 90 times its text size (
[[1],[1],…]is the worst case), so the byte limit alone does not bound memory. The engine estimates the parsed size with a byte scan before parsing. - Clauses.
matchcounts one clause per analyzed term,multi_matchone per term per field. - Work units. A unit is one posting decoded, one column value scanned or one document visited; natively a unit costs about 1.2 to 1.7 ns, so the budget is roughly 40 to 55 ms. The contract’s faceted benchmark query uses well under one million.
- Response size. A full
size: 10000page of the sample catalog is about 4.9 MB compact and 7.5 MB with?pretty, so the largest page the result window allows still fits.
Rejected features
Section titled “Rejected features”These are answered with 400 illegal_argument_exception and a reason starting [ferroseek] unsupported:.
Search body options: highlight, suggest, explain, version, seq_no_primary_term, stored_fields, docvalue_fields, script_fields, collapse, rescore, min_score, timeout, terminate_after, indices_boost, stats, profile, pit, runtime_mappings, fields, slice, ext.
Queries: fuzzy, wildcard, regexp, dis_max, function_score, query_string, simple_query_string, nested, script, script_score, boosting, combined_fields, geo_distance, more_like_this; named queries (_name); custom analyzers; multi_match types most_fields, cross_fields and bool_prefix. Parameter-level details are on Queries.
Aggregations: see Aggregations and facets.
Sorting: custom missing values, mode other than min and max, unmapped_type, numeric_type, format, nested.
APIs: anything not in the route table: no writes, no scroll, no point in time, no index or cluster management.
Mappings: only text, keyword, numeric, boolean and date fields with no options; numeric, boolean and date fields must be single-valued.
Where Ferroseek is stricter
Section titled “Where Ferroseek is stricter”Some requests that Elasticsearch accepts are rejected:
- URL parameters. Only
?prettyis accepted. Elasticsearch also takes search options such assizeorqin the query string; Ferroseek rejects them so that a request means the same thing everywhere. Put options in the body. - Clause count per request. Lucene 8 checks the 1,024-clause limit per
boolquery; Ferroseek counts every clause in the request, acrossquery,post_filterand aggregation filters. Very wide query trees that Elasticsearch accepts can be rejected. - Body size of 256 KiB against Elasticsearch’s 100 MB.
- Every limit marked “Ferroseek only” in the table above.
- Date math. Dates must be ISO 8601 or epoch milliseconds;
now-7dand similar expressions are not parsed.
Known scoring difference
Section titled “Known scoring difference”With typo tolerance, one match query with several terms can expand two different terms to the same dictionary term. For example, on name with fuzziness: "AUTO", "toe tte" expands to {hoe, tee, tote} and {tee, tote}. Lucene merges the duplicate clauses while rewriting the query and, when it does, rebuilds the clause list in HashMap order, which depends on JVM identity hash codes. Ferroseek uses first-occurrence order instead.
The only thing this can change is which document frequency a merged term is scored with, and only when two other terms of the same query also share an expansion. In that case scores can differ slightly from Elasticsearch’s, and the order of hits with close scores can differ with them. The conformance corpus includes queries with shared expansions, such as conformance/queries-stretch/320-fuzzy-two-tokens-shared-expansions.json.