Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion docs/glossary.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,7 @@ Instead of a heavy state management library like Redux, MiniSearch uses a minima

### Reranker

A secondary search stage that takes initial results from SearXNG and re-orders them based on relevance to the query using a cross-encoder model (`jina-reranker-v1-tiny-en`) running in-process via ONNX Runtime.
A secondary search stage that takes initial results from SearXNG and re-orders them based on relevance to the query using a multilingual cross-encoder model (`mmarco-mMiniLMv2-L12-H384-v1`) running in-process via ONNX Runtime.

- **Implementation**: Loads the model's ONNX export with `onnxruntime-node`; no child process
- **Health Check**: Polls `/health` endpoint via `getRerankerStatus`
Expand Down
55 changes: 32 additions & 23 deletions docs/reranking.md
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@ The reranking subsystem consists of three components:

The `rerankerServiceHook` starts the reranker during server initialization:

1. Downloads the model and tokenizer from HuggingFace if not present (`jinaai/jina-reranker-v1-tiny-en`)
1. Downloads the model and tokenizer from HuggingFace if not present (`cross-encoder/mmarco-mMiniLMv2-L12-H384-v1`)
2. Creates an ONNX Runtime inference session
3. Performs a warmup inference (`query: "test"`, `documents: ["test document"]`) to ensure the graph is initialized
4. Sets `isReady = true`
Expand All @@ -29,30 +29,30 @@ There is no child process, port, or health endpoint: inference runs inside the N

`downloadFileFromHuggingFaceRepository` compares each cached file against the size the Hub reports, so a file downloaded only in part is replaced instead of being loaded. Without that check a truncated model made every later startup fail with `Protobuf parsing failed`, and only deleting `server/models/` by hand recovered.

Downloads land on a sibling `.part-<pid>` path and are renamed once the whole body is on disk, which keeps an interrupted write from being visible where the next startup would trust it. Files are also checked individually, so a missing tokenizer is fetched on its own while the 130MB model stays cached.
Downloads land on a sibling `.part-<pid>` path and are renamed once the whole body is on disk, which keeps an interrupted write from being visible where the next startup would trust it. Files are also checked individually, so a missing tokenizer is fetched on its own while the 119MB model stays cached.

The size check costs one metadata request per file at startup (roughly 600ms for the three of them). When the Hub is unreachable the cached files are used as they are, so an offline server still starts with a warm cache.

### Execution Providers

The session is created with `["webgpu", "cpu"]`, with no configuration to set. WebGPU is roughly 3x faster than CPU (24ms against 77ms for 30 documents) and agrees with it to within float32 rounding (1e-6, identical ordering), so it is preferred where it works.
The session is created with `["cpu"]`, with no configuration to set, and there is no second attempt to fall back from.

When that session fails to be created, the reason is logged and a second session is created with `["cpu"]`. The retry is what makes hosts without a GPU work: a trailing `cpu` in the provider list is not a fallback chain, because ONNX Runtime only falls back per operator once a provider has been registered. A provider that fails to initialize at all rejects `InferenceSession.create` outright, and the entries behind it never get a chance. On Hugging Face Spaces, for instance, Dawn cannot find `libvulkan.so.1`, WebGPU then reports `No supported adapters`, and before the retry existed that left the reranker permanently unready with every search falling back to unranked results. The providers of each attempt are logged.

Note that ONNX Runtime's WebGPU provider here is native, part of the `onnxruntime-node` binary. It is not the browser API, so it needs neither a browser nor Deno.
The model ships as a dynamically quantized graph, and the WebGPU provider has no kernels for its integer matmuls. It registers happily and then hands every one of them back to the CPU, paying a round trip each time: 812ms against 172ms of the same work, with scores drifting by as much as 1.15 and reordering results. GPU acceleration is therefore not on the table for this model, which also removes the reason the two-attempt session logic existed.

| Provider | Availability in the Node binding | Notes |
|----------|----------------------------------|-------|
| `cpu` | Everywhere | Fallback, retried as its own session |
| `webgpu` | Windows, Linux x64, macOS | Preferred; experimental in ONNX Runtime, and needs a real GPU adapter plus a Vulkan loader on Linux |
| `cpu` | Everywhere | The only provider used |
| `webgpu` | Windows, Linux x64, macOS | Not used: no integer-matmul kernels, so a quantized graph runs ~4.7x slower than on CPU and its scores drift |
| `cuda` | Linux x64 (CUDA v12) | Not used: the binaries are not bundled, and would need `npm install onnxruntime-node --onnxruntime-node-install=cuda12` |
| `coreml` | macOS | Not used: slower than CPU for this model's dynamic shapes |

There is no GPU provider for Linux arm64, so those hosts always run on CPU.
### One Document per Call

Documents are scored one at a time rather than in batches, which is what keeps a score a property of its own `(query, document)` pair.

### Batching
Dynamic quantization computes each activation tensor's scale from that tensor's own range at runtime. Padding rows out to a shared width puts the pad positions inside that range, and while the attention mask keeps them out of attention it cannot keep them out of the scale, so the quantization of the real tokens shifts. Batches of 10 moved logits by up to 1.29 at ordinary snippet lengths and reordered 2 of 10 fixtures depending only on which documents shared a batch. The fp32 export of the same model shows a difference of exactly zero, so this is a property of quantization and not of the graph.

Documents are scored in batches of 10. `onnxruntime-node` wraps a synchronous native call, so scoring all 30 results at once would block the event loop for the full duration. Batching yields between calls, capping the stall at roughly 27ms rather than 78ms, at the cost of about 4% more wall time. Scores are identical either way, because padding is per batch but the attention mask excludes it.
Scoring one pair at a time removes the padding, and with it the coupling. It also suits the reason batching was introduced in the first place: `onnxruntime-node` wraps a synchronous native call, and one pair holds the event loop for about 13ms instead of the 50ms a batch of 10 takes. The cost is throughput, roughly 12% more wall time for 30 documents on two threads and up to 47% more where there are cores to spare.

### Shutdown

Expand All @@ -70,6 +70,8 @@ Each result is formatted as `` `${title}\n${snippet}` ``: the cased title and sn

Documents are sent to the reranker whole. Truncation happens by tokens inside the reranker, not by characters here: `score()` caps each encoded `(query, document)` pair at `MAX_SEQUENCE_LENGTH` tokens, dropping tokens from the end of the document only, so the query and the trailing separator are preserved.

`MAX_SEQUENCE_LENGTH` is 512 because the model has 514 learned position embeddings and XLM-RoBERTa reserves two of them. It is a limit, not a preference: a 513-token pair fails outright with `indices element out of data bounds` at the position-embedding gather. Web snippets land around 33-63 tokens per pair, so truncation is rare in practice.

### Unicode Sanitization

`sanitizeUnicodeSurrogates()` validates Unicode surrogate pairs in input strings. Invalid surrogates are replaced with the Unicode replacement character (`�`). This prevents failures when processing malformed UTF-8 from web search results.
Expand Down Expand Up @@ -115,23 +117,31 @@ Reranking is applied to both text and image search results. For image results, t

| Property | Value |
|----------|-------|
| Model | jina-reranker-v1-tiny-en |
| Format | ONNX (fp32) |
| HuggingFace Repo | jinaai/jina-reranker-v1-tiny-en |
| Type | Cross-encoder reranker |
| Size | 4 layers, 33M parameters |
| Max sequence length | 8192 tokens (ALiBi); capped at 2048 via `MAX_SEQUENCE_LENGTH` |
| Storage | `server/models/jinaai/jina-reranker-v1-tiny-en/` |
| Model | mmarco-mMiniLMv2-L12-H384-v1 |
| Format | ONNX, dynamically quantized (`onnx/model_quint8_avx2.onnx`, 119MB) |
| HuggingFace Repo | cross-encoder/mmarco-mMiniLMv2-L12-H384-v1 |
| Type | Cross-encoder reranker (XLM-RoBERTa, one relevance logit) |
| Size | 12 layers, 118M parameters, of which 96M are the 250k-token vocabulary |
| Max sequence length | 512 tokens, the model's own limit, enforced via `MAX_SEQUENCE_LENGTH` |
| Storage | `server/models/cross-encoder/mmarco-mMiniLMv2-L12-H384-v1/` |

It is trained on mMARCO, the machine-translated multilingual MS MARCO, and is genuinely multilingual: a 250k-token sentencepiece vocabulary shared across languages, which is also where most of the parameter count goes.

The model it replaced, `jina-reranker-v1-tiny-en`, was chosen on the belief that an English model ranked other languages well enough. Measured against 240 MIRACL queries across English, Spanish, French, German, Russian and Japanese, it scored 0.4636 nDCG@10, below the 0.4963 of the un-reranked first-stage order: on non-English content it was not adding ranking signal. This model scores 0.7992 on the same set. Part of the mechanism is the tokenizer, since an English WordPiece vocabulary spends 2.17x as many tokens on Russian and about 1.4x on Spanish, German and Japanese, so the English model paid more compute to see worse-fragmented subwords.

### Choosing the Quantized Export

The repository ships one build per CPU kernel family from the same weights: `qint8_arm64`, `qint8_avx512`, `qint8_avx512_vnni` and `quint8_avx2`. `quint8_avx2` is the one used, because unsigned activations sidestep the signed-int8 saturation that x64 without VNNI otherwise has to work around, and it measured no slower than the arm64 build when run on arm64. The `qint8_arm64` and `qint8_avx512` builds are bit-identical in output and score 0.8039, marginally above `quint8_avx2`, but that margin is inside the noise of a 240-query set and does not buy predictable behaviour on hosts whose instruction set cannot be known in advance.

Despite being an English model, it ranks non-English results (Portuguese, for example) well in practice, which is why it is preferred over larger alternatives.
Quantization costs nothing measurable here: 0.7992 against 0.7973 nDCG@10 for the fp32 export, for a quarter of the download and half the latency. That is a different conclusion from the one that held for the previous 33M-parameter model, where the `q8` export did measurably degrade ranking; the loss does not carry over to a model with 12 layers and a large embedding table.

Quantized variants are deliberately not used. The `q8` export measurably degrades ranking quality on this 33M-parameter model, and the `fp16` export fails to load in `onnxruntime-node`.
The `O4` (fp16) export loads without complaint but is slower than fp32 on CPU, so it is not used either.

## Testing

`server/rerankerService.test.ts` runs in the default suite and covers Unicode sanitization plus the execution-provider fallback, with the ONNX Runtime session, the tokenizer, and the model download all mocked.
`server/rerankerService.test.ts` runs in the default suite and covers Unicode sanitization, token truncation, and that every document is sent as its own unpadded row, with the ONNX Runtime session, the tokenizer, and the model download all mocked.

`server/rerankerService.integration.test.ts` loads the real model and asserts ranking quality against English and Portuguese fixtures. It downloads ~130MB, so it is excluded from the default suite:
`server/rerankerService.integration.test.ts` loads the real model and asserts ranking quality against English and Portuguese fixtures. It downloads ~136MB, the model plus a 17MB sentencepiece tokenizer, so it is excluded from the default suite:

```sh
npx vitest run --config vitest.integration.config.ts
Expand All @@ -142,7 +152,6 @@ npx vitest run --config vitest.integration.config.ts
| Scenario | Behavior |
|----------|----------|
| Reranker not ready | Falls back to unranked SearXNG results |
| GPU provider unavailable | Logged, then the session is created again with `["cpu"]` |
| Model fails to load | Logged by the hook; reranker stays unready and search returns unranked results |
| Cached file of the wrong size | Logged, then downloaded again before the session is created |
| Download shorter than the reported size | Throws before anything is written, so the cache keeps no partial file |
Expand Down
10 changes: 5 additions & 5 deletions server/rerankerService.integration.test.ts
Original file line number Diff line number Diff line change
Expand Up @@ -2,8 +2,8 @@

/**
* Exercises the real reranker model end to end, including the multilingual
* behaviour that motivated picking jina-reranker-v1-tiny-en. Downloads ~130MB
* on first run, so it is excluded from the default suite:
* behaviour that motivated picking mmarco-mMiniLMv2-L12-H384-v1. Downloads
* ~136MB on first run, so it is excluded from the default suite:
*
* npx vitest run --config vitest.integration.config.ts
*/
Expand Down Expand Up @@ -200,9 +200,9 @@ describe("reranker service", () => {
.map(({ index }) => index);

// Every relevant result must outrank every irrelevant one. The model
// clears this with a score gap of at least 0.99 between the two groups,
// so it is not sensitive to the ~1e-6 difference between the CPU and
// WebGPU execution providers.
// clears this with a score gap of at least 5.7 between the two groups, so
// the assertion is about ranking behaviour and not about a threshold the
// model happens to sit on.
const topIndices = ordered
.slice(0, fixture.relevant.length)
.sort((a, b) => a - b);
Expand Down
Loading
Loading