Integrations

Sharded static indexes

Serve a large search index as lazily-fetched shards, with rankings identical to the full index.

@zbsearch/static splits a ZBSearch index into a set of static files the browser fetches lazily: a visitor downloads a small manifest and term dictionary up front, and each query then pulls only the few kilobytes of postings and documents it actually touches. A 10,000-page site answers its first search within a couple of hundred kilobytes, where the single-file index would weigh megabytes.

The trade is transfer size for round trips, and quality is kept out of the trade entirely: the client runs ZBSearch's own search() against the fetched data, so every ranking - BM25 scores, prefix expansion, typo tolerance, thresholds, boosts - is bit-identical to what the full in-memory index returns. The test suite asserts exactly that, query by query.

The Docusaurus, Starlight and VitePress plugins use this package automatically once a site's index outgrows a single file - there is nothing to install or configure there. Reach for @zbsearch/static directly when you are wiring search into something those plugins do not cover.

Installation

npm install @zbsearch/static

The file set

buildStaticIndex turns your records into a deployable directory:

FileRoleLoaded
manifest.jsonGlobal statistics (document count, average field lengths), the shard table, the format version and the build idUp front
dictionary.jsonEvery indexed term as a compact trie, without postingsUp front, ideally prefetched on focus
postings/<build>/*.binBinary postings shards, packed by sorted term ranges to roughly 40 KB eachPer query, only the shards covering its terms
fragments/<build>/*.jsonStored documents in small groupsOnly for the results being shown - everything up to the requested page

<build> is a hash of everything the client combines: the dictionary, every shard and fragment, and the manifest's own identity (language, schema, statistics and the document ID mapping), so even a rebuild that changes only the records' id values gets a new build id. The manifest names it, the dictionary carries it, and shards and fragments live under it. A redeploy or a stale cache therefore cannot combine files from two different builds: the client refuses a dictionary from another build outright, and asks for shards and fragments by the manifest's build id, where a mismatch shows up as a missing file instead of silently different results.

The dictionary staying resident is what keeps search quality intact: once it is loaded, prefix expansion and typo correction walk the local trie with no further network involved, so a misspelled query costs the same single round trip as a correct one. Each postings entry carries its term frequency and exact field length inline, which is why one shard fetch is enough to score a term exactly.

Building

import { buildStaticIndex } from '@zbsearch/static';
import { mkdir, writeFile } from 'node:fs/promises';
import path from 'node:path';

const { files, stats } = await buildStaticIndex({
  records,
  schema: { title: 'string', content: 'string' },
  language: 'english',
});

for (const [file, bytes] of files) {
  const target = path.join('dist/search', file);
  await mkdir(path.dirname(target), { recursive: true });
  await writeFile(target, bytes);
}

console.log(`${stats.termCount} terms in ${stats.shardCount} shards, ${stats.totalBytes} bytes`);
OptionDefaultDescription
records-The documents to index
schema-Property types; string properties only for now
language'english'Language used to tokenize and stem the index
targetShardBytes40960Byte budget per postings shard, before compression
fragmentGroupSize16Documents per fragment file

Records without an id receive sequential ones, which the client regenerates from the document count alone - zero bytes on the wire. Records that bring their own id keep it, and the mapping ships in the manifest.

Searching

import { createStaticSearchClient } from '@zbsearch/static';

const client = createStaticSearchClient({ baseUrl: '/search/' });

// Optional: fetch manifest and dictionary while the visitor is still reaching for the input.
searchButton.addEventListener('focus', () => client.preload());

const results = await client.search({ term: 'sharded quey', tolerance: 1, limit: 10 });

search() accepts the familiar ZBSearch options: term, properties, exact, tolerance, prefix, threshold, boost, relevance, limit, offset, where and preflight. Results have the same shape search(db, ...) returns, documents included.

A query travels like this:

  1. The term is tokenized and resolved against the resident dictionary - exact words, prefix expansions and typo corrections, all locally.
  2. The shards covering those words are fetched, in parallel, skipping any already in memory.
  3. ZBSearch's own search() runs against the merged postings.
  4. Fragments are fetched for the results up to the requested page (offset + limit), and the hits come back hydrated. Exact-match queries fetch every candidate's document, because the engine verifies exact against the stored text.

Every fetched shard and document stays merged for the rest of the session, so refining a query usually costs nothing further, and repeating one costs no requests at all. client.stats() reports requests, bytes and shard counts if you want to watch it happen.

fetchBytes swaps the transport - tests serve the file set straight from memory:

const client = createStaticSearchClient({
  fetchBytes: async (file) => files.get(file)!,
});

Compared with Pagefind

Pagefind established the lazy-loading model for static-site search; @zbsearch/static keeps that model and pairs it with ZBSearch's search quality. The repository ships a head-to-head benchmark (npm run benchmark:pagefind in benchmarks/) that indexes the identical corpus with both engines, serves both bundles from the same gzip-capable server and drives them in headless Chromium. What ZBSearch brings, measured:

  • Typos keep working. On queries with one character dropped from a word (the benchmark's typo battery covers deletions only, not insertions or substitutions) ZBSearch finds the target document 98% of the time, Pagefind 3% - the resident dictionary corrects a misspelling locally before anything is fetched, so a typo'd query costs the same single round trip as a correct one.
  • A cheaper first search. A cold visitor's first query completes in ~92 KB of transfer on a 1,500-page corpus and ~130 KB on a 10,000-page one - engine, index and results included - against Pagefind's ~123 KB and ~147 KB.
  • Fewer round trips per query. 0.4 requests per query on the 1,500-page corpus against 1.4; on real-world connections, round trips are what a visitor feels.
  • Flat tail latency at scale. At 10,000 pages the p95 query time is 8 ms against 71 ms.
  • Faster builds. Measured end to end, from records to files on disk for both engines, the build runs about 6-7× faster: 0.7 s against 5.1 s at 10,000 pages, 0.14 s against 0.9 s at 1,500.
  • Better prefix recall. Search-as-you-type finds the target within the top five for 97% of 4-character prefixes, against 88%.

Current limits

  • Only string properties are indexed; other property values still ride along on the stored documents.
  • facets, groupBy, distinctOn, sortBy and vector search are rejected with a clear error rather than returning something incomplete. They are on the roadmap.
  • The index is immutable once built - rebuild and redeploy, as with any static asset.

The builder and client version together through the manifest's format version: a client refuses an index written by an incompatible format rather than misreading it. Within a format, the build id keeps every file of a deploy together.

On this page