Guides · 5 min read

Why AgentKits Memory's Hybrid Search Has to Solve Japanese Twice

BM25 misses meaning. Vector search misses exact error strings. AgentKits Memory fuses both — but for Japanese, Chinese, and Korean queries, the keyword half of that fusion needs a second decision underneath it.

Why AgentKits Memory's Hybrid Search Has to Solve Japanese Twice

Two Queries That Break Two Different Ways

Ask an AI coding assistant’s memory for “authentication pattern” and you want it to find the entry that says “use JWT with refresh tokens” even though none of those words match literally — that’s a meaning problem, and keyword search is bad at it. Ask the same memory for the exact string ModuleNotFoundError: No module named 'networkx' and you want the one entry that contains that exact error, not five entries that are vaguely about Python — that’s an identity problem, and vector search is bad at it, because embedding similarity blurs precise strings into a neighborhood of “similar-ish” text.

AgentKits Memory, the local-first persistent memory layer for AI coding assistants we’ve written about before, doesn’t pick one. Its HybridSearchEngine runs SQLite FTS5 keyword search and local embedding-based vector search on every query, then fuses the two scores. That decision isn’t a hunch — it matches what retrieval research keeps finding when it measures the two approaches separately. On the WANDS e-commerce retrieval benchmark, a tuned hybrid setup reached 0.7497 NDCG against 0.6983 for BM25 alone and 0.6953 for vector search alone — each side wins some queries and loses others, and fusing them beats picking a favorite. (Denser AI)

How the Fusion Actually Works

The implementation is a straightforward weighted blend, not a black box: keyword score times 0.3 plus semantic score times 0.7 by default, both normalized to a 0–1 range before combining. BM25 ranking comes from FTS5’s own bm25() function; semantic similarity comes from cosine distance between a locally-generated query embedding and stored entry embeddings — no cloud embedding API, using the multilingual-e5-small ONNX model the README documents. When only one side finds a match, that entry keeps its single score rather than getting penalized for the other side’s silence.

The default 70/30 weighting toward semantics makes a real bet: for a memory system where entries are mostly prose decisions and summaries rather than product SKUs, meaning usually matters more than exact phrasing — but the 30% keyword floor is what keeps an exact error string or function name from getting lost in embedding space, which is precisely the failure mode dense-only retrieval is documented to have.

The CJK Problem Hiding Inside “Keyword”

Here’s the part that isn’t generic RAG advice: “keyword search” assumes you can tell where one word ends and the next begins, and Japanese, Chinese, and Korean text doesn’t give you that for free — there are no spaces between words. FTS5’s standard unicode61 tokenizer, built around word boundaries, simply doesn’t segment CJK text into anything searchable. AgentKits Memory’s answer is to default to FTS5’s trigram tokenizer, which indexes overlapping 3-character sequences instead of words — a strategy specifically built to work when you can’t reliably find word boundaries in the first place.

That default isn’t free of trade-offs, and the code is honest about where they show up. Character n-gram indexing for Japanese is documented in information-retrieval literature as prone to “spurious results” — unrelated words that happen to share an n-gram surface as false matches — while the alternative, proper morphological word segmentation, trades that for the opposite problem: a segmenter’s dictionary is always incomplete, so it silently misses whatever it hasn’t learned. (Whoosh docs on n-gram indexing; general Japanese-tokenization surveys make the same trade-off explicit for both directions). AgentKits Memory’s HybridSearchEngine codes around this rather than picking one side blindly: it negotiates the best tokenizer the local SQLite build actually supports (trigram, then porter, then unicode61, checked live at initialization), and for CJK queries shorter than 3 characters — too short for a trigram to even form — it drops to a plain LIKE scan instead of returning nothing. For teams that want real word-level Japanese segmentation instead of the n-gram compromise, the package exposes an optional lindera-sqlite backend as an explicit upgrade path, rather than forcing the trade-off on everyone by default.

Why the Savings Have to Be Protected, Not Just Won

Running two search engines and reconciling two tokenizer strategies is wasted engineering if the result gets exhaled straight into the model’s context window as full-text dumps. That’s what the accompanying TokenEconomicsTracker and the search engine’s 3-layer design (compact → timeline → full) are actually for: layer one returns roughly 50 tokens per hit — id, score breakdown, a 100-character snippet, an estimated token count — so an assistant can look at ten candidates for the price of one full entry before deciding which ones are worth fetching in full. The README documents this progressive-disclosure pattern saving roughly 70% of tokens against fetching everything up front, and separately claims up to 87% in some configurations; the point either way is that fusing two retrieval signals only pays off if the layer sitting on top of it doesn’t immediately spend the savings back.

The Honest Version

None of this makes AgentKits Memory’s search “solved” in some final sense — trigram tokenization is a compromise that’s been chosen and documented as one, not a claim that it matches a proper morphological analyzer. What it demonstrates is a specific kind of design discipline: a hybrid search engine that fuses keyword and semantic scores is table-stakes RAG architecture in 2026, but treating “keyword” as a single well-defined operation across English and CJK text — and building an explicit fallback ladder instead of assuming unicode61 covers everyone — is the part that only shows up once you actually try to make search work for a query written in Japanese.


AgentKits Memory is available on GitHub and npm as @aitytech/agentkits-memory. Questions about how it fits your workflow? Reach out at [email protected].

Explore Our Open Source

We build and maintain open-source tools for developers. Check out our repositories on GitHub.

View on GitHub

Related Articles