Skip to content

Proof

Benchmarks, with the caveats on the same page

Retrieval numbers for the engine, each with its run and method declared. What the scores measure, and what they don't, is up here, not in the footer.

LoCoMo · recall@10 · full pipeline

96.58%

1918/1986 questions · run 2026-07-29 ·0 LLM calls in ranking · ~0.84s per query on CPU

+2.1 pp over the published open-source leader (94.5% · PMB),measured on the same benchmark.

  • 96.58% evidence recall@10 on LoCoMo. +2.1 pp vs published open-source leader (94.5%)
  • 80.8% end-to-end binary accuracy on BEAM-100K (323/400 correct across 20 conversations; reader GLM-5.3 Flash (Ox Alpha); judge GLM-5.2; effort=max; run 2026-08-23; ~44% slower than default effort; the reader is the customer's choice — results vary with the model (stronger readers can score higher). ValorBrain doesn't sell LLM calls: the memory layer is the product)
  • LoCoMo retrieval: ~0.84s per query on CPU · 0 LLM calls in ranking

Full per-conversation table

Canonical run with 1986 questions across 10 conversations. The table total is the headline above (same math, no cherry-picked average).

Recall@10 per conversation on LoCoMo, run 2026-07-29
ConversationDocumentsHitsQuestionsR@10
261919119996%
301910210597.1%
413218819397.4%
422925226096.9%
432923624297.5%
442815415897.5%
473118019094.7%
483023023996.2%
492518719695.4%
503019820497.1%
Total2721918198696.57603222557906%

BEAM-100K · end-to-end answers · binary accuracy

80.8%

323/400 correct answers across 20 conversations. Unlike LoCoMo: here the score is the final answer delivered to the user, not the retrieved passage.

Benchmark configuration: reader GLM-5.3 Flash (Ox Alpha); judge GLM-5.2; effort=max; run 2026-08-23; ~44% slower than default effort; the reader is the customer's choice — results vary with the model (stronger readers can score higher). ValorBrain doesn't sell LLM calls: the memory layer is the product

What we don't measure

LoCoMo measures retrieval; BEAM-100K measures end-to-end answers on synthetic conversations. Organizational memory (permission, origin, multi-person + multi-agent) is product capability outside both scores.

What the scores cover

  • Long-horizon conversational evidence retrieval (LoCoMo recall@10)
  • End-to-end multi-session memory answers (BEAM-100K binary accuracy)
  • Ranking without calling an LLM in the retrieval path

Outside both scores

  • Multi-tenant isolation or role-based permission filtering
  • Team handoffs, knowledge gaps, or veracity/origin badges
  • Production Ask citation fidelity or latency under live tenant workloads