LoCoMo · recall@10 · full pipeline
96.58%
1918/1986 questions · run 2026-07-29 ·0 LLM calls in ranking · ~0.84s per query on CPU
+2.1 pp over the published open-source leader (94.5% · PMB),measured on the same benchmark.
- 96.58% evidence recall@10 on LoCoMo. +2.1 pp vs published open-source leader (94.5%)
- 80.8% end-to-end binary accuracy on BEAM-100K (323/400 correct across 20 conversations; reader GLM-5.3 Flash (Ox Alpha); judge GLM-5.2; effort=max; run 2026-08-23; ~44% slower than default effort; the reader is the customer's choice — results vary with the model (stronger readers can score higher). ValorBrain doesn't sell LLM calls: the memory layer is the product)
- LoCoMo retrieval: ~0.84s per query on CPU · 0 LLM calls in ranking
Full per-conversation table
Canonical run with 1986 questions across 10 conversations. The table total is the headline above (same math, no cherry-picked average).
| Conversation | Documents | Hits | Questions | R@10 |
|---|---|---|---|---|
| 26 | 19 | 191 | 199 | 96% |
| 30 | 19 | 102 | 105 | 97.1% |
| 41 | 32 | 188 | 193 | 97.4% |
| 42 | 29 | 252 | 260 | 96.9% |
| 43 | 29 | 236 | 242 | 97.5% |
| 44 | 28 | 154 | 158 | 97.5% |
| 47 | 31 | 180 | 190 | 94.7% |
| 48 | 30 | 230 | 239 | 96.2% |
| 49 | 25 | 187 | 196 | 95.4% |
| 50 | 30 | 198 | 204 | 97.1% |
| Total | 272 | 1918 | 1986 | 96.57603222557906% |