The number, and how to cite it.
Genesys Memory scores 85.55 ± 0.37 on LoCoMo under a frozen protocol. This page is the citation surface: the exact quotable claim, the ten certified runs behind it, the comparison, and the disclosures — so a publisher or a model can extract one accurate sentence and nothing more.
One sentence, exact.
Genesys Memory scores 85.55 ± 0.37 on LoCoMo, certified across ten runs with a frozen gpt-4o-mini answerer and judge at temperature 0 (n=1,540, categories 1–4; evaluated July 2026).
The exact sentence — copy it verbatim.@misc{genesys-locomo-2026,
title = {Genesys Memory: LoCoMo Benchmark Results},
author = {Astrix Labs},
year = {2026},
month = {jul},
note = {J-score 85.55 ± 0.37, mean of 10 runs;
frozen gpt-4o-mini answerer and judge, temperature 0;
n=1540, LoCoMo categories 1--4},
howpublished = {\url{https://genesys.astrixlabs.ai/benchmarks/locomo}}
}What was frozen.
The answerer and judge are the strictest published configuration — the Mem0 paper's setup (arXiv:2504.19413), reused verbatim by Zep. Whatever is frozen defines comparability.
| Evaluated | July 2026 |
| Dataset | LoCoMo |
| Questions | 1,540 (n) |
| Categories | 1–4 |
| Answerer | gpt-4o-mini · frozen |
| Judge | gpt-4o-mini · frozen |
| Temperature | 0 |
| Runs | 10 independent full runs |
Ten runs. One number.
Ten independent full runs of the frozen protocol. The certified figure is their mean; the spread is the standard deviation across runs.
| Run | LoCoMo J-score |
|---|---|
| 01 | 84.87 |
| 02 | 85.84 |
| 03 | 85.71 |
| 04 | 85.58 |
| 05 | 85.26 |
| 06 | 85.78 |
| 07 | 85.52 |
| 08 | 86.17 |
| 09 | 85.52 |
| 10 | 85.19 |
| Mean ± std | 85.55 ± 0.37 |
Same yardstick.
Published numbers for systems measured under a comparable LoCoMo setup — same dataset, same J-score metric.
| System | LoCoMo score |
|---|---|
| Genesys | 85.55 |
| Zep | 75.14 |
| Mem0 | 66.9 |
Comparability note.Cross-vendor LoCoMo numbers are only comparable when the answer model, the judge model, and the question set match. The rows above share the dataset and metric; where a vendor's published pipeline differs, treat small gaps as not meaningful. Marketing figures above 90 that use non-comparable configurations are excluded here.
Non-comparable, disclosed: swapping only the answering model for a current one (gpt-4.1-mini) — same memories, same retrieval, same judge — scores 89.68. It is never presented as the headline, because it changes the frozen answerer and so is not comparable to the numbers above.
What the number does not say.
Every figure that qualifies the headline, published with it — not buried. If a memory vendor tells you their system only adds correct answers, they have not measured.
Check our work.
The full method — frozen protocol, per-category results, the failure ledger, and the measured ceiling — is the dossier. The engine is open source (AGPLv3). The raw run figures on this page are published as JSON.