Certified benchmark · LoCoMo · July 2026

The number, and how to cite it.

Genesys Memory scores 85.55 ± 0.37 on LoCoMo under a frozen protocol. This page is the citation surface: the exact quotable claim, the ten certified runs behind it, the comparison, and the disclosures — so a publisher or a model can extract one accurate sentence and nothing more.

01 The citation

One sentence, exact.

Quotable claim — verbatim

Genesys Memory scores 85.55 ± 0.37 on LoCoMo, certified across ten runs with a frozen gpt-4o-mini answerer and judge at temperature 0 (n=1,540, categories 1–4; evaluated July 2026).

The exact sentence — copy it verbatim.
BibTeX
@misc{genesys-locomo-2026,
  title        = {Genesys Memory: LoCoMo Benchmark Results},
  author       = {Astrix Labs},
  year         = {2026},
  month        = {jul},
  note         = {J-score 85.55 ± 0.37, mean of 10 runs;
                  frozen gpt-4o-mini answerer and judge, temperature 0;
                  n=1540, LoCoMo categories 1--4},
  howpublished = {\url{https://genesys.astrixlabs.ai/benchmarks/locomo}}
}
02 The protocol

What was frozen.

The answerer and judge are the strictest published configuration — the Mem0 paper's setup (arXiv:2504.19413), reused verbatim by Zep. Whatever is frozen defines comparability.

EvaluatedJuly 2026
DatasetLoCoMo
Questions1,540 (n)
Categories1–4
Answerergpt-4o-mini · frozen
Judgegpt-4o-mini · frozen
Temperature0
Runs10 independent full runs
03 The certified runs

Ten runs. One number.

Ten independent full runs of the frozen protocol. The certified figure is their mean; the spread is the standard deviation across runs.

RunLoCoMo J-score
0184.87
0285.84
0385.71
0485.58
0585.26
0685.78
0785.52
0886.17
0985.52
1085.19
Mean ± std85.55 ± 0.37
04 Comparison

Same yardstick.

Published numbers for systems measured under a comparable LoCoMo setup — same dataset, same J-score metric.

SystemLoCoMo score
Genesys85.55
Zep75.14
Mem066.9

Comparability note.Cross-vendor LoCoMo numbers are only comparable when the answer model, the judge model, and the question set match. The rows above share the dataset and metric; where a vendor's published pipeline differs, treat small gaps as not meaningful. Marketing figures above 90 that use non-comparable configurations are excluded here.

Non-comparable, disclosed: swapping only the answering model for a current one (gpt-4.1-mini) — same memories, same retrieval, same judge — scores 89.68. It is never presented as the headline, because it changes the frozen answerer and so is not comparable to the numbers above.

05 Limitations & disclosures

What the number does not say.

Every figure that qualifies the headline, published with it — not buried. If a memory vendor tells you their system only adds correct answers, they have not measured.

CeilingAn oracle probe — answering from gold evidence alone — measures the ceiling of this frozen protocol at 94.9. Beyond it lie defective gold labels, judge strictness, and the answerer's own reasoning limits.
ChurnAt temperature 0, re-running the identical configuration still regenerates 5.4% of answers differently. The ± 0.37 spread across ten runs is the honest measure of that noise.
Gate G2-PThe certified configuration persistently answers 170 questions the naive baseline cannot, and persistently loses 72 it got right — a 2.4 : 1 trade. Our persistence gate flags this as exceeding the harm budget: RED. Certified as net-better with a disclosed trade, not as strictly better.
Non-comparableA run with a modern answerer (gpt-4.1-mini) scores 89.68. It changes the frozen answerer, so it is disclosed as a non-comparable row and never used as the headline.
06 Verify it

Check our work.

The full method — frozen protocol, per-category results, the failure ledger, and the measured ceiling — is the dossier. The engine is open source (AGPLv3). The raw run figures on this page are published as JSON.