memoception

Measured. Same judge. Reproducible.

Every number here comes from our own harness on the public LoCoMo and LongMemEval datasets — including the results that went against us. Our rules:

The field

The basis column is the one that matters: ours is measured and re-judgeable; everyone else's is self-reported on their own harness (one 84% claim was later corrected to 58.44% by outside review).

systemLoCoMo answer accuracyLongMemEvalbasis
👑 Memoception70.2% overall · 62.7% excl-adv · 96.2% abstention96.4% retrieval@10✅ measured — re-judgeable
Mem066.9% excl-adv (paper) · 92.5 composite (marketing)94.4 (marketing)self-reported
Supermemory95% Recall@15 (aggregated)self-reported
Zep84% → corrected to 58.44%self-reported, corrected
Mastra95% (research page)self-reported
70.2%
LoCoMo overall · complete set
96.4%
LongMemEval retrieval@10
96.2%
adversarial abstention
~28ms
search p50 · local

Head-to-head — identical harness, identical judge

The only comparison we publish as a ranking: both systems ingest the same conversation, same prompts, same judge (gpt-4o-mini both sides; Mem0 OSS v2.0.12, conversation 0, 199 questions).

overall accuracy
Memoception: 60.3%Memoception60.3%Mem0 OSS: 51.8%Mem0 OSS51.8%
excl. adversarial
Memoception: 52%Memoception52%Mem0 OSS: 39.5%Mem0 OSS39.5%

Also: search p50 ~28ms local vs 442ms remote · zero-cost ingestion tier. Caveats we publish: one conversation; Mem0's FAISS mode disables their hybrid keyword search.

LoCoMo by category

Full 1,986 questions, claude-haiku-4.5 answerer + judge, zero dropped. Adversarial questions count as correct only when the system abstains.

abstention: 96.2%abstention96.2%single-hop: 73.5%single-hop73.5%temporal: 67%temporal67%multi-hop: 37.6%multi-hop37.6%open-domain: 27.1%open-domain27.1%

Same engine under three judges — internally comparable, full runs each: 70.2% haiku-4.5 · 67.7% mistral-small · 63.3% gpt-4o-mini¹

¹ 1,740/1,986 — 246 dropped on rate limits, counted and excluded, never silent.

LongMemEval & cost

What did NOT work

Negative results stay in the record — that's what makes the positive ones believable.

Reproduce it

Questions or corrections: hello@memoception.com