Skip to content
Conversational AISociety

Technology. Human experience.
The conversations in between.

← Back to latest key findings
Engineering · Evaluation

LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory

A 500-question benchmark testing information extraction, reasoning across sessions, time, knowledge updates and abstention.

What it contributes

Organises memory evaluation around distinct abilities and connects performance to indexing, retrieval and reading choices. It supplies a practical vocabulary for diagnosing memory failures.

Read with care

Performance on questions grounded in conversation histories does not establish relationship quality, privacy protection or safe real-world behaviour. Model comparisons reflect the versions and conditions tested.

Why it belongs here

For engineers choosing memory architectures and evaluators looking beyond simple recall.

Source status

ICLR 2025 conference paper; arXiv v2 dated 4 March 2025.

Source reading depth

Abstract, version history and ICLR publication status checked. This is a source note, not a full critical review.

Source check: 1 October 2026. Reading cautions are editorial interpretation, not quotations from the authors.