APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AI
Pratyay Banerjee, Masud Moshtaghi, Shivashankar Subramanian, Amita Misra, Ankit Chadha
Agent memoryACL 2026 (Main Conference, Long Papers), pages 16470–16489 · Amazon
PDF arXiv DOI Publisher page BibTeX
APEX-MEM uses an append-only property graph of temporally grounded, entity-centric events and a multi-tool retrieval agent to resolve evolving information at query time.
Research contributionIntroduced an append-only property graph memory that represents conversations as temporally grounded, entity-centric events and resolves evolving information at retrieval time.
Abstract
Large language models still struggle with reliable long-term conversational memory: simply enlarging context windows or applying naïve retrieval often introduces noise and destabilizes responses. We present APEX-MEM, a conversational memory system that combines three key innovations: (1) a property graph which uses domain-agnostic ontology to structure conversations as temporally grounded events in an entity-centric framework, (2) append-only storage that preserves the full temporal evolution of information, and (3) a multi-tool retrieval agent that understands and resolves conflicting or evolving information at query time, producing a compact and contextually relevant memory summary. This retrieval-time resolution preserves the full interaction history while suppressing irrelevant details. APEX-MEM achieves 88.88% accuracy on LOCOMO’s Question Answering task and 86.2% on LongMemEval, outperforming state-of-the-art session-aware approaches and demonstrating that structured property graphs enable more temporally coherent long-term conversational reasoning.
Why it matters
Chatbots that must remember months of conversation usually either stuff everything into the prompt, which adds noise, or overwrite old facts, which destroys the record of what changed and when. APEX-MEM keeps every version of a fact with its date and decides which one applies at question time. Practitioners get better answers on time-sensitive questions, at the cost of a heavier build step and many tool calls per query.
Key results
- Raises overall accuracy on the LOCOMO conversational-memory benchmark from 85.38% for MIRIX, the previous best system, to 88.88% for APEX-MEM with GPT5 as the question-answering agent, a gain of 3.50 percentage points.
- Raises the LongMemEval overall score from 74.6% for Nemori and 72.5% for a session-aware SimpleSearch RAG baseline to 86.2% for APEX-MEM with Claude 4.5 Sonnet, gains of 11.6 and 13.7 percentage points.
- On LOCOMO temporal questions, accuracy is 90.63% for APEX-MEM versus 65.62% for MIRIX and 75.71% for Mem0, which the authors attribute to never overwriting superseded facts.
- Null or negativeIn the tool ablation on LOCOMO with Claude 4.5 Haiku, adding the hybrid Search tool raises overall accuracy from 79.45% to 87.00% but lowers temporal accuracy from 82.29% to 79.17%, so the GraphSQL-only setup is better on the paper's headline reasoning category.
- Null or negativeOn SealQA-Hard, APEX-MEM with GPT5 reaches 40.15% accuracy against 38.6% for plain GPT5 with a web-search tool, a margin of about 1.5 points; the paper's claimed 5.55 point gain is measured instead against O3 at 34.6%.
- Null or negativeOn LOCOMO the plain full-context GPT4o baseline scores 87.52% overall, above the 86.35% APEX-MEM reaches when it also uses GPT4o as the question-answering agent, so the memory system only pulls ahead by switching to a stronger backend model.
Every number above appears in the paper. Results that were null or went the wrong way are included and marked.
Topics
This paper studies long-term conversational memory through property graph memory, temporal reasoning, retrieval-time resolution, a domain-agnostic ontology, entity-centric events, append-only knowledge graphs, multi-tool retrieval agents, LOCOMO, LongMemEval, SealQA, GraphSQL text-to-SQL, LLM agent memory, and retrieval-augmented generation.
What this paper does not show
Every accuracy number here comes from an LLM judge, not from human raters. Append-only storage is the paper's headline idea and we never ablated it directly, so the evidence for it is indirect. Full-context GPT-4o beats our GPT-4o configuration on LOCOMO, 87.52 against 86.35, and adding search hurts temporal accuracy. On SealQA-Hard our 5.55 point margin is measured against O3, but the same table has GPT-5 with web search at 38.6, which makes the honest margin about 1.5 points. Table 1 and Table 3 also disagree on the Haiku overall score, 84.92 against 87.00, so quote the GPT-5 number and not that one.
How to cite
Pratyay Banerjee, Masud Moshtaghi, Shivashankar Subramanian, Amita Misra, and Ankit Chadha. 2026. APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AI. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 16470–16489.
@inproceedings{banerjee2026apexmem,
title = {{APEX-MEM}: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AI},
author = {Banerjee, Pratyay and Moshtaghi, Masud and Subramanian, Shivashankar and Misra, Amita and Chadha, Ankit},
booktitle = {Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)},
year = {2026},
pages = {16470--16489},
publisher = {Association for Computational Linguistics},
eprint = {2604.14362},
archivePrefix = {arXiv},
primaryClass = {cs.CL}
}
Related work of mine
- APEX-EM: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience ReplayEMNLP 2026, accepted, to appear
- Routed Graph Handoff: Adaptive Format Selection for Multi-Agent LLM DelegationEMNLP 2026, accepted, to appear
arXiv version. Cited 14 times as counted by Google Scholar on 2026-09-15.