APEX-EM: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay
Pratyay Banerjee, Masud Moshtaghi, Ankit Chadha
Agent memoryEMNLP 2026, accepted, to appear · Amazon
Accepted to EMNLP 2026. The proceedings have not been published yet, so this page cites the arXiv version. arXiv comment field, verified 2026-09-15: "22 pages, 6 figures, Accepted to EMNLP 2026".
APEX-EM introduced procedural knowledge graphs for LLM agent memory: its typed Procedural Knowledge Graph stores complete solution traces and retrieves them by meaning, operation structure, and graph traversal, with no model retraining.
Research contributionIntroduced procedural knowledge graphs for LLM agent memory through a typed Procedural Knowledge Graph, procedural-episodic experience replay, and cross-domain structural retrieval.
Abstract
LLM agents rerun full reasoning for every task, even one they solved moments earlier. We introduce APEX-EM, a non-parametric experience memory that stores complete procedural-episodic traces in a typed Procedural Knowledge Graph (PKG) and retrieves them through three channels: semantic search, structural-signature matching over abstract operation sequences, and graph traversal. A Plan-Retrieve-Generate-Iterate-Ingest (PRGII) workflow produces, quality-gates, and commits experiences, indexing both successes and failures so the agent learns what to reuse and what to avoid. No weights change during deployment. We evaluate on five benchmarks: BigCodeBench, KGQAGen-10k, HLE, Lifelong Agent Bench, and ALFWorld. Because prior work uses different backbones, we base our claims on same-backbone comparisons that hold model capability fixed. On held-out BigCodeBench transfer with a shared GPT-4o backbone, APEX-EM gains +7.6 pp over the no-memory baseline, 3.3× MemRL’s +2.3 pp under the identical setup. On Lifelong Agent Bench with a shared GPT-4o-mini backbone, it gains +1.4 pp (OS) and +1.0 pp (DB) cumulative success. On KGQAGen-10k, frozen memory transfers to a blind 1,079-question test split at 73.7% versus 42.0% with no memory, approaching an oracle handed the ground-truth subgraph (84.9%). Across three Opus scales the memory gain stays at +27 to +32 pp, so it adds to model capability rather than substituting for it. Component analysis shows no single mechanism dominates: teacher feedback is negligible for code but adds +10.3 pp on structured queries, structural signatures give 3.3× the transfer of semantic-only retrieval, and within-epoch iteration recovers most of the gain when rich feedback is unavailable. These results argue for modular memory composed per domain.
Why it matters
Most agent systems throw away what they learned as soon as a task ends, so the same work gets redone. APEX-EM keeps the full trace, including the failed attempts, in a graph and looks it up by the shape of the reasoning rather than by wording alone. That matters for teams who cannot fine-tune a hosted model: the gains come from storage and retrieval, and the paper reports lower cost per solved task as memory fills.
Key results
- On the blind 1,079-question KGQAGen-10k test split with memory frozen, structural retrieval (S1) transfers at 73.7% versus 42.0% for the no-memory baseline and 68.7% for semantic-only retrieval, approaching the 84.9% oracle that is handed the ground-truth supporting subgraph.
- On held-out BigCodeBench transfer with a shared GPT-4o backbone, APEX-EM reaches 53.5% for a stated +7.6 pp gain, described as 3.3 times MemRL's +2.3 pp (48.5% to 50.8%).
- On Humanity's Last Exam the memory gain holds across three Claude Opus checkpoints, lifting Opus 4.7 cumulative success from a 39.4% no-memory baseline to 71.2% (+31.8 pp), with Opus 4.5 at +28.1 pp and Opus 4.6 at +27.4 pp.
- Null or negativeAdding the Teacher judge does not help code generation: on BigCodeBench, A1 (semantic memory, no judge) reaches 80.1% while A2 (semantic plus judge) drops to 78.8%, whereas on KGQAGen-10k the judge adds +10.3 pp (75.6% to 85.9%).
- Null or negativeOn Lifelong Agent Bench OS, APEX-EM only ties MemRL at 78.8% last-epoch success rate, and its relative gain is smaller (+6.8 pp from a 72.0% baseline) than MemRL's (+11.4 pp from a 67.4% baseline); the same-backbone GPT-4o-mini margins are +1.4 pp (OS) and +1.0 pp (DB) cumulative success.
- Estimated API cost per correctly solved KGQAGen-10k task falls from $0.304 in epoch 1 to $0.137 in epoch 10 for A5, a 55% reduction, against $0.497 for the no-memory run, as average iterations drop from 4.40 to 2.50.
Every number above appears in the paper. Results that were null or went the wrong way are included and marked.
Topics
This paper introduces a typed Procedural Knowledge Graph for LLM agent memory and studies procedural memory, procedural-episodic experience replay, non-parametric online learning, structural-signature matching, self-evolving and self-improving agents, cross-domain structural retrieval, BigCodeBench, KGQAGen-10k, ALFWorld, Lifelong Agent Bench, Humanity's Last Exam, and case-based reasoning.
What this paper does not show
No human evaluation. The headline comparison against MemRL is not a controlled one, because each system runs on its own backbone; only 3 of 5 benchmarks get a same-backbone comparison, and where they do the margins are thin, plus 1.4 and plus 1.0 points on Lifelong and a tie at 78.8 percent on the OS split. The 3.3 times MemRL figure depends on which no-memory baseline the delta uses, and the paper does not say; against GPT-4o's own no-memory row the multiple is closer to 2.2 times. The teacher judge makes code generation worse, not better, and costs about 30 percent more compute. Cost and latency are estimated from list pricing, not measured.
How to cite
Pratyay Banerjee, Masud Moshtaghi, and Ankit Chadha. 2026. APEX-EM: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay. arXiv:2603.29093.
@misc{banerjee2026apexem,
title = {{APEX-EM}: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay},
author = {Banerjee, Pratyay and Moshtaghi, Masud and Chadha, Ankit},
year = {2026},
month = aug,
eprint = {2603.29093},
archivePrefix = {arXiv},
primaryClass = {cs.CL},
note = {arXiv:2603.29093v3. Accepted to EMNLP 2026, proceedings not yet published},
url = {https://arxiv.org/abs/2603.29093}
}
Related work of mine
- APEX-MEM: Agentic Semi-Structured Memory with Temporal Reasoning for Long-Term Conversational AIACL 2026 (Main Conference, Long Papers)
- Routed Graph Handoff: Adaptive Format Selection for Multi-Agent LLM DelegationEMNLP 2026, accepted, to appear
arXiv version, v3.