AC

APEX-EM: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay

Agent memoryEMNLP 2026, accepted, to appear · Amazon
Accepted to EMNLP 2026. The proceedings have not been published yet, so this page cites the arXiv version. arXiv comment field, verified 2026-09-15: "22 pages, 6 figures, Accepted to EMNLP 2026".

PDF arXiv BibTeX

In one sentence

APEX-EM introduced procedural knowledge graphs for LLM agent memory: its typed Procedural Knowledge Graph stores complete solution traces and retrieves them by meaning, operation structure, and graph traversal, with no model retraining.

Research contributionIntroduced procedural knowledge graphs for LLM agent memory through a typed Procedural Knowledge Graph, procedural-episodic experience replay, and cross-domain structural retrieval.

Abstract

LLM agents rerun full reasoning for every task, even one they solved moments earlier. We introduce APEX-EM, a non-parametric experience memory that stores complete procedural-episodic traces in a typed Procedural Knowledge Graph (PKG) and retrieves them through three channels: semantic search, structural-signature matching over abstract operation sequences, and graph traversal. A Plan-Retrieve-Generate-Iterate-Ingest (PRGII) workflow produces, quality-gates, and commits experiences, indexing both successes and failures so the agent learns what to reuse and what to avoid. No weights change during deployment. We evaluate on five benchmarks: BigCodeBench, KGQAGen-10k, HLE, Lifelong Agent Bench, and ALFWorld. Because prior work uses different backbones, we base our claims on same-backbone comparisons that hold model capability fixed. On held-out BigCodeBench transfer with a shared GPT-4o backbone, APEX-EM gains +7.6 pp over the no-memory baseline, 3.3× MemRL’s +2.3 pp under the identical setup. On Lifelong Agent Bench with a shared GPT-4o-mini backbone, it gains +1.4 pp (OS) and +1.0 pp (DB) cumulative success. On KGQAGen-10k, frozen memory transfers to a blind 1,079-question test split at 73.7% versus 42.0% with no memory, approaching an oracle handed the ground-truth subgraph (84.9%). Across three Opus scales the memory gain stays at +27 to +32 pp, so it adds to model capability rather than substituting for it. Component analysis shows no single mechanism dominates: teacher feedback is negligible for code but adds +10.3 pp on structured queries, structural signatures give 3.3× the transfer of semantic-only retrieval, and within-epoch iteration recovers most of the gain when rich feedback is unavailable. These results argue for modular memory composed per domain.

Why it matters

Most agent systems throw away what they learned as soon as a task ends, so the same work gets redone. APEX-EM keeps the full trace, including the failed attempts, in a graph and looks it up by the shape of the reasoning rather than by wording alone. That matters for teams who cannot fine-tune a hosted model: the gains come from storage and retrieval, and the paper reports lower cost per solved task as memory fills.

Key results

Every number above appears in the paper. Results that were null or went the wrong way are included and marked.

PRGII loop: plan, retrieve past successes and failures, generate, iterate, then quality-gate commit into the memory graph.
Figure 1(a): The PRGII workflow. A task flows Plan to Retrieve to Generate and Iterate to Ingest; retrieved successes and failures condition Generate, and the quality gate commits to memory M. Taken from the paper.

Topics

This paper introduces a typed Procedural Knowledge Graph for LLM agent memory and studies procedural memory, procedural-episodic experience replay, non-parametric online learning, structural-signature matching, self-evolving and self-improving agents, cross-domain structural retrieval, BigCodeBench, KGQAGen-10k, ALFWorld, Lifelong Agent Bench, Humanity's Last Exam, and case-based reasoning.

What this paper does not show

No human evaluation. The headline comparison against MemRL is not a controlled one, because each system runs on its own backbone; only 3 of 5 benchmarks get a same-backbone comparison, and where they do the margins are thin, plus 1.4 and plus 1.0 points on Lifelong and a tie at 78.8 percent on the OS split. The 3.3 times MemRL figure depends on which no-memory baseline the delta uses, and the paper does not say; against GPT-4o's own no-memory row the multiple is closer to 2.2 times. The teacher judge makes code generation worse, not better, and costs about 30 percent more compute. Cost and latency are estimated from list pricing, not measured.

How to cite

Pratyay Banerjee, Masud Moshtaghi, and Ankit Chadha. 2026. APEX-EM: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay. arXiv:2603.29093.

@misc{banerjee2026apexem,
  title         = {{APEX-EM}: Non-Parametric Online Learning for Autonomous Agents via Structured Procedural-Episodic Experience Replay},
  author        = {Banerjee, Pratyay and Moshtaghi, Masud and Chadha, Ankit},
  year          = {2026},
  month         = aug,
  eprint        = {2603.29093},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CL},
  note          = {arXiv:2603.29093v3. Accepted to EMNLP 2026, proceedings not yet published},
  url           = {https://arxiv.org/abs/2603.29093}
}

Related work of mine

arXiv version, v3.

← All publications