Skip to content

ResearchSeptember 19, 202618 min read

The Shadow Memory Stack: What 900+ Sources Say About the Hole in the Agent Memory Market

Eight deep-research cycles and 900+ sources on agent memory: frameworks solve memory inside their own perimeter, practitioners improvise the rest, and continuity (the thing that makes teams of agents not redo work) has no owner. Benchmarks, hidden costs, the nine unpublished gaps, and the Brazilian regulatory reality.

ValorBrain team

We ran eight deep-research cycles (an agent with a supervisor and per-source subagents, search through a self-hosted SearXNG, GLM models, human curation), collecting 900+ sources and citing about 190 after mechanical citation validation. The subject: the state of agent memory. Architectures, benchmarks, latency and cost, the five main vendors, the demand side, forgetting and consolidation, LGPD, and the Brazilian market. This post is the synthesis, and the thesis fits in one sentence: the market sells storage and calls it memory; the thing that makes a team of agents not redo work (continuity) has no owner.

Before the content, the method, because research without method is just a blog post: every claim below links to a primary or documented source, accessed on September 19, 2026. Where the evidence did not close (and it did not close in several places), we say so instead of filling the gap with generalities. And we work on a memory engine; this research existed because we had to decide what to build and how to sell it. Read accordingly.

Every framework solves memory at home

Agent frameworks made real progress in 2025-2026, inside their own perimeter. CrewAI has native memory with four types (short-term, long-term, entities, external), shared across the agents of a crew [2]. LangGraph persists state by thread_id and offers long-term memory via a store, with the LangMem layer for cross-thread continuity [3]. The OpenAI Agents SDK hands off between agents, transferring conversation history with fine control over how much history crosses [4].

The pattern is always the same: memory exists inside the framework's own construct: the crew, the thread, the conversation. A CrewAI crew does not read a LangGraph store. A Codex agent cannot see a Claude agent's state. When we searched the official docs of all three for any bridge between universes, the result was negative across the board, and not for lack of trying: each page defines memory; none defines interoperability.

Two well-known corporate retrospectives illustrate what happens in that void. Cognition (the Devin company) published "Don't Build Multi-Agents," arguing that parallel, unsynchronized contexts lead to conflicting decisions [5]. Anthropic answered with the engineering of their own multi-agent research system, where coordination works (inside a single architecture they designed) [6]. Both readings are compatible with our conclusion: multi-agent works when someone owns the common memory; it breaks when every agent carries its own.

The shadow stack is a symptom, not a solution

Back to the stack from the opening. It is not hobbyist eccentricity: it is what the community does because the product does not exist. Practitioners document the regime openly: accumulated memory files, manual snapshots, hand-maintained overwrite rules [1]. What this stack lacks is everything that would make it a system: entity deduplication across different agents' memories, arbitration when two team members believe contradictory things, handoff serialization that survives a framework switch, governance of who writes what.

The cost of not having it already has a famous case study. On July 21, 2025, Replit's agent deleted a company's production database during a change it had described as frozen, fabricated data to cover the failures, and reported success [7]. The episode is usually told as a "rogue agent" story. Our reading is different: it is the most serious documented materialization of divergence between what the agent believes and what the human knows, precisely the absence of synchronization between team state and reality that the shadow stack pretends to solve.

There is more evidence the pain scales with time horizons. Chroma measured and named it, context rot: reading quality degrades with context size, so "throw everything in the prompt" is not continuity, it is postponement [8]. At ultra-long horizons, SWE-Marathon (June 2026) catalogs hundreds of agent failure modes in long-duration tasks [9]. And the Hacker News thread "Agentic Coding Is a Trap" documents accumulated rework as a structural failure mode, not an accident [10].

The benchmark circus

If continuity engineering is the hole, benchmarks are the circus ring. The most instructive episode in the category came from Zep itself: in January 2025 their temporal knowledge graph paper contested Mem0's LoCoMo SOTA, showing that a full-context baseline (feeding the whole conversation to the model, no engine at all) beats most commercially available systems (~73% vs. ~68% J-score), and that swapping only the judge model moves the scoreboard by up to ~10 points [11]. Then came the rematch: a public issue reopened Zep's own 84% claim and corrected it to 58.44% under corrected evaluation [12]. Meanwhile Hacker News circulated "AI Startup Caught Cheating on Benchmark Papers" [13]. Nobody came out of that story clean; everyone came out with marketing.

The lesson is not "benchmarks are useless." It is that vendor numbers open the conversation and never close it. LoCoMo has known limitations [33], LongMemEval tries to fix them [34], BEAM was born to scale where memory value becomes measurable [31][32], and Mastra showed a different architectural path with 95% on LongMemEval [35], but every result comes with the harness of whoever published it. Buyers should always ask: who ran it, with which judge, compared against what.

The same applies to cost. Independent measurement of eight memory systems across 2,176 tasks landed far above what pricing pages suggest (on the order of US$ 341 per 1,000 answers [15]) because real memory cost includes the LLM that extracts memories, the embedding of every fact, and the BYOK pattern that shifts the bill to the customer without shifting predictability. And there is silence: Letta, supermemory, MemOS, LangGraph and the pgvector-based solutions publish no latency numbers at all. The overhead of memory MCP servers only started being measured academically in 2026 [16][17]. In infrastructure, what goes unpublished is usually the number the market does not want you to see.

A side note, because it ties cost to quality: an open issue on Mem0's repository reports that 97.8% of ingested memories were junk [14]. Memory that only accumulates is not an asset; it is a liability with interest, the subject of the consolidation section below.

Academia formalized it; the market doesn't sell it

While the marketing fights in the circus ring, 2026 was the year academic research formalized exactly the problems the products ignore. A June survey dedicated to "persistent memory, state, and governance in LLM agents" legitimized the field as a discipline [27]. A September paper on tamper-evident evidence for agent runs appeared, but targets runs, not memory storage itself [28]. SIGIR 2026 demonstrated the problem we care most about here: deleting the data from storage does not delete its downstream influence: the "incomplete forgetting" that makes right to erasure and persistent memory a structural tension [29].

In products, the most sophisticated mechanism we found in a primary source is Graphiti's temporal edge invalidation, in Zep's paper: facts get validity windows, new invalidates old [19]. That is real bi-temporality, and temporal only. There is no notion of source authority (can a human contradict an agent? can an import contradict a human?), no automated curation of the team's lessons, no auditable trail of who changed what. A cognitive-science-inspired memory survey maps what theory suggests (consolidation, replay, interleaving [30]), and Letta even shipped something close with sleep-time compute [23], but these are isolated exceptions in a market still fighting over who extracts more facts from a chat.

The segment's dominant pattern, meanwhile, is known: an Apache 2.0 core scoped as a library, a paid cloud, and no complete free self-hosted product. The most instructive case is Zep again: the only player that kept a complete self-hosted Community Edition discontinued it: deprecated in April 2025, with further removals in February 2026; only the Graphiti framework remains open source [36][37][38]. Defensive licensing across data infrastructure tells the same story in cycles: MongoDB went SSPL in 2018 [39], Redis left BSD in 2024 (the Valkey fork was born) and reverted to AGPL in 2025 [40], Elastic abandoned open source in 2021 (OpenSearch was born) and re-adopted it in 2024 [41]. Source-available buys time, creates forks, and the market converges back. For anyone deciding how to sell self-hosted today, the history points one way: it is a paid tier with support (where what you sell is a response-time SLA, never uptime) or it is nothing.

The buyer, and Brazil

On the demand side, the category is young and already produced relevant rounds: Cognee raised a US$ 7.5M seed [47] and Supermemory US$ 2.6M with Cloudflare and Google executives on the cap table [48], not counting the giants embedding memory into their platforms, which is the most threatening move: OpenAI with ChatGPT memory [44], Anthropic with Claude memory [45], Bedrock with agent memory [46]. When the memory layer becomes a model-provider feature, the independent vendor has to sell what the platform cannot: continuity across models and frameworks, and governance over data that crosses tenant boundaries.

In Brazil, it is worth correcting a pitch that became a mantra: LGPD does not require local data. International transfer is legitimate under the standard clauses of Resolution CD/ANPD 19/2024, with the grace period ending in August 2025 [49][50]; CMN 4.893 requires cyber risk governance, not on-premises [51]. The AI bill (PL 2338/2023) is still stuck in Congress [52] and the EU AI Act was postponed by the omnibus [53]. The correct sales argument for Brazilian enterprise is contractual safeguard and verifiable architecture, not "on-prem because LGPD." That said, enforcement arrives even without an AI law: European GDPR has already produced enforcement against chatbot memory [54], and the technical risks are documented: semantic retrieval returning data that "was deleted" [55][56].

The nine gaps nobody published

The inventory below comes from our continuity research; every item was verified by absence, with the search location recorded. No framework or commercial product documents solving:

  1. Native shared memory across distinct frameworks — the bridge between a CrewAI crew, a LangGraph store and an OpenAI SDK handoff does not exist in any of the three vendors' official docs [2][3][4][1].
  2. Entity identity and deduplication across different agents' memories — today it is a practitioner's manual rule [1].
  3. Arbitration of contradictions between team beliefs — contradictions show up only as failure diagnostics in the retrospectives [5][6].
  4. Duplicate-work detection across agents on different frameworks [10].
  5. Universal handoff serialization — the only documented contract is filterable conversation history [4]; real formats in production are ad hoc markdown [1].
  6. Team lessons-and-mistakes memory with automated curation — a manual, individual version exists; nothing for the team [1].
  7. Synchronizing human state (Slack, Jira, PRs, approvals) with the team's memory — the absence the Replit incident turned into a headline [7][8].
  8. Return briefings for absent humans — nothing beyond manual snapshots [1].
  9. Memory governance — who writes, who deletes, per-project scope; in the shadow stack, nonexistent by construction [1].

After this text was written, we went back and verified these five names against primary sources. The result strengthened the diagnosis: none of the five mechanisms examined (AutoGen, AG2, MetaGPT, MCP, A2A) offers persistent, cross-run, cross-framework team memory as a first-class primitive: in AutoGen, memory is per-agent with persistence delegated to integrations; in MetaGPT the pool dies with the run; and MCP (stable spec 2025-11-25) and A2A (now a Linux Foundation project, v1.0.0, maintained by a committee with AWS, Cisco, Google, IBM, Microsoft, Salesforce, SAP and ServiceNow) leave shared state out of scope by design [57][58][59][60][61].

What it would take to actually solve this

Add the sections up and the spec writes itself. A memory system for teams of agents (heterogeneous agents, different frameworks, working with humans over weeks-long work) would need: memory accessible by protocol rather than framework coupling; entity identity and deduplication across writers; contradiction arbitration with an explicit authority hierarchy; serialized handoff that survives a tool switch; automated curation of lessons with deliberate forgetting of dead memory; synchronization with human state where it lives; return briefings readable by a person; and an auditable provenance trail on every fact.

We wrote that list as a research finding. It also describes, item by item, what we have spent the last months building. That is not coincidence and not a moral to the story: it is the reason we ran the eight cycles: to find out whether what we were building had a market that would confirm or refute it. The sources confirmed the hole. If you disagree with any point, the links are below. That is what they are for.


References

Accessed and validated on September 19, 2026. The full set of collected sources (900+) lives in our own ValorBrain instance; below, the ones cited in the text.

[1] r/AI_Agents, "Has anyone actually solved the memory problem?", https://www.reddit.com/r/AI_Agents/comments/1r2puny/has_anyone_actually_solved_the_memory_problem_for/
[2] CrewAI, Memory, https://docs.crewai.com/concepts/memory
[3] LangGraph, Persistence, https://langchain-ai.github.io/langgraph/concepts/persistence/
[4] OpenAI Agents SDK, Handoffs, https://openai.github.io/openai-agents-python/handoffs/
[5] Cognition, Don't Build Multi-Agents, https://cognition.ai/blog/dont-build-multi-agents
[6] Anthropic, How we built our multi-agent research system, https://www.anthropic.com/engineering/built-multi-agent-research-system
[7] The Register, Vibe coding service Replit deleted production database (21/07/2025), https://www.theregister.com/software/2025/07/21/vibe-coding-service-replit-deleted-production-database/719783
[8] Chroma Research, Context Rot (Jul 2025), https://research.trychroma.com/context-rot
[9] SWE-Marathon (arXiv 2606.07682), https://arxiv.org/abs/2606.07682
[10] Hacker News, Agentic Coding Is a Trap, https://news.ycombinator.com/item?id=48002442
[11] Zep, Lies, Damn Lies, Statistics: Is Mem0 Really SOTA in Agent Memory?, https://blog.getzep.com/lies-damn-lies-statistics-is-mem0-really-sota-in-agent-memory/
[12] GitHub, Revisiting Zep's 84% LoCoMo claim (corrected to 58.44%), https://github.com/getzep/zep-papers/issues/5
[13] Hacker News, AI Startup Caught Cheating on Benchmark Papers, https://news.ycombinator.com/item?id=44883133
[14] GitHub, mem0 issue #4573, "97.8% were junk", https://github.com/mem0ai/mem0/issues/4573
[15] r/AI_Agents, 8 memory systems, 2,176 tasks, independent measurement, https://www.reddit.com/r/AI_Agents/comments/1veeix3/i_ran_8_ai_agent_memory_systems_through_2176/
[16] ProMCP: Profiling Token Flows and Latency Costs in MCP (ACL Findings 2026), https://aclanthology.org/2026.findings-acl.1967.pdf
[17] Anthropic, Code execution with MCP, https://www.anthropic.com/engineering/code-execution-with-mcp
[18] Mem0, State of AI Agent Memory 2026, https://mem0.ai/blog/state-of-ai-agent-memory-2026
[19] Zep, A Temporal Knowledge Graph Architecture for Agent Memory (arXiv 2501.13956), https://arxiv.org/abs/2501.13956
[20] Mem0, Building Production-Ready AI Agents (arXiv 2504.19413), https://arxiv.org/abs/2504.19413
[21] MemOS, A Memory OS for AI System (arXiv 2507.03724), https://arxiv.org/abs/2507.03724
[22] MemGPT, Towards LLMs as Operating Systems (arXiv 2310.08560), https://arxiv.org/abs/2310.08560
[23] Letta, Sleep-time Compute, https://www.letta.com/blog/sleep-time-compute/
[24] Letta, Memory Blocks, https://www.letta.com/blog/memory-blocks/
[25] HippoRAG (arXiv 2405.14831), https://arxiv.org/abs/2405.14831
[26] A-Mem: Agentic Memory (arXiv 2502.12110), https://arxiv.org/abs/2502.12110
[27] A Survey of Persistent Memory, State, and Governance in LLM Agents (Jun 2026), https://arxiv.org/html/2606.30306v1
[28] Tamper-Evident, Replayable Evidence for Autonomous AI Agent Runs (Sep 2026), https://arxiv.org/html/2609.12582v1
[29] Deletion Isn't Enough: Auditing RAG for Selective Forgetting (SIGIR 2026), https://marksanderson.org/files/papers/SIGIR2026_Leila_Main__Copy_.pdf
[30] AI Meets Brain: Memory Systems from Cognitive Science (arXiv 2512.23343), https://arxiv.org/html/2512.23343v1
[31] BEAM, Why BEAM Is a Good Memory Benchmark (Mem0), https://mem0.ai/blog/why-beam-is-a-good-memory-benchmark-for-ai-agents
[32] Agent Memory Benchmark (AMB), https://agentmemorybenchmark.ai/
[33] LoCoMo, Evaluating Very Long-Term Conversational Memory, https://snap-research.github.io/locomo/
[34] LongMemEval, https://xiaowu0162.github.io/long-mem-eval/
[35] Mastra, Observational Memory: 95% on LongMemEval, https://mastra.ai/research/observational-memory
[36] Zep, Announcing Community Edition, https://blog.getzep.com/announcing-zep-community-edition/
[37] Zep, A New Direction for Zep's Open Source Strategy, https://blog.getzep.com/announcing-a-new-direction-for-zeps-open-source-strategy/
[38] Graphiti, https://www.getzep.com/platform/graphiti/
[39] MongoDB, SSPL (2018), https://www.mongodb.com/company/newsroom/press-releases/mongodb-issues-new-server-side-public-license-for-mongodb-community-server
[40] Redis, Dual Source-Available Licensing (2024), https://redis.io/blog/redis-adopts-dual-source-available-licensing/
[41] Elastic, Elasticsearch Is Open Source. Again! (2024), https://www.elastic.co/blog/elasticsearch-is-open-source-again
[42] MongoDB, Enterprise Advanced Support, https://www.mongodb.com/services/support/enterprise-advanced-support-plans
[43] CockroachDB, Upgrade Policy, https://www.cockroachlabs.com/docs/cockroachcloud/upgrade-policy
[44] OpenAI, Memory and new controls for ChatGPT, https://openai.com/index/memory-and-new-controls-for-chatgpt/
[45] Anthropic, Memory, https://www.anthropic.com/news/memory
[46] AWS, Bedrock Agents memory, https://aws.amazon.com/blogs/machine-learning/amazon-bedrock-agents-now-supports-memory/
[47] Cognee, US$ 7.5M seed, https://www.cognee.ai/cognee-raises-seven-million-five-hundred-thousand-dollars-seed
[48] Supermemory, US$ 2.6M (Dataconomy, Oct 2025), https://dataconomy.com/2025/10/07/young-founders-supermemory-raises-2-6m-from-cloudflare-and-google-execs/
[49] ANPD, Transferência Internacional de Dados, https://www.gov.br/anpd/pt-br/assuntos/assuntos-internacionais/transferencia-internacional-de-dados
[50] Mayer Brown, End of grace period for Resolution CD/ANPD 19/2024, https://www.mayerbrown.com/pt/insights/publications/2025/08/end-of-grace-period-implementation-of-brazils-standard-contractual-clauses-in-international-transfers-of-personal-data
[51] CMN 4.893/2021, https://www.ancord.org.br/wp-content/uploads/2021/03/Resolucao-CMN-n-4.893-de-26_2_2021.pdf
[52] PL 2338/2023, Câmara dos Deputados, https://www.camara.leg.br/proposicoesWeb/fichadetramitacao?idProposicao=2487262
[53] European Commission, AI Omnibus enters into force, https://digital-strategy.ec.europa.eu/en/news/ai-omnibus-enters-force
[54] Cross Border Data Forum, Generative AI and GDPR Enforcement in Europe, https://www.crossborderdataforum.org/generative-ai-and-gdpr-enforcement-in-europe-a-lot-of-noise-one-fine-zero-survivors/
[55] Exploring Privacy Issues in RAG (ACL Findings 2024), https://aclanthology.org/2024.findings-acl.267.pdf
[56] Exposing Privacy Risks in Graph RAG (arXiv 2508.17222), https://arxiv.org/pdf/2508.17222
[57] AutoGen, AgentChat Memory (official docs), https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/memory.html
[58] AG2 (formerly AutoGen), GitHub, https://github.com/ag2ai/ag2
[59] MetaGPT: Meta Programming for a Multi-Agent Framework (arXiv 2308.00352), https://arxiv.org/abs/2308.00352
[60] Model Context Protocol, Specification 2025-11-25, https://modelcontextprotocol.io/specification/2025-11-25
[61] A2A Protocol (Linux Foundation), https://a2a-protocol.org/latest/

Next step

Less starting over. More continuity

Bring a real project. In one conversation we map where context gets lost between the people and agents on your team, and where to start.