Personal Agents and Memory | Liao Jiayi

Personal Agents and Memory

Liao Jiayi Liao Jiayi #AI#LLM#Agent

A survey of Agent Memory techniques and industry practice—from external memory and in-model state to parameter updates—and how memory can make agents intelligent and personalized.

Translated from Chinese with AI · Read the original

I previously wrote Continual Learning: The Next Frontier of Intelligence?. The question I cared about then was whether, beyond pretraining and post-training, models could continually absorb feedback and acquire new knowledge and capabilities as people use them.

Introduction

This article explores several approaches to Agent Memory, including external memory, internalizing memory within the model, continual parameter learning, and hybrid architectures. I want both to understand the latest frontier and to assess whether this research can realistically reach production.

From a user’s perspective, the intelligence expected of a Personal Agent is straightforward: we want an agent that “understands you” as a person would. Memory is central because most information in daily life does not come from an encyclopedia such as Wikipedia; it comes from immediate feedback about events as they happen. In that sense, it resembles how people browse Douyin and TikTok.

This has made me wonder whether recommender systems and Personal Agents have something in common. Memory in a recommender system is essentially a collection of features, which is why such systems often have a Feature Store. Those features become important inputs to every inference.

The industry currently takes three broad approaches to Memory:

  1. External memory: preserve history in a file system as individual Markdown documents;
  2. Model memory: as in TTT (Test-Time Training), use every input/output interaction as a hook for updating memory;
  3. Hybrid approaches: external memory (offline) plus model memory (real time).

We will first examine these approaches, then return to the recommender-system analogy.

External Memory

The basic external-memory workflow is fairly direct:

image

File-Based Memory

Claude Code Memory

Claude Code’s official memory documentation clearly separates two mechanisms: users write persistent instructions, while Claude records lessons from corrections, preferences, and work.

  • Instructions can live in CLAUDE.md files at several scopes, including managed policy, user, project, and local.
  • Auto-memory is stored in a project-specific memory directory with MEMORY.md as its entry point. At startup, the first 200 lines or 25 KB are loaded, while other topic files are read on demand.
  • The file hierarchy defines scope and loading relationships; it does not express a concept of “compression.”

Dive into Claude Code: The Design Space of Today’s and Future AI Agent Systems gives a detailed interpretation of Agent Memory in Claude Code. It describes context management rather than fine-tuning the model with memory—in essence, an external model harness.

Files also create substantial problems. For example:

  1. Internal inconsistency: deepseek-harness, a recently popular project built entirely through AI-assisted coding, makes extensive use of documentation. Many of its issues already concern errors in .md files. If it keeps evolving this way, I suspect the project will eventually become a maintenance nightmare—even though it comes from DeepSeek.
  2. Poor indexing: compared with embedding-based recall, document retrieval can be implemented in many ways, but it is difficult to guarantee that the system will retrieve what the user actually wants every time. As an experiment, ask Claude Code about a task you worked on a month ago and see whether it can still describe it accurately.

Everything is Context / AFS

Everything is Context / AFS systematically organizes memory into a structured file-system hierarchy:

Area Purpose Lifecycle
/context/history Stores historical events and interactions Immutable, append-only
/context/memory/agentID Stores processed experience, summaries, and indexes Updatable and persistent
/context/pad/taskID Temporary workspace for the current task Short-lived and auditable

This separation resembles logs, derived views, and temporary state in engineering systems. Retaining raw history makes it possible to regenerate a faulty summary, while intermediate hypotheses from the current task do not have to enter long-term memory immediately.

My view: the immediate value of this approach is engineering-oriented. It is inspectable and reconstructable, and it distinguishes temporary information from stable information—all genuine requirements for long-running systems. Whether this directory structure improves task success rates still requires separate experiments.

Vector, Hierarchical, and Graph Memory

Mem0: Vector Retrieval

Mem0: Building Production-Ready AI A=gents with Scalable Long-Term Memory

Background/problem: in long-running conversations, useful information is often scattered across many separate sessions.

  1. Putting the entire history into context increases cost and latency;
  2. Chunking and retrieving raw data can introduce irrelevant content, duplicate facts, and outdated information.

Approach (illustrated below):

  1. Extract facts: combine the new conversation turn, recent messages, and historical summaries so an LLM can identify information worth retaining.
  2. Find existing memories: use embeddings to retrieve semantically similar stored facts.
  3. Decide how to update: have the LLM compare new and old information and select ADD, UPDATE, DELETE, or NOOP.
  4. Retrieve on demand: when a new question arrives, fetch relevant memories and place their text in the model context.

image

My view: this is a simple and highly practical Memory design, and it is already widely used.

MemoryOS: Hierarchical Memory

Memory OS of AI Agent

Background/problem: the most recent conversation turns need detailed retention, older discussions need topic-based retrieval, and user preferences must accumulate over time. Applying the same storage and access policy to all three can mix topics, crowd out important information, and produce an inconsistent user profile.

Approach: a three-tier architecture

Tier Stored content Access method
Short-term memory (STM) Recent conversations and their contextual relationships Included in the current context as a whole
Mid-term memory (MTM) Topic-organized conversation pages and summaries Retrieve the topic first, then the specific conversation
Long-term personal memory (LPM) User and agent profiles, preferences, and facts Use profiles directly; retrieve facts by relevance

image

Graph

Zep: A Temporal Knowledge Graph Architecture for Agent Memory

Background: the most common form of RAG for LLMs is designed mainly for static information and cannot accurately represent relationships that change frequently. For example, “the user once worked at Company A” and “the user now works at Company B” can both be true.

Approach: Graphiti organizes memory into three components:

  • Episode Graph: an original conversation episode -> an entity mentioned in that conversation
  • Entity Graph: Entity -> Entity
  • Community Graph: a community formed by a group of entities

The most interesting part of this approach is how the graph is constructed:

  1. An LLM extracts entities from the raw input. For example, “Xiao Wang works at Company A” contains the two entities “Xiao Wang” and “Company A.”
  2. Large numbers of entity connections can be clustered with graph algorithms.
  3. During retrieval, the system can start from an entity embedding and return episodes to a specified depth.

Abstracting Memory from Trajectories

From Storage to Experience

Background: existing LLM Memory designs are fragmented and do not evolve over time. Approach: a three-layer Memory design

  1. Storage: preserve raw trajectories
  2. Reflection: reflect on the raw trajectories
  3. Experience: extract generalizable experience from those reflections

image

LongMemEval-V2

Background: remembering user preferences and becoming familiar with a specific work system are different capabilities. An Agent that uses business applications over a long period must also learn page structures, the consequences of actions, task workflows, and common pitfalls. Approach: this design resembles a combination of evaluation and memory. Rather than storing isolated facts, it records the full chain of causes and consequences in an agent trajectory.

image

Sleep-Time Compute / Dreaming: Predicting User Needs

Sleep-time Compute

Published in April 2025.

Background: scaling test-time computation can improve model reasoning, but it increases latency and cost. Approach: this paper proposes predicting user needs offline and precomputing the corresponding chains of thought, somewhat like offline test-time scaling. How are user needs predicted?

  1. Use context as a clue
  2. Prompt the model to generate tasks
  3. Infer, then store, the results

image

My view: this is well suited to some fixed scenarios. In common data-analysis tasks, for example, the available analytical methods and aggregation dimensions are limited, so useful computations can be prepared in advance using business characteristics and a finite set of analytical methods. This may be a promising direction.

Letta’s Current Memory & Dreaming Documentation

Letta is the company commercializing MemGPT and focuses on Continual Learning and memory-enabled Agents. Its methodology resembles Sleep-Time Compute, except that its offline computation operates on a user’s past interaction trajectories rather than raw text. Sleep-Time Compute is more oriented toward knowledge questions and answers.

Karpathy’s 2025 Interview

Several key points:

  1. Pretraining provides knowledge absorption and question-answering ability, but what it acquires is fundamentally a form of “fuzzy memory.”
  2. Context offers more direct access to information and is more reliable than fuzzy memory.
  3. To turn experience into weights, he imagines a model reviewing interactions, analyzing them repeatedly, generating synthetic training data, and distilling the result back into its weights. Personalized updates might use localized parameters such as LoRA. The goal is for interactions to produce learning that outlives the current context.

Note: there is still no consensus solution.

Memory During Model Inference

TTT(Test-Time Training)

Training here means that the model can still be trained during testing or inference, rather than updating model weights only through the conventional backward-training stage.

The main idea combines properties of RNNs and Attention:

  • RNNs can compress memory by continually folding historical information into an RNN unit of a fixed dimension. This loses information over time, which is why mechanisms such as LSTM gates control forgetting, although those gates depend on human-designed model structure.
  • Attention performs an O(L^2) attention computation over all tokens. It avoids information loss but creates a computational-complexity problem for very long contexts.

TTT combines properties of both, with Linear and MLP variants:

  • TTT-Linear: without backward propagation, it places a small linear model over hidden states and updates that model’s weights in an RNN-like manner as new hidden states arrive. The small model’s weights become a compressed representation of all historical inputs—the hidden states.
  • TTT-MLP: similar to the Linear variant, but with nonlinear layers that can uncover interactions within hidden states; its model weights are updated through backward propagation.

Titans: Learning to Memorize at Test Time

Titans: Learning to Memorize at Test Time is Google work from January 2025.

It is fundamentally similar to TTT, but introduces more complex MLP and model-architecture designs, including momentum, adaptive forgetting, and information admission.

Nested Learning / HOPE: Memory at Different Frequencies

Nested Learning, published at NeurIPS in December 2025, is another Google Research project from the same group.

It adds layers operating at different time scales to the existing model architecture, much like short-window and long-window features in recommender systems. Short-window features change in real time, while long-window features can update on T+1 or T+7 schedules. The goal is to preserve reasoning over historical knowledge while allowing recent knowledge to be absorbed quickly.

  • Nested Learning: treats the model and its optimization process as interconnected learning problems at multiple levels, each with its own information flow and update frequency.
  • It combines a Titans variant capable of learning its own update mechanism with CMS (Continuum Memory System), a continuous multi-timescale memory system.

image

The Surprising Effectiveness of Test-Time Training for Few-Shot Learning

My view: published in 2025, the core idea is to synthesize a batch of data when a new situation appears and then train it into the model with SFT. This is not fundamentally different from fine-tuning as practiced today.

Macaron-V1: Evolving the Model and Harness Together

On the Scaling of PEFT δ-mem: Efficient Online Memory for Large Language Models Macaron-V1

I previously discussed these papers here: https://www.liaojiayi.com/blog/continual-learning#mindlab

Continually Training New Memory into a Model Is Not Lossless

Consider these three papers together:

  • Sequential Knowledge Editing
    • This September 2026 paper measures perplexity on Qwen2.5. Injecting either false information or genuine updates causes clear changes in perplexity even when the MMLU score remains unchanged.
  • AlphaEdit
    • An October 2024 exploration: AlphaEdit projects updates into the null space of the old-knowledge representation matrix. Its conclusion is that a matrix-mapping constraint can successfully ingest new knowledge, but it cannot guarantee completely lossless question answering; the outcome depends on the old-knowledge training data.
  • Model Editing at Scale
    • A successful single knowledge edit does not mean the same model can accept updates indefinitely. By the time it answers a new fact correctly, previously written facts and other capabilities may already have begun to degrade.

My view: many studies show that directly training an unfrozen pretrained model quickly damages its general capabilities. This suggests that current LLMs are poorly suited to real-time updates and continual learning. Even LoRA can create a knowledge discontinuity because it is not jointly trained with the backbone.

Add a Memory Layer Directly to the Model?

Continual Learning via Sparse Memory Finetuning

Inference: use the model’s 12th layer as a memory layer. The previous layer’s output becomes a query sent to multiple slots in this layer. The system selects the top-k slots by key similarity, retrieves their values, and projects them into the output. Each slot has two roles:

  • key: used to measure similarity to the query
  • value: used for output projection

Training: within each batch, apply TF-IDF to select the top-t slots for weight updates, so new knowledge is written only into a designated small subset of slots.

Hybrid Memory

EvolveR / EMPO²: Expanding Memory Data While Training the Model

  • EvolveR constructs a loop:
    1. Extract natural-language principles from successful and failed trajectories, deduplicate and merge them, and maintain their quality.
    2. On new tasks, the Agent independently retrieves principles or external knowledge to produce new trajectories.
    3. Update policy parameters with GRPO using rewards for answer correctness and behavioral formatting.
  • EMPO² observes that reflection stored in memory helps solve problems but also makes the model heavily dependent on that reflection. It therefore constructs three data combinations:
    • Sample without memory and update under the same conditions;
    • Sample with memory and retain the prompt during the update;
    • Sample with memory, remove the prompt, and perform an off-policy update.

The GPT-4o Sycophancy Incident: User Preference Does Not Necessarily Mean Better Behavior

OpenAI’s May 2, 2025 postmortem provides a concrete case of feedback optimization. It reviews the excessive agreeableness introduced by the April 2025 GPT-4o update. OpenAI concluded that user-feedback rewards combined with several training changes may have weakened mechanisms that constrained sycophancy, while evaluations and A/B tests failed to detect the issue in time. The update was ultimately rolled back. OpenAI’s later discussion of behavior in sensitive conversations separately addresses issues such as emotional reliance. For an assistant that remembers a user over the long term, these behaviors require longitudinal observation; a one-turn judgment of “which response the user prefers” may not represent long-term outcomes.

Recommender Systems and Personal Agents

Different Optimization Objectives

Recommender systems optimize interaction behavior such as CTR. Once content controls are in place, the objective is generally to maximize user engagement. A Personal Agent optimized for human preference, however, can easily reproduce incidents like GPT’s sycophancy failure. Personal Agents therefore need to focus more on whether they complete the user’s task. From this perspective, the underlying problem is still whether the Agent is capable enough to meet the need, not whether it is sufficiently personalized.

Long-Delayed Negative Feedback Resembles Agent Feedback

Sparse negative feedback is a major problem in recommendation. For advertising CVR, for example, conversion events are far rarer than clicks and may arrive after a long delay. Common remedies include (1) balancing positive and negative samples and (2) building a delayed negative-feedback loop. Agent systems face a similar problem: negative feedback often consists of the user leaving when a task is not completed. Such feedback is scarce, ambiguous, and delayed, suggesting that similar techniques might help with memory construction or sample collection.

Feature Engineering May Apply to Memory Weights

As noted above, sparse features in recommendation are often divided into short and long windows. Human memory is similar: recent memories are clearer and change more frequently. Current LLM retrieval methods such as RAG do not assign temporal weights, making this another direction worth investigating.