Google Cloud 讲解在 AlloyDB 与 Memorystore for Valkey 中实现 AI Agent 长期记忆
How to implement long-term AI agent memory in AlloyDB and Memorystore for Valkey
Google Cloud 发布教程,讲解如何用 Memorystore for Valkey 做短期会话缓存、AlloyDB AI 做长期持久记忆的两层记忆架构。
Enterprise AI agents need persistent memory to execute complex, multi-day workflows and long-horizon tasks. In this blog, we examine how a 2-tier memory architecture using Memorystore for Valkey for short-term buffer memory and AlloyDB AI for long-term persistent memory can help reduce token spend by up to 70%, while maintaining critical data and enterprise guardrails.
Imagine building a personalized travel agent designed to help users book vacations. The user interacts with the agent many times over the course of several days, asking questions that range from brainstorming itineraries to actual purchase intent. To provide a truly seamless experience, this agent must remember flight preferences (e.g. “I only want non-stop flights”), hotel budgets, and dietary restrictions (e.g. “I need Gluten Free dining options”) established in previous sessions. More importantly, it has to hold onto these core facts even when the conversation gets deep into the weeds of sightseeing recommendations and itinerary planning. The agent must ensure that all the follow-up questions and exciting details about places to visit doesn’t cause it to forget or overwrite the user's fundamental requirements and decisions.
Stateful workflows inherently conflict with LLM’s stateless nature
AI agents are being used for multi-turn conversations and long-running workflows, but large language models (LLMs) remain stateless across sessions. When a user returns to an agent days later, the model starts with an empty context window. Unless your application is built to reconstruct past context using long-term agent memory, your users have to explain their goals and context all over again, resulting in a frustrating and fragmented experience.
With million-token context windows now the norm, a common shortcut to this problem is "context stuffing" - dumping everything you can fit into the prompt at every turn, from raw chat histories to tool execution logs. In fact, this was a common pattern in the early days of AI model usage. But this shortcut quickly creates issues at scale: token costs multiply with every message, response times can drag out past 30+ seconds for otherwise simple prompts, and the model starts suffering from "lost in the middle" degradation, overlooking critical instructions buried in mountains of prompt text.
Another common workaround is to use rolling summaries; asking an LLM to periodically compress older messages into a summary paragraph. While this trims prompt size, LLM-based summarization is inherently lossy. After a few rounds of compression, subtle but important details get filtered out as background noise. A few turns later, your agent quietly breaks the exact constraints you set earlier. Clearly, you need a more scalable approach for keeping a memory from previous conversations or multi-turn tasks, but without degrading the experience or creating new bottlenecks.
Implementing a 2-tier memory architecture
To build reliable, cost-effective enterprise agents that respect your guardrails and constraints, we recommend a 2-tier memory architecture:
-
Short-term session buffer: Caches active conversation turns in memory using a token-bounded sliding window, allowing the context window size to remain stable across multiple rounds, even across devices, while keeping the latest messages fresh. This tier requires sub-millisecond, high-throughput lookups on every turn, making Memorystore for Valkey well suited for maintaining active session state.
-
Long-term persistent memory: Stores important facts, user preferences, and episodic facts across sessions. This tier requires transactional integrity, data governance, and hybrid retrieval across relational data and vectors - capabilities provided natively by AlloyDB AI.
While it’s possible to store both long-term memory and active session buffers directly in a relational database, a two-tier architecture with an in-memory cache provides better performance and scalability. Active session buffers are highly ephemeral and require sub-millisecond, high-throughput updates on every single conversation turn. Handling these rapid-fire writes in an in-memory key-value cache prevents write amplification and table bloat in your relational database, which would otherwise require frequent row deletions and intensive vacuuming. This is analogous to adding a caching tier in front of your database to offload high-frequency lookups for hot rows. The division of labor keeps your primary database lean and responsive, allowing it to focus its resources on what it is designed for: transactional consistency, complex hybrid vector search, and long-term analytical query execution.
By pairing short-term caching in Memorystore for Valkey with native AlloyDB AI capabilities, you can run entity extraction, memory compaction, cross-session memory, and hybrid retrieval directly inside the database tier, keeping active prompts lean, fast, and cost-efficient.
Understanding the four memory types
To organize long-term agent state effectively, we further divide agent memory into four complementary types across the short-term and long-term storage tiers we described above:
|
Memory type |
What it stores |
Storage layer |
Lifespan |
|---|---|---|---|
|
Buffer (Short-term) |
Recent raw conversation turns |
Memorystore for Valkey |
Active session |
|
Summary memory |
Compressed history of older turns, commonly referred to as “compaction” |
Memorystore for Valkey |
Multi-turn window |
|
Episodic memory |
Past actions, events, and tool outputs |
AlloyDB for PostgreSQL (Hybrid retrieval with structured SQL + full-text search + vector) |
Permanent |
|
Entity & rule memory |
User preferences, constraints, and vetoes |
AlloyDB for PostgreSQL (Hybrid retrieval with structured SQL + full-text search + vector) |
Permanent |
By isolating short-term conversation context from structured long-term rules, your agent retrieves relevant context on demand without filling token windows with raw interaction logs.
The tiered memory architectural blueprint
The diagram below outlines the read and write paths connecting the application orchestration layer, the Memorystore for Valkey short-term buffer, and the AlloyDB AI long-term repository:

The system operates across two coordinated execution paths:
-
The read path: When a user asks a question, the application fetches the active sliding window from Memorystore for Valkey, runs in-database query normalization using
ai.generate, and queries AlloyDB usingai.hybrid_searchto retrieve scoped entity rules and relevant episodic facts. -
The write path: After generating the response, the turn is immediately cached in Memorystore for Valkey. An asynchronous background queue worker extracts structured entities from the exchange and writes them directly into AlloyDB, where transactional auto-embeddings immediately compute and store vector representations in the database.
Business impact and ROI
In our benchmark testing across multi-turn development dialogues (45+ turns with heavy tool executions), separating short-term caching from long-term persistence delivered measurable cost and performance improvements compared to naive context stuffing:
|
Metric / dimension |
Naive context stuffing |
Tiered memory (AlloyDB + Valkey) |
Net business impact |
|---|---|---|---|
|
Active prompt size (turn 45) |
747,033 tokens |
83,262 tokens |
88.9% smaller prompt |
|
Turn 45 response latency |
33.5 seconds |
6.7 seconds |
80.0% faster response |
|
Per-turn response wait time |
33.5 seconds |
4.2s – 6.7s |
36% to 80% reduction |
|
Cumulative session tokens |
17.9M tokens |
4.09M tokens |
72.0% token & cost savings |
|
Rule & constraint recall |
Degrades over turns |
Does not degrade |
ACID-preserved recall |
Note: The metrics above reflect internal benchmark results from a simulated developer workload that mimics a real-world enterprise AI pair-programming assistant interacting with a developer over multiple sessions, projects, and context switches. Actual savings and latencies vary based on prompt structure, query frequency, and data volume.
These results demonstrate how tiered memory changes agent unit economics: instead of an escalating cost curve on every additional turn, prompt sizes remain bounded, reducing ongoing LLM API expenses while keeping response times fast.
Core AlloyDB AI technical advantages
AlloyDB AI reduces the operational overhead of implementing persistent agent memory by embedding core AI functions directly into the database engine:
-
Transactional in-database auto-embeddings (ai.initialize_embeddings): AlloyDB automatically generates vector embeddings for text columns using a native integration with Agent Platform (formerly Vertex AI), generating up to 3,000 embeddings per second. Using
incremental_refresh_mode => 'transactional', AlloyDB keeps embeddings up to date as source data changes within the same transaction, removing the need for custom embedding pipelines, external schedulers, or complex retry logic. -
In-database generative AI functions (ai.generate): AlloyDB allows you to execute foundation models, such as Gemini, directly from SQL queries. You can use this for in-database query decomposition - breaking compound user questions into single-aspect sub-queries and resolving relative time phrases (like "last session") into explicit identifiers - without making separate roundtrips from your application.
-
Native hybrid search with built-in Reciprocal Rank Fusion (ai.hybrid_search): AlloyDB provides a built-in SQL function that executes Reciprocal Rank Fusion (RRF) directly inside the engine. It combines vector cosine similarity (
<=>over HNSW or ScaNN indexes) with PostgreSQL full-text search (using BM25, RUM, or GIN) in a single database call, blending semantic matching with exact keyword retrieval while supporting metadata filter pushdown (filter_condition) for improved performance and deterministic scope isolation. -
Direct Agent Platform integration with IAM credentials: AlloyDB connects directly to Agent Platform foundation models over Google Cloud's private network using Google Cloud IAM service account roles and database authentication, avoiding the need to store, rotate, or pass API keys in application code.
-
Unified operational, vector, and governance engine: AlloyDB consolidates relational business data, vector embeddings, full-text indexes, and enterprise permissions in a single ACID-compliant PostgreSQL database, avoiding data drift and integration complexity across separate operational and vector databases.
Key implementation patterns
Below are the core database patterns used to configure the 2-tier memory architecture. For complete, runnable Python and SQL scripts, refer to the companion AlloyDB Agent Memory Codelab.
1. Setting up schema and auto-embeddings
In AlloyDB, install the necessary extensions and define the agent_entities table with structured metadata, a generated tsvector column for full-text search, and a vector embedding column.
Then, generate embeddings using ai.initialize_embeddings. Using transactional mode ensures the embeddings are kept up to date as source data changes.
Finally, create the HNSW vector index and the RUM full-text search index to ensure your hybrid searches are fast and efficient.
2. Querying long-term memory with native hybrid search
On the read path, retrieve relevant long-term entities using ai.hybrid_search. This native SQL function executes Reciprocal Rank Fusion (RRF) directly in AlloyDB, seamlessly reranking and combining vector similarity search and full-text keyword search results in a single database query:
3. In-database memory compaction with AI functions
To manage long-term storage growth without writing custom pruning scripts, you can run automated extraction and compaction queries directly in AlloyDB to identify the important parts which can benefit from being stored. This pattern uses a SQL common table expression (CTE) with ai.generate to consolidate older episodic entries into a high-density summary:
For example, a long interaction regarding the complexities of changing flights with kids can result in a summary of “I prefer non-stop flights”.
4. Connecting the 2-tier memory architecture to an ADK agent
With the core 2-tier memory architecture configured, you can now extend the default ADK Memory provider to use this 2-tier memory architecture as shown in the accompanying Codelab (e.g. ADKTieredMemoryProvider). To attach the long-term memory to an ADK agent, you simply provide it as a tool (e.g. longterm_memory_tool) like this:
Enterprise governance and multi-tenant security
Running agent memory in enterprise production environments requires strict security boundaries and access controls. First of all, we need to maintain scope and multi-tenant isolation. By indexing user_id, project_id, and scope columns in agent_entities and enforcing PostgreSQL Row-Level Security (RLS), you can isolate memory stores across departments, teams, and individual users within the same database cluster. Parameterized Secure Views (PSV) offer another layer of deterministic application-level security, helping you protect against malicious prompts and overly-broad SQL queries.
In addition, we need to provide automated memory lifecycle management as shown in the implementation pattern above. Combining scheduled SQL compaction queries with time-based partition pruning helps maintain predictable database footprint and query latencies over time. See the accompanying Codelab for more details on this approach.
Summary and next steps
Decoupling active context windows from persistent storage is a practical approach to building production-ready AI agents. By pairing Memorystore for Valkey for sub-millisecond session caching with AlloyDB AI for transactional long-term storage, you can achieve substantial token cost savings and faster response times while maintaining strict business rules throughout long-horizon tasks and many-turn agentic experiences.
To get started:
-
Step through the complete hands-on tutorial in the companion AlloyDB Agent Memory Codelab to deploy the working 2-tier memory architecture.
-
Learn more about database-side machine learning features in the AlloyDB AI documentation.
-
Explore guides on generating auto vector embeddings and running hybrid vector search.
Posted in
来源:Google Cloud:Databases · cloud.google.com