Google Cloud on October 2, 2026 published a technical guide describing how to give AI agents long-term memory by combining Memorystore for Valkey as a short-term session cache with AlloyDB AI as a persistent, transactional store. The post, written by AlloyDB AI engineering manager Itai Rosenblatt and product manager Paul Ramsey, includes SQL and Python patterns plus internal benchmark figures.
What it does
Google says large language models remain stateless across sessions, so an agent that a user returns to days later starts with an empty context window unless the application reconstructs past context.
The company describes two common workarounds and their drawbacks. "Context stuffing" multiplies token costs, can drag response times past 30+ seconds for otherwise simple prompts and causes "lost in the middle" degradation, according to the post, while rolling LLM summaries are described as inherently lossy.
Google's recommended alternative is a 2-tier memory architecture: a token-bounded sliding window of recent turns cached in Memorystore for Valkey, and important facts, user preferences and episodic facts stored in AlloyDB AI. The company further divides memory into four types — buffer, summary memory, episodic memory, and entity and rule memory — with the first two in Valkey and the last two in AlloyDB for PostgreSQL.
Before you start
The guide assumes AlloyDB with the google_ml_integration, vector and rum extensions installed, plus Memorystore for Valkey for session caching. Google says AlloyDB connects to Agent Platform (formerly Vertex AI) foundation models over Google Cloud's private network using IAM service account roles and database authentication, avoiding the need to store, rotate or pass API keys in application code.
Google notes that complete, runnable Python and SQL scripts live in a companion AlloyDB Agent Memory Codelab, so the blog's excerpts are partial rather than a full implementation.
Steps as Google describes them
Step 1: Set up schema and auto-embeddings. Google says to enable the AI extensions, create an agent_entities table with structured metadata, a generated tsvector column for full-text search and a VECTOR(768) embedding column, then call ai.initialize_embeddings with model_id 'text-embedding-005', incremental_refresh_mode => 'transactional' and batch_size => 10. The guide then creates an HNSW vector index and a RUM full-text index.
Step 2: Query long-term memory with native hybrid search. On the read path, Google says to retrieve relevant long-term entities using ai.hybrid_search, a SQL function that executes Reciprocal Rank Fusion directly in AlloyDB, combining vector similarity and full-text keyword results in a single query. The sample weights the text input at 0.6 and the vector input at 0.4 and applies a filter_condition on user_id, project_id and scope.
Step 3: Compact memory in the database. The post shows a common table expression that selects older events and passes them to ai.generate with model_id => 'gemini-3.5-flash' to produce a dense summary, which is then inserted into agent_entities with an ON CONFLICT update.
Step 4: Attach the memory to an agent. Google says you can extend the default ADK Memory provider (for example, ADKTieredMemoryProvider) and pass long-term memory to an ADK Agent as a tool.
For governance, Google recommends indexing user_id, project_id and scope, enforcing PostgreSQL Row-Level Security, and using Parameterized Secure Views as an additional application-level layer against malicious prompts and overly broad SQL queries.
What the company says about results
Google reported internal benchmark results from a simulated developer workload of 45+ turns with heavy tool executions. It said the active prompt at turn 45 fell from 747,033 tokens with naive context stuffing to 83,262 tokens, an 88.9% smaller prompt, and that turn 45 latency dropped from 33.5 seconds to 6.7 seconds.
Cumulative session tokens fell from 17.9M to 4.09M, which the company described as 72.0% token and cost savings. Google cautioned that actual savings and latencies vary based on prompt structure, query frequency and data volume. These are the company's own figures and have not been independently verified.
What to do
- Follow the complete hands-on tutorial in the companion AlloyDB Agent Memory Codelab to deploy the working 2-tier memory architecture, as Google recommends.
- Read the AlloyDB AI documentation for database-side machine learning features.
- Review Google's guides on generating auto vector embeddings and running hybrid vector search.
- Google says new Google Cloud customers get $300 in free credits and that AlloyDB offers a 30-day free trial instance for testing your own workload.
Key facts and where they come from
- Google recommends a two-tier split between Valkey and AlloyDB AI, claiming up to 70% lower token spend.
a 2-tier memory architecture using Memorystore for Valkey for short-term buffer memory and AlloyDB AI for long-term persistent memory can help reduce token spend by up to 70%
- Benchmark prompt size at turn 45 dropped from 747,033 to 83,262 tokens.
Active prompt size (turn 45)
747,033 tokens
83,262 tokens
88.9% smaller prompt
- Benchmark latency at turn 45 fell from 33.5 seconds to 6.7 seconds.
Turn 45 response latency
33.5 seconds
6.7 seconds
80.0% faster response
- AlloyDB generates embeddings via Agent Platform at a stated rate of up to 3,000 per second.
AlloyDB automatically generates vector embeddings for text columns using a native integration with Agent Platform (formerly Vertex AI), generating up to 3,000 embeddings per second.
- ai.hybrid_search runs Reciprocal Rank Fusion inside the database engine.
AlloyDB provides a built-in SQL function that executes Reciprocal Rank Fusion (RRF) directly inside the engine.
- Google suggests Row-Level Security and Parameterized Secure Views for multi-tenant isolation.
enforcing PostgreSQL Row-Level Security (RLS), you can isolate memory stores across departments, teams, and individual users within the same database cluster
- The benchmark used a simulated developer workload, and Google says results vary.
Actual savings and latencies vary based on prompt structure, query frequency, and data volume.
