Building Persistent Agent Memory with NVIDIA NeMo Agent Toolkit and Amazon S3 Vectors
- Authors

- Name
- Nino
- Occupation
- Senior Tech Editor
Autonomous LLM agents are rapidly evolving from simple conversational bots to complex multi-step reasoning engines. However, statefulness remains one of the primary bottlenecks in production environments. Standard context windows quickly become saturated, expensive, and latency-heavy when agents perform long-running orchestration, iterative search, or financial report analysis.
To build production-grade autonomous workflows, developers must provide agents with a high-performance, cost-effective persistent memory layer. By pairing the NVIDIA NeMo Agent Toolkit (NAT) with Amazon S3 Vectors deployed on Amazon Elastic Kubernetes Service (Amazon EKS), you can create a enterprise-ready memory infrastructure capable of storing and retrieving context across millions of interactions.
When powering complex agent topologies, reliable model access is critical. Developer teams frequently route their multi-agent LLM requests through unified inference platforms like n1n.ai to maintain high uptime, optimize latency, and lower API cost overhead across models such as DeepSeek-V3, Claude 3.5 Sonnet, and GPT-4o.
In this comprehensive guide, we examine the inner architecture of the NVIDIA NeMo Agent Toolkit memory subsystem, demonstrate how to implement Amazon S3 Vectors as a custom memory provider, and deploy a real-world multi-agent investment research workflow on Amazon EKS.
Deep Dive into the NVIDIA NeMo Agent Toolkit (NAT) Memory Subsystem
The NVIDIA NeMo Agent Toolkit (NAT) provides an open, highly modular framework designed to construct, orchestrate, and deploy high-throughput multi-agent systems. Unlike basic chain execution frameworks, NAT separates agent logic into discrete control loops: Perception, Planning, Memory, and Action Execution.
NAT models memory across three primary abstractions:
- Episodic Memory: Stores immediate execution traces, tool output histories, and short-term dialogue dynamics within the current execution session.
- Semantic Memory: Maintains long-term conceptual knowledge, domain-specific documents, financial reports, or code documentation indexed via vector embeddings.
- Procedural Memory: Manages agent execution workflows, system instructions, and tool definitions for dynamic policy enforcement.
To decouple vector storage hardware from agent code, NAT exposes a standard interface class BaseMemoryProvider. Custom memory backends must implement state insertion, vector search similarity queries, and lifecycle management functions.
+-------------------------------------------------------------------------+
| NVIDIA NeMo Agent Toolkit |
| |
| +-------------------+ +------------------+ +-------------------+ |
| | Perception Module | | Reasoning Engine | | Action Execution| |
| +---------+---------+ +--------+---------+ +---------+---------+ |
| | | | |
| +----------------------+-----------------------+ |
| | |
| v |
| +---------------------------+ |
| | NAT Memory Subsystem | |
| | (BaseMemoryProvider API) | |
| +-------------+-------------+ |
+------------------------------------|------------------------------------+
| Abstract Vector Query / Store
v
+-------------------------------------------------------------------------+
| Amazon S3 Vectors (EKS Hosted) |
| |
| +-------------------+ +------------------+ +-------------------+ |
| | S3 Object Storage | <--> | Vector Index Driver| <-->| EKS Pod Workers | |
| +-------------------+ +------------------+ +-------------------+ |
+-------------------------------------------------------------------------+
Why Amazon S3 Vectors for Agent Memory?
Traditional vector databases often impose significant cost trade-offs at scale. Maintaining standard in-memory vector indices (e.g., HNSW in RAM) for millions of long-term agent interaction histories becomes cost-prohibitive. Amazon S3 Vectors provides a cloud-native architecture separating compute from vector storage:
- Cost Efficiency: High-density vector embeddings reside in durable object storage, reducing raw infrastructure memory costs by up to 70% compared to dedicated RAM clusters.
- Infinite Horizontal Scale: S3 handles petabyte-scale storage, while indexing compute scales independently inside Amazon EKS pods.
- Predictable Latency: Optimized disk-backed indexing strategies ensure query retrieval latency remains
< 35msfor top-k vector similarity searches.
Building the Custom S3 Vector Memory Provider
To integrate Amazon S3 Vectors into NAT, we implement a custom Python subclass inheriting from BaseMemoryProvider. This class translates NAT memory payload objects into embedding vectors and commits both the vector indices and original unstructured payloads to Amazon S3.
Below is the concrete code implementation for the S3VectorMemoryProvider:
import uuid
from typing import List, Dict, Any, Optional
import boto3
import numpy as np
from nemora.memory.base import BaseMemoryProvider, MemoryRecord
class S3VectorMemoryProvider(BaseMemoryProvider):
def __init__(
self,
bucket_name: str,
index_prefix: str,
embedding_dim: int = 1536,
aws_region: str = "us-west-2"
):
super().__init__()
self.bucket_name = bucket_name
self.index_prefix = index_prefix
self.embedding_dim = embedding_dim
self.s3_client = boto3.client("s3