Building Persistent Agent Memory with NVIDIA NeMo Agent Toolkit and Amazon S3 Vectors
- Authors

- Name
- Nino
- Occupation
- Senior Tech Editor
As autonomous AI agents evolve from simple prompt-response chains into sophisticated, long-running decision engines, state management becomes a primary engineering challenge. Standard context windows, even when expanded to millions of tokens, are insufficient for enterprise workflows that demand durable, multi-session memory, high retrieval throughput, and fine-grained access controls.
To achieve continuous context awareness without ballooning inference costs, production systems rely on decoupled memory architectures. The NVIDIA NeMo Agent Toolkit (NAT) provides a flexible framework for building, orchestrating, and scaling agentic microservices. When paired with Amazon S3 Vectors—which leverages object storage for ultra-scalable, cost-effective vector search—and deployed on Amazon Elastic Kubernetes Service (Amazon EKS), developers can construct a enterprise-grade memory subsystem.
In this comprehensive technical guide, we will explore the architectural foundation of NAT's memory interface, build a production-ready custom S3 Vectors memory provider, deploy it on EKS, and walk through a complex multi-agent financial research use case. High-performance agent fleets require low-latency LLM routing across multiple foundational models; developers can streamline their underlying model access by leveraging aggregators like n1n.ai for reliable API delivery.
Technical Architecture: Deconstructing NAT Memory Subsystem
The NVIDIA NeMo Agent Toolkit structures agent execution across distinct layers: task orchestration, tool interaction, and state persistence. The memory subsystem acts as the semantic backbone, allowing agents to write operational histories, index unstructured background data, and retrieve historical decisions across sessions.
+-----------------------------------------------------------------------+
| Amazon EKS Cluster Microservices |
| |
| +---------------------+ +-------------------+ +---------------+ |
| | Market Ingest Agent | | Compliance Agent | | Synth Agent | |
| +----------+----------+ +---------+---------+ +-------+-------+ |
| | | | |
| +------------------------+---------------------+ |
| | |
| [NVIDIA NeMo Agent Toolkit] |
| +-------------------------+ |
| | Memory Subsystem Abstr. | |
| +------------+------------+ |
| | |
+-------------------------------------|---------------------------------+
| AWS SDK / Boto3
v
+-------------------------------+
| Amazon S3 Vectors |
| +-------------------------+ |
| | Vector Index & Metadata | |
| +-------------------------+ |
+-------------------------------+
NAT categorizes agent memory into three foundational tiers:
- Episodic Memory: Captures turn-by-turn conversation dynamics and intermediate execution logs for the duration of a single transaction.
- Semantic Memory: Stores domain knowledge, ingested documents, and embeddings accessible via retrieval-augmented generation (RAG).
- Procedural Memory: Retains operational strategies, tool usage instructions, and self-reflection loops learned from prior task iterations.
By implementing Amazon S3 Vectors as the storage provider for Semantic and Procedural memory, we gain horizontal scalability, near-limitless capacity, and cost optimization over traditional in-memory vector databases.
Implementing Amazon S3 Vectors as a NAT Custom Memory Provider
NAT exposes an extensible abstract base class, BaseMemoryProvider, allowing developers to write custom connectors. Below, we implement a production custom memory provider that bridges NAT with Amazon S3 Vectors using Python and boto3.
1. Provider Core Implementation
import os
import json
import logging
from typing import List, Dict, Any, Optional
import boto3
from botocore.exceptions import BotoCoreError, ClientError
# Native NAT Interfaces
from nemo_agent_toolkit.memory.base import BaseMemoryProvider, MemoryRecord
logger = logging.getLogger(__name__)
class S3VectorMemoryProvider(BaseMemoryProvider):