Agentic Retrieval with LangChain and Amazon Bedrock Knowledge Bases
- Authors

- Name
- Nino
- Occupation
- Senior Tech Editor
Retrieval-Augmented Generation (RAG) has transformed how enterprise applications interact with proprietary data. However, standard single-shot RAG systems often fail when presented with multi-part, comparative, or ambiguous queries. When a prompt requires pulling data across disparate documents or synthesizing sub-insights sequentially, single-shot vector searches return incomplete or irrelevant contexts.
To overcome these limitations, developers are turning to Agentic Retrieval. By pairing orchestration frameworks like LangChain with managed vector repositories like Amazon Bedrock Knowledge Bases, you can empower Large Language Models (LLMs) to iteratively decompose queries, execute dynamic tool calls, reflect on retrieved context, and synthesize comprehensive answers.
In high-throughput enterprise architectures, managing model endpoints across cloud vendors requires low latency and high availability. Leveraging unified API solutions such as n1n.ai enables teams to seamlessly route agent orchestrations between models like Claude 3.5 Sonnet, DeepSeek-V3, and OpenAI o3 with zero infrastructure friction.
This article provides a practical comparison between single-shot RAG and agentic retrieval using LangChain and Amazon Bedrock Knowledge Bases. We will review code implementations, trace execution events, evaluate multi-part query handling, and perform a cost-latency audit.
Single-Shot RAG vs. Agentic Retrieval Architecture
To understand why single-shot retrieval degrades on complex questions, we must examine the architectural differences between the two patterns.
Single-Shot RAG Flow
- Input: User submits a complex prompt (e.g., "Compare our 2023 cloud hosting expenditure against our cybersecurity compliance updates in 2024.").
- Embedding & Vector Search: The entire prompt is converted into a single dense vector embedding.
- Top-K Retrieval: The vector database matches the prompt against index chunks using cosine similarity or HNSW indexing.
- Context Injection: Top-K retrieved passages are stuffed directly into the system context window.
- Generation: The LLM generates an answer based solely on this single retrieval pass.
The Failure Mode: The single-embedding vector representation merges two distinct topics (hosting expenditure and compliance updates). The retrieved context usually captures one topic well while missing the other entirely.
Agentic Retrieval Flow
- Plan & Decompose: An LLM agent evaluates the prompt and determines that two separate search actions are required.
- Sub-Query Generation: The agent formulates Sub-Query A ("2023 cloud hosting expenditure") and queries the retriever tool.
- Context Evaluation: The agent inspects the returned documents. If sufficient, it proceeds; if incomplete, it refines the search query.
- Sequential Execution: The agent formulates Sub-Query B ("2024 cybersecurity compliance updates") and invokes the retriever tool again.
- Synthesis: The agent aggregates scratchpad context from all tool calls to output a structured, complete answer.
+-----------------------------------------------------------------------------------+
| USER INPUT QUERY |
+-----------------------------------------------------------------------------------+
|
+------------------------+------------------------+
| |
v v
[ Single-Shot RAG Path ] [ Agentic Retrieval Path ]
| |
+---------------------------+ +---------------------------+
| Embed Full Query Once | | Agent Planner LLM |
+---------------------------+ +---------------------------+
| |
+---------------------------+ +---------------------------+
| Vector DB Top-K Search | | Action: Call Retrieval A |
+---------------------------+ +---------------------------+
| |
+---------------------------+ +---------------------------+
| Pass Context to LLM | | Evaluate Context Scratchpad|
+---------------------------+ +---------------------------+
| |
+---------------------------+ +---------------------------+
| Final Generation | | Action: Call Retrieval B |
+---------------------------+ +---------------------------+
| |
v +---------------------------+
(Incomplete Context) | Final Synthesis LLM Call |
+---------------------------+
|
v
(Complete Synthesis)
Implementation: Building Both Paths with LangChain and Amazon Bedrock
Let's set up both single-shot and agentic architectures programmatically. We use Amazon Bedrock Knowledge Bases as our managed retriever and LangChain for orchestration.
Environment Prerequisites
Ensure you have installed the required dependencies:
pip install langchain langchain-aws langchain-community boto3 botocore
1. Single-Shot RAG Implementation
In single-shot RAG, we construct an AmazonBedrockKnowledgeBaseRetriever and pipe its output directly into a standard chain using LangChain Expression Language (LCEL).
import os
from langchain_aws import AmazonBedrockKnowledgeBaseRetriever, ChatBedrock
from langchain_core.prompts import ChatPromptTemplate
from langchain_core.runnables import RunnablePassthrough
from langchain_core.output_parsers import StrOutputParser
# Configuration
KNOWLEDGE_BASE_ID = "YOUR_BEDROCK_KB_ID"
REGION_NAME = "us-east-1"
# Initialize Bedrock KB Retriever
retriever = AmazonBedrockKnowledgeBaseRetriever(
knowledge_base_id=KNOWLEDGE_BASE_ID,
retriever_config=\{
"vectorSearchConfiguration": \{
"numberOfResults": 4
\}
\},
region_name=REGION_NAME
)
# Initialize LLM
llm = ChatBedrock(
model_id="anthropic.claude-3-5-sonnet-20240620-v1:0