Inside Microsoft Foundry Agent Endpoint: Versions, Canary Rollouts, and Publishing to Teams & Copilot
- Authors

- Name
- Nino
- Occupation
- Senior Tech Editor
Building a prompt agent inside a developer playground is straightforward. You iterate on system instructions, attach a vector store for RAG, test a few multi-turn edge cases, and verify that the output format meets your schema. However, transitioning from a localized prototype to an enterprise-grade deployment introduces immediate operational questions:
- How do you deploy version 2 of an agent without breaking existing client applications or active conversations?
- How do you safely route 5% of production traffic to a candidate prompt or model upgrade while retaining instant rollback capabilities?
- How do you expose a single logical agent across multiple interfaces—a REST API endpoint, a Microsoft Teams channel, a Microsoft 365 Copilot plugin, and an Agent-to-Agent (A2A) network—without duplicating business logic?
Microsoft Foundry Agent Service solves these production challenges through a decoupled object model. Instead of binding runtime behavior directly to a static prompt string or deployment deployment script, Foundry separates Identity (the stable Agent entity), State & Logic (immutable Agent Versions), Routing (the Version Selector), Transport (Multi-protocol endpoints), and Governance (Authorization schemes and Publishing catalog).
When scaling complex multi-agent architectures, enterprise teams often combine agent governance platforms with flexible API aggregation infrastructure. Platforms like n1n.ai provide unified access to leading underlying foundation models—such as OpenAI o3, Claude 3.5 Sonnet, and DeepSeek-V3—ensuring high availability, cost control, and low-latency backends for your agent orchestration layer.
In this comprehensive architectural guide, we will break down the mechanics of the Microsoft Foundry Agent endpoint, explore hands-on implementation patterns for canary traffic splitting, and walk through enterprise publishing pipelines for Microsoft Teams and Copilot.
1. The Core Hierarchy: Four Nested Entities
To operate Microsoft Foundry agents reliably at scale, you must understand the four structural layers that govern runtime execution:
+-----------------------------------------------------------------------+
| Foundry Project |
| (Logical RBAC, Resource Grouping, Billing Attribution Scope) |
| |
| +---------------------------------------------------------------+ |
| | Agent | |
| | (Stable Identity, Name, Permanent URL Identifier) | |
| | | |
| | +-------------------------------------------------------+ | |
| | | Agent Version | | |
| | | (Immutable Snapshot: Instructions, Model, Tools) | | |
| | +-------------------------------------------------------+ | |
| | | |
| | +-------------------------------------------------------+ | |
| | | Agent Endpoint | | |
| | | (Version Selector, Protocols, Auth Schemes) | | |
| | +-------------------------------------------------------+ | |
| +---------------------------------------------------------------+ |
+-----------------------------------------------------------------------+
- Foundry Project: The overarching container defining RBAC, access policies, telemetry tracking, and resource allocation.
- Agent: The consumer-facing identity. It holds a persistent name, icon, description, and canonical URL. External callers interact with the Agent identity rather than tying themselves to underlying prompt code.
- Agent Version: An immutable snapshot of the agent's configuration. Modifying instructions, switching baseline models (e.g., upgrading from GPT-4o to a fast model via n1n.ai), or adding function tools automatically mints a new numerical version (e.g.,
v1,v2,v3). Previous versions remain frozen in perpetuity. - Agent Endpoint: The network listener and request processor. Located at a stable URI, the endpoint uses dynamic routing configuration to resolve which underlying Agent Version processes an incoming request.
The canonical endpoint URL follows this structure:
https://{account}.services.ai.azure.com/api/projects/{project}/agents/{agent}/endpoint/protocols/{protocol}
2. Anatomy of a Request: The Four-Stage Pipeline
When a payload hits an Agent Endpoint, the Microsoft Foundry agent runtime processes the request through four sequential evaluation gates before performing any LLM inference or executing tool code:
Inbound Payload
│
▼
┌───────────────────────────┐
│ 1. Protocol Negotiation │ ──► Resolves wire format (Responses, Activity, A2A, MCP)
└─────────────┬─────────────┘
│
▼
┌───────────────────────────┐
│ 2. Authorization Check │ ──► Validates Entra ID, BotServiceRbac, or BotServiceTenant
└─────────────┬─────────────┘
│
▼
┌───────────────────────────┐
│ 3. Version Resolution │ ──► Evaluates version_selector rules (Latest, Pinned, Canary)
└─────────────┬─────────────┘
│
▼
┌───────────────────────────┐
│ 4. Session & State Key │ ──► Binds thread state via user_isolation_key / chat_isolation_key
└─────────────┬─────────────┘
│
▼
LLM Inference & Tool Loop
- Protocol Negotiation: The URI pattern determines the active protocol adapter (e.g.,
/protocols/responsesor/protocols/activityprotocol). - Authorization Check: The runtime enforces token or channel-level authorization based on active schemes.
- Version Resolution: The
version_selectordetermines which immutable Agent Version config (prompt, tools, bindings) receives the execution context. - Session Isolation Resolution: Thread state is loaded based on identity headers or custom isolation keys.
3. Protocol Surface Matrix
Foundry endpoints support multiple wire protocols concurrently. This allows developers to use a single centralized agent definition across diverse integration targets:
| Protocol | Schema / Specification | Ideal Use Case | Primary Calling Client |
|---|---|---|---|
| Responses | OpenAI Responses API format | Web apps, backend microservices, custom SDKs | Native Python/TS apps, API gateways |
| Activity Protocol | Bot Framework Activity schema | Enterprise Chat Channels | Microsoft Teams, M365 Copilot |
| Invocations | Lightweight REST invocation | Serverless functions, event-driven webhooks | Azure Functions, Event Grid |
| A2A (Preview/GA) | Agent-to-Agent communication protocol | Multi-agent orchestration frameworks | External peer agents, AutoGen, CrewAI |
| MCP (Preview) | Model Context Protocol | Exposing agent capability as a tool | MCP-compatible client host systems |
By leveraging API aggregation services like n1n.ai for multi-model fallback routines, enterprise developers can maintain ultra-high availability for backend agents while serving Teams users over Activity Protocol and backend services over Responses API simultaneously.
4. Implementation: Canary Rollouts & Endpoint Configuration
Instead of relying on risky "blue-green" cutovers or manual portal changes, Foundry enables weighted traffic splitting via FixedRatio rules in the version_selector config.
The snippet below demonstrates how to configure an endpoint using the official Python SDK to route 90% of live traffic to known-good Version 4 and 10% of traffic to candidate Version 5, while concurrently exposing Responses, Activity, and A2A protocols.
import os
from azure.ai.projects import AIProjectClient
from azure.identity import DefaultAzureCredential
from azure.ai.projects.models import (
AgentEndpointConfig,
FixedRatioVersionSelectionRule,
VersionSelector,
ProtocolConfiguration,
ResponsesProtocolConfiguration,
ActivityProtocolConfiguration,
A2AProtocolConfiguration,
EntraAuthorizationScheme,
BotServiceRbacAuthorizationScheme,
)
# Initialize client using Entra ID DefaultAzureCredential
PROJECT_ENDPOINT = os.environ.get("AZURE_FOUNDRY_PROJECT_ENDPOINT")
AGENT_NAME = "enterprise-compliance-agent"
project_client = AIProjectClient(
endpoint=PROJECT_ENDPOINT,
credential=DefaultAzureCredential(),
)
with project_client:
# 1. Define Version Selector: 90% traffic to v4, 10% traffic to v5 (Canary)
version_selector = VersionSelector(
version_selection_rules=[
FixedRatioVersionSelectionRule(agent_version="4