Building MCP Servers & Clients
How to build, test, deploy and consume MCP servers with current tooling: the official Python SDK v2 (MCPServer, formerly FastMCP), the other official SDKs, the standalone FastMCP framework, the MCP Inspector, and the ways agents connect to servers - your own loop, a model provider's hosted MCP connector, or an agent framework. The code samples are condensed from the module's lab, which runs against mcp 2.2 on protocol 2026-07-28.
- Build an MCP server with the Python SDK v2 - tools, structured output, resources, prompts and an elicitation resolver - and run it over stdio and Streamable HTTP
- Test a server in-process and inspect it with the MCP Inspector
- Convert MCP tool definitions into a model's function-calling format and execute calls through an MCP client
- Choose between a hand-written loop, a provider-hosted MCP connector and a framework integration
- Apply a design checklist for production MCP servers
The Tooling Landscape
| Tool | What it is | Use it for |
|---|---|---|
| Official SDKs - Tier 1: TypeScript, Python, Go, C#; others include Java, Kotlin, Rust, Ruby, Swift, PHP | Maintained under the MCP project; Tier 1 SDKs support 2026-07-28 and legacy clients | Most servers and clients |
Python SDK v2 (pip install mcp) | High-level MCPServer (renamed from FastMCP in v2) plus the low-level Server; a high-level Client | Python servers and clients |
FastMCP (standalone, pip install fastmcp) | The project whose high-level API was adopted into the official SDK in 2024, now developed independently | Extras: composing and proxying servers, generating servers from OpenAPI specs, auth provider integrations, MCP Apps |
MCP Inspector (npx @modelcontextprotocol/inspector) | A browser UI that connects to a server and lets you list and call everything by hand | Debugging servers before any model is involved |
Version note. Much MCP code online targets Python SDK v1 (from mcp.server.fastmcp import FastMCP). In v2 that import fails with a pointer to the migration guide: the class is mcp.server.mcpserver.MCPServer, transport options moved from the constructor to run(), model fields are snake_case (result.is_error, tool.input_schema), and Context is injected as a parameter. Pin mcp<2 for old code, or migrate.
A Server
from typing import Annotated
from pydantic import BaseModel, Field
from mcp.server import CacheHint
from mcp.server.mcpserver import (AcceptedElicitation, Elicit, ElicitationResult, MCPServer, Resolve)
from mcp.server.mcpserver.exceptions import ToolError
from mcp.types import ToolAnnotations
mcp = MCPServer("shop", version="0.3.0",
instructions="Customer-service tools for an outdoor-gear shop. Read shop://policy first.",
cache_hints={"tools/list": CacheHint(ttl_ms=300_000, scope="private")})
class Order(BaseModel): # a return type -> outputSchema + structuredContent
order_id: str
status: str
total: float
@mcp.tool(annotations=ToolAnnotations(readOnlyHint=True, openWorldHint=False))
def get_order(order_id: str) -> Order:
"""Get one order's status and total. Check the owner and status before any change."""
if order_id not in DB:
raise ToolError(f"No order {order_id!r}") # -> result with isError: true
return Order(**DB[order_id])
class ConfirmRefund(BaseModel):
confirm: bool = Field(description="Issue this refund?")
async def confirm_large(order_id: str, item_id: int) -> ConfirmRefund | Elicit[ConfirmRefund]:
amount = refund_amount(order_id, item_id)
return ConfirmRefund(confirm=True) if amount <= 100 else Elicit(f"Refund {amount:.2f}?", ConfirmRefund)
@mcp.tool(annotations=ToolAnnotations(readOnlyHint=False, destructiveHint=True, idempotentHint=False))
def refund_item(order_id: str, item_id: int,
confirmation: Annotated[ElicitationResult[ConfirmRefund], Resolve(confirm_large)]) -> dict:
"""Refund one line of a delivered order. Refunds over 100.00 ask the user to confirm."""
match confirmation:
case AcceptedElicitation(data=ConfirmRefund(confirm=True)):
return do_refund(order_id, item_id)
raise ToolError("The user did not confirm the refund; nothing was refunded.")
@mcp.resource("shop://policy", mime_type="text/markdown")
def policy() -> str:
return POLICY_MARKDOWN
@mcp.prompt(title="Handle a customer request")
def customer_request(email: str, request: str) -> str:
return f"I'm {email}. {request}"
if __name__ == "__main__":
mcp.run() # stdio
# mcp.run(transport="streamable-http", host="127.0.0.1", port=8765) # http://127.0.0.1:8765/mcp
What the SDK does for you: derives inputSchema from type hints and the description from the docstring; derives outputSchema from a Pydantic return type and returns structuredContent plus a text copy; turns ToolError into an isError result; and runs the Resolve dependency before the tool body - returning an InputRequiredResult when it yields Elicit, then calling the tool with the user's answer on the client's retry. That resolver pattern works over the stateless 2026-07-28 protocol; the older ctx.elicit() call only works for legacy (2025-11-25) connections, because it sends a server-initiated request. For state that must survive the round trip, the SDK can protect requestState with an AES-GCM codec.
A Client
from mcp import Client
from mcp.client.stdio import StdioServerParameters
from mcp.types import ElicitResult
async def ask_user(context, params):
print(params.message) # the host shows which server asks what
return ElicitResult(action="accept", content={"confirm": True})
target = "http://127.0.0.1:8765/mcp" # or StdioServerParameters(command="python",
# args=["shop_server.py"]), or an MCPServer
async with Client(target, elicitation_callback=ask_user) as client:
tools = await client.list_tools() # tools.ttl_ms, tools.cache_scope
result = await client.call_tool("refund_item", {"order_id": "O1001", "item_id": 1})
print(result.is_error, result.structured_content)
Client picks the protocol era automatically (mode="auto"), sends the per-request _meta, and runs the MRTR loop for you: on an InputRequiredResult it calls your callback for each input request and retries with the answers, up to input_required_max_rounds. Passing an MCPServer object connects in-process - the fastest way to unit-test a server.
Connecting MCP Tools to an Agent
Your own loop
A host converts MCP tool definitions into its model's function-calling format and routes calls back through the client. This is all the "MCP support" a loop needs:
def to_openai_tools(mcp_tools):
return [{"type": "function", "function": {
"name": t.name, "description": t.description or "", "parameters": t.input_schema}}
for t in mcp_tools]
# inside the loop, for each tool call the model makes:
r = await client.call_tool(call.function.name, json.loads(call.function.arguments))
body = r.structured_content if r.structured_content is not None else r.content[0].text
messages.append({"role": "tool", "tool_call_id": call.id,
"content": json.dumps({"ok": not r.is_error, "result": body})})
With several servers, prefix tool names with a server id to avoid collisions, and keep a map from prefixed name back to the right client.
Provider-hosted connectors
Model APIs can call remote MCP servers themselves, so your code never runs the loop for those tools:
| API | Configuration | Notes |
|---|---|---|
Claude Messages API (beta mcp-client-2025-11-20) | mcp_servers=[{"type": "url", "url": ..., "name": "shop", "authorization_token": ...}] plus a {"type": "mcp_toolset", "mcp_server_name": "shop"} entry in tools | Per-tool enable/disable via toolset configs |
| OpenAI Responses API | {"type": "mcp", "server_label": "shop", "server_url": ..., "allowed_tools": [...], "require_approval": "always"} in tools | Approval required by default; the model's call pauses for your approval response |
Hosted connectors only reach servers on the public internet, and the provider sees the tool traffic - check that fits your data policies. For local servers or full control, use your own client.
Frameworks
LangGraph/LangChain (MCP adapters), the OpenAI Agents SDK, the Claude Agent SDK, Google ADK and others accept MCP servers as tool sources directly - see Agent Frameworks.
Testing and Deploying
Test in layers. (1) Unit-test tools as plain functions. (2) In-process Client(server) tests for schemas, errors, elicitation paths and the tool list's stability. (3) The MCP Inspector for manual exploration. (4) Agent-level evaluation - does a model use your tools correctly? - with the methods in Agent Evaluation Basics; the lab measures it.
Deploy HTTP servers like stateless web services. Because 2026-07-28 has no sessions, run several identical instances behind a load balancer; keep cross-call state in a shared store keyed by user-bound handles. Terminate TLS, validate Origin, authenticate every request as a resource server (Authorization), set request timeouts and rate limits, and emit OpenTelemetry traces (propagated through _meta). Distribute stdio servers as versioned packages, and document the exact launch command.
Server design checklist
- Tools shaped around user tasks, not API endpoints; few and well described; names prefixed by domain
- Strict input schemas (
additionalProperties: false); Pydantic return types for structured output - Execution errors as
isErrorresults that say what to do next - Honest annotations; confirmation (elicitation) for destructive or expensive actions
- Resources for context the application should choose; prompts for user workflows
- Deterministic tool order; sensible
ttlMs;listChangedonly when the list really changes - Authorization with least-privilege scopes; tool lists filtered by scope; no token passthrough
- Pagination and size limits on large outputs; no secrets in outputs or logs
Check Yourself
- Old code does
from mcp.server.fastmcp import FastMCPand fails on mcp 2.x. What is the fix? - Why does the lab's refund tool use a Resolve dependency instead of calling ctx.elicit() inside the tool?
- What is the quickest way to test an MCP server's tools without starting a process or a network listener?
- When would you use a provider-hosted MCP connector rather than your own client?
Exercises
Extend the lab's server with a track_shipment(order_id) tool (read-only, structured output with carrier and eta) and a shop://customers/{customer_id}/orders resource template. Test both with an in-process Client.
Solution
Add a Pydantic model Shipment(order_id, carrier, eta) and a @mcp.tool(annotations=ToolAnnotations(readOnlyHint=True)) returning it; raise ToolError for orders that aren't shipped. Add @mcp.resource("shop://customers/{customer_id}/orders", mime_type="application/json") returning json of list_orders. Test: list_resource_templates shows the template; read_resource("shop://customers/C2/orders") returns two orders; call_tool on a pending order returns is_error.
Write a host that connects to two MCP servers (the lab's shop and a second tiny server with a tool also called get_order) and exposes both to the model without collisions.
Solution
Keep {prefixed_name: (client, original_name)}: prefix with a server id you choose ("shop__get_order", "crm__get_order"), since server names aren't guaranteed unique. Convert each tool with the prefixed name; on a call, look up the client and original name and call it there. Note the prefix in descriptions if it helps the model distinguish them.
Study Notes
- Tier 1 SDKs (TS, Python, Go, C#) support 2026-07-28; Python SDK v2: FastMCP → MCPServer, run() takes transport options, snake_case fields
- Standalone FastMCP adds composition, proxying, OpenAPI generation, auth integrations, apps
- MCPServer: type hints → inputSchema; Pydantic returns → outputSchema + structuredContent; ToolError → isError; Resolve + Elicit → MRTR elicitation
- Client: auto era detection, per-request _meta, automatic MRTR retries; in-process Client(server) for tests
- Hosts convert MCP tools to function-calling schemas; prefix names across servers
- Hosted connectors: Claude mcp_servers + mcp_toolset; OpenAI {"type": "mcp", ...} with approvals
- Deploy HTTP servers statelessly behind a load balancer with Origin validation, OAuth, limits, tracing
References
- MCP Python SDK v2 documentation and v1 → v2 migration guide
- MCP SDKs overview and MCP Inspector
- FastMCP documentation
- Anthropic, MCP connector (docs, 2026)
- OpenAI, MCP servers and connectors (docs, 2026)
Last reviewed: 2026-09