AWS Bedrock
Bedrock is AWS's managed door into foundation models - Anthropic, Meta, Amazon's own models, and others - without you having to run any GPU infrastructure yourself. If your GCP experience is with Vertex AI, Bedrock covers the same ground with AWS's naming and API shapes.
Bedrock is AWS's foundation-model access and orchestration layer: model hosting/invocation (Converse API), managed RAG (Knowledge Bases), safety filtering (Guardrails), and agent orchestration (AgentCore). It's the AWS-native equivalent of Vertex AI's model access + RAG Engine + Vertex AI Guardrails, with different provisioning and IAM conventions.
- Call any Bedrock model through the Converse API and explain what the unified schema buys you
- Describe what Knowledge Bases, Guardrails and AgentCore each manage, and their Google Cloud and Azure counterparts
- Pick cost and scale levers - batch, prompt caching, cross-Region inference, provisioned throughput - for a workload
- Choose between fine-tuning, distillation and Custom Model Import for customization
- RAG Fundamentals for the Knowledge Bases section
Converse API and Model Access
The Converse API is one consistent way to talk to any model Bedrock hosts - Anthropic's Claude, Meta's Llama, Amazon's Titan/Nova - so switching providers doesn't mean rewriting your integration code.
Converse/ConverseStream is a unified request/response shape across all Bedrock model providers - system prompts, multi-turn messages, tool use, and streaming all follow one schema regardless of which underlying model you call. Since September 2025, serverless models are enabled by default in every account (Anthropic models still need a one-time use-case form); administrators restrict access with IAM policies and Service Control Policies instead of a per-model opt-in. Invocation is billed per token like any hosted LLM API.
Knowledge Bases (Managed RAG)
Knowledge Bases is Bedrock's built-in RAG - point it at your documents in S3, and it handles chunking, embedding, and retrieval for you, without you standing up your own vector database.
Knowledge Bases handles ingestion (chunking strategy, embedding model choice), vector storage (a quick-create OpenSearch Serverless collection, or your own store such as Aurora PostgreSQL, S3 Vectors, Pinecone or Redis), and retrieval, exposed either as a direct Retrieve/RetrieveAndGenerate API call or as a tool an agent can call. This maps to Vertex AI's RAG Engine - same problem (managed retrieval pipeline instead of hand-rolled chunking/embedding/indexing), different vendor plumbing. See RAG Fundamentals for the underlying concepts this wraps.
Guardrails and AgentCore
Guardrails filter what goes in and out of a model call - blocking denied topics, redacting sensitive data - independent of whatever model you're using. AgentCore is Bedrock's newer layer for running full agents (not just single model calls) with session/memory management built in.
Guardrails apply configurable policies (denied topics, content filters, PII redaction, word filters) to both the input prompt and the model output, model-agnostically - the same guardrail config applies whether the underlying call goes to Claude or Llama. AgentCore provides managed agent runtime primitives: session state, memory, and tool orchestration for longer-running agentic workloads, roughly analogous to what a self-hosted LangGraph deployment provides but managed by AWS. AgentCore reached general availability in October 2025 and is framework-agnostic: its components (Runtime, Memory, Gateway for turning APIs into agent tools, Identity, Observability, Code Interpreter and Browser tools, plus Policy and Evaluations) can host agents built with any framework or model. See Production Agents for the underlying patterns.
Cost, Scale and Customization
| Lever | What it does | When to use |
|---|---|---|
| Batch inference | Submit a file of requests for asynchronous processing at a discount to on-demand pricing (for supported models) | Offline enrichment, evaluation runs, bulk summarization |
| Prompt caching | Caches a repeated prompt prefix so later requests read it at reduced price and latency | Long system prompts, RAG context, agent tool definitions |
| Cross-Region inference profiles | Route requests across Regions to absorb traffic bursts | Spiky traffic that hits per-Region throughput limits |
| Provisioned throughput | Reserved model capacity for a term | Steady high volume; required for some custom models |
| Fine-tuning and distillation | Customize supported models on your data, or distill a large model into a smaller one | Narrow tasks where a smaller tuned model is cheaper |
| Custom Model Import | Bring your own open-weight fine-tune (supported architectures) and serve it on Bedrock | You trained with the Fine-Tuning Lab workflow and want managed serving |
Study Notes
Must-know for interviews:
- Converse API is model-agnostic within Bedrock - one schema for Claude, Llama, Titan, Nova
- Knowledge Bases = Bedrock's managed RAG, roughly equivalent to Vertex AI's RAG Engine
- Guardrails apply independent of which model is behind the call - not baked into any one model
- AgentCore is Bedrock's managed agent runtime layer, newer than Converse/Knowledge Bases
- Serverless models are enabled by default (since Sept 2025; Anthropic needs a one-time form) - restrict with IAM/SCPs
- Cost and scale levers: batch inference, prompt caching, cross-Region inference profiles, provisioned throughput
- Customization: fine-tuning, distillation, and Custom Model Import for your own open-weight fine-tunes
Check Yourself
- A nightly job summarizes 200,000 support tickets; results are needed by morning. Which Bedrock lever cuts cost most directly?
- How is access to serverless Bedrock models controlled today?
- Why would a team choose Bedrock's Converse API over calling each model provider's native SDK directly?
- What's the AWS-side equivalent of Vertex AI's RAG Engine?
- Do you still need to request model access in Bedrock?
- Do Bedrock Guardrails apply per-model or account-wide?
Exercises
An assistant on Google Cloud uses Gemini through Model Garden, RAG Engine over Cloud Storage documents, Model Armor for prompt screening, and an ADK agent on Agent Runtime. Name the AWS component for each, and one design decision that changes in the move.
Solution
- Model access: Bedrock
Converse/ConverseStream(Claude, Llama, Nova and others; Gemini is not on Bedrock, so the model changes and prompts need re-evaluation). - RAG: Bedrock Knowledge Bases over S3, with a chosen vector store (OpenSearch Serverless, Aurora PostgreSQL, S3 Vectors).
- Prompt screening: Bedrock Guardrails (denied topics, content filters, PII, prompt-attack filter) applied to input and output.
- Agent runtime: AgentCore Runtime (framework-agnostic - the ADK agent can run there), with AgentCore Memory, Gateway and Identity.
Design decisions that change: the model itself (re-run your eval suite), the vector-store choice and its cost model, and IAM - each component gets its own role with least privilege.
References
- AWS, Amazon Bedrock User Guide
- AWS, Amazon Bedrock simplifies access with automatic enablement of serverless foundation models (2025)
- AWS, Amazon Bedrock AgentCore is now generally available (2025)
Last reviewed: 2026-09