Contents
Map

10 · Cloud Platforms

AWS Bedrock

View as:

AWS Bedrock

Bedrock is AWS's managed door into foundation models - Anthropic, Meta, Amazon's own models, and others - without you having to run any GPU infrastructure yourself. If your GCP experience is with Vertex AI, Bedrock covers the same ground with AWS's naming and API shapes.

Bedrock is AWS's foundation-model access and orchestration layer: model hosting/invocation (Converse API), managed RAG (Knowledge Bases), safety filtering (Guardrails), and agent orchestration (AgentCore). It's the AWS-native equivalent of Vertex AI's model access + RAG Engine + Vertex AI Guardrails, with different provisioning and IAM conventions.

Learning objectives 30 min
By the end of this page you will be able to:
  • Call any Bedrock model through the Converse API and explain what the unified schema buys you
  • Describe what Knowledge Bases, Guardrails and AgentCore each manage, and their Google Cloud and Azure counterparts
  • Pick cost and scale levers - batch, prompt caching, cross-Region inference, provisioned throughput - for a workload
  • Choose between fine-tuning, distillation and Custom Model Import for customization
Prerequisites

Converse API and Model Access

The Converse API is one consistent way to talk to any model Bedrock hosts - Anthropic's Claude, Meta's Llama, Amazon's Titan/Nova - so switching providers doesn't mean rewriting your integration code.

Converse/ConverseStream is a unified request/response shape across all Bedrock model providers - system prompts, multi-turn messages, tool use, and streaming all follow one schema regardless of which underlying model you call. Since September 2025, serverless models are enabled by default in every account (Anthropic models still need a one-time use-case form); administrators restrict access with IAM policies and Service Control Policies instead of a per-model opt-in. Invocation is billed per token like any hosted LLM API.


Knowledge Bases (Managed RAG)

Knowledge Bases is Bedrock's built-in RAG - point it at your documents in S3, and it handles chunking, embedding, and retrieval for you, without you standing up your own vector database.

Knowledge Bases handles ingestion (chunking strategy, embedding model choice), vector storage (a quick-create OpenSearch Serverless collection, or your own store such as Aurora PostgreSQL, S3 Vectors, Pinecone or Redis), and retrieval, exposed either as a direct Retrieve/RetrieveAndGenerate API call or as a tool an agent can call. This maps to Vertex AI's RAG Engine - same problem (managed retrieval pipeline instead of hand-rolled chunking/embedding/indexing), different vendor plumbing. See RAG Fundamentals for the underlying concepts this wraps.


Guardrails and AgentCore

Guardrails filter what goes in and out of a model call - blocking denied topics, redacting sensitive data - independent of whatever model you're using. AgentCore is Bedrock's newer layer for running full agents (not just single model calls) with session/memory management built in.

Guardrails apply configurable policies (denied topics, content filters, PII redaction, word filters) to both the input prompt and the model output, model-agnostically - the same guardrail config applies whether the underlying call goes to Claude or Llama. AgentCore provides managed agent runtime primitives: session state, memory, and tool orchestration for longer-running agentic workloads, roughly analogous to what a self-hosted LangGraph deployment provides but managed by AWS. AgentCore reached general availability in October 2025 and is framework-agnostic: its components (Runtime, Memory, Gateway for turning APIs into agent tools, Identity, Observability, Code Interpreter and Browser tools, plus Policy and Evaluations) can host agents built with any framework or model. See Production Agents for the underlying patterns.


Cost, Scale and Customization

LeverWhat it doesWhen to use
Batch inferenceSubmit a file of requests for asynchronous processing at a discount to on-demand pricing (for supported models)Offline enrichment, evaluation runs, bulk summarization
Prompt cachingCaches a repeated prompt prefix so later requests read it at reduced price and latencyLong system prompts, RAG context, agent tool definitions
Cross-Region inference profilesRoute requests across Regions to absorb traffic burstsSpiky traffic that hits per-Region throughput limits
Provisioned throughputReserved model capacity for a termSteady high volume; required for some custom models
Fine-tuning and distillationCustomize supported models on your data, or distill a large model into a smaller oneNarrow tasks where a smaller tuned model is cheaper
Custom Model ImportBring your own open-weight fine-tune (supported architectures) and serve it on BedrockYou trained with the Fine-Tuning Lab workflow and want managed serving

Study Notes

Must-know for interviews:

  • Converse API is model-agnostic within Bedrock - one schema for Claude, Llama, Titan, Nova
  • Knowledge Bases = Bedrock's managed RAG, roughly equivalent to Vertex AI's RAG Engine
  • Guardrails apply independent of which model is behind the call - not baked into any one model
  • AgentCore is Bedrock's managed agent runtime layer, newer than Converse/Knowledge Bases
  • Serverless models are enabled by default (since Sept 2025; Anthropic needs a one-time form) - restrict with IAM/SCPs
  • Cost and scale levers: batch inference, prompt caching, cross-Region inference profiles, provisioned throughput
  • Customization: fine-tuning, distillation, and Custom Model Import for your own open-weight fine-tunes

Check Yourself

Check yourself
0 / 6 answered
  1. A nightly job summarizes 200,000 support tickets; results are needed by morning. Which Bedrock lever cuts cost most directly?
  2. How is access to serverless Bedrock models controlled today?
  3. Why would a team choose Bedrock's Converse API over calling each model provider's native SDK directly?
  4. What's the AWS-side equivalent of Vertex AI's RAG Engine?
  5. Do you still need to request model access in Bedrock?
  6. Do Bedrock Guardrails apply per-model or account-wide?

Exercises

Exercise - Map a GCP design to AWS

An assistant on Google Cloud uses Gemini through Model Garden, RAG Engine over Cloud Storage documents, Model Armor for prompt screening, and an ADK agent on Agent Runtime. Name the AWS component for each, and one design decision that changes in the move.

Solution
  • Model access: Bedrock Converse / ConverseStream (Claude, Llama, Nova and others; Gemini is not on Bedrock, so the model changes and prompts need re-evaluation).
  • RAG: Bedrock Knowledge Bases over S3, with a chosen vector store (OpenSearch Serverless, Aurora PostgreSQL, S3 Vectors).
  • Prompt screening: Bedrock Guardrails (denied topics, content filters, PII, prompt-attack filter) applied to input and output.
  • Agent runtime: AgentCore Runtime (framework-agnostic - the ADK agent can run there), with AgentCore Memory, Gateway and Identity.

Design decisions that change: the model itself (re-run your eval suite), the vector-store choice and its cost model, and IAM - each component gets its own role with least privilege.

References

Last reviewed: 2026-09

⚡AI-assisted content - always verify, always explore multiple perspectives·