Tag: amazon-bedrock
All the articles with the tag "amazon-bedrock".
Tenancy, Rate Limits and PII Masking on AgentCore Gateway
Three articles of per-tenant controls — budgets, caches, rate limits — all keyed on a header any caller could set. Replacing it with a validated token costs 0.53 ms. Along the way: why a request-rate limit cannot protect a token quota, why a cache hit was silently stealing quota it never spent, and why the PII screen that costs 50x less is the one you must not use.
Semantic Caching for LLMs on AWS
An exact LLM cache almost never hits, and the obvious fix — serve the answer to the nearest stored prompt — is worse than no cache: the near-misses that change the answer sit closer in embedding space than the paraphrases that don't. A verification stage fixes the ranking and leaves a 5% floor of confidently wrong answers that you cannot measure in production. Built on DynamoDB vector search, measured three ways, and mostly an argument for not running it.
Customizing inference on AgentCore Gateway: a plugin pipeline for routing, budgets and guardrails
Every model call through Amazon Bedrock AgentCore Gateway passes one place where you can check a budget, clamp tokens, run a guardrail, pick the model, cache the answer and meter the cost. A reference architecture that makes each of those a swappable plugin, deployed and measured, plus seven things about interceptors the documentation does not tell you.
Scaling LiteLLM on AWS
Updated:A reference architecture and operating guide for running LiteLLM as an enterprise LLM gateway on AWS with Amazon Bedrock, plus the FinOps, governance and security work that scales with it, and a managed alternative on AgentCore Gateway.