Tag: architecture
All the articles with the tag "architecture".
Tenancy, Rate Limits and PII Masking on AgentCore Gateway
Three articles of per-tenant controls — budgets, caches, rate limits — all keyed on a header any caller could set. Replacing it with a validated token costs 0.53 ms. Along the way: why a request-rate limit cannot protect a token quota, why a cache hit was silently stealing quota it never spent, and why the PII screen that costs 50x less is the one you must not use.
Semantic Caching for LLMs on AWS
An exact LLM cache almost never hits, and the obvious fix — serve the answer to the nearest stored prompt — is worse than no cache: the near-misses that change the answer sit closer in embedding space than the paraphrases that don't. A verification stage fixes the ranking and leaves a 5% floor of confidently wrong answers that you cannot measure in production. Built on DynamoDB vector search, measured three ways, and mostly an argument for not running it.
Sizing Aurora Connections for Autoscaled Services on AWS
A service that scales on request rate scales its connection count with it, and Aurora's ceiling is fixed by instance memory. The arithmetic, the formula behind max_connections, why lowering the pool limit is a trap, and what RDS Proxy does and does not fix.
Customizing inference on AgentCore Gateway: a plugin pipeline for routing, budgets and guardrails
Every model call through Amazon Bedrock AgentCore Gateway passes one place where you can check a budget, clamp tokens, run a guardrail, pick the model, cache the answer and meter the cost. A reference architecture that makes each of those a swappable plugin, deployed and measured, plus seven things about interceptors the documentation does not tell you.