• Open

    Query claims in natural language with Amazon Bedrock Knowledge Bases
    This technical how-to builds a conversational claims assistant on Amazon Bedrock Knowledge Bases that answers natural-language questions with citations. It covers ingesting claim documents from Amazon S3, querying with the AgenticRetrieveStream API, multi-turn follow-ups, metadata filters, and contextual grounding guardrails.  ( 122 min )
    Build a multi-agent music production pipeline on Amazon Bedrock AgentCore Runtime Instances
    Amazon Bedrock AgentCore Runtime Instances gives multi-agent workflows AWS managed EC2 infrastructure with GPUs, persistent volumes, and multi-day sessions. In this post, we deploy a three-agent music production pipeline where the agents colocate on one GPU instance, share a filesystem, and hand work to each other to produce a finished track.  ( 123 min )
    Amazon Bedrock expands Claude model availability to in-country inferencing in India
    Anthropic's Claude Opus 5, Claude Sonnet 5, and Claude Haiku 4.5 are now available in India through Amazon Bedrock geographic cross-Region inference. You can access these models while processing data within the India Regions, and get started from the Amazon Bedrock console or with the Messages, InvokeModel, and Converse APIs.  ( 116 min )
    Introducing Anthropic models on Amazon Bedrock for in-region inference in Seoul and Singapore
    Amazon Bedrock now supports Anthropic's Claude Opus 5 and Claude Sonnet 5 with in-region inference in Seoul, and Claude Sonnet 5 in Singapore. If you have local data processing requirements in South Korea or Singapore, you can now use these Anthropic models at scale, with inference processed entirely within the Region you call.  ( 116 min )

  • Open

    Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock
    GPT-6.1 Sol is now generally available on Amazon Bedrock, bringing stronger reasoning to coding, computer use, and professional workloads that run frequently.  ( 115 min )
    Prompt engineering fundamentals for Amazon Quick
    Prompt engineering in Amazon Quick shapes how accurately its AI-powered features respond to your requests. Part 1 of a two-part series covers the foundational principles and reusable frameworks (specificity, context-setting, few-shot examples, and the CRISPE framework) for consistent, high-quality results across Amazon Quick.  ( 120 min )
    Prompt engineering by Quick component: Patterns and pitfalls
    Part 2 of our Amazon Quick prompt engineering series goes component by component. Learn the prompt patterns that get the best results from Amazon Quick Research, Quick Flows, Quick Sight, chat agents, and action integrations, plus the common pitfalls to avoid.  ( 121 min )
    Building an AI-powered contract intelligence platform with Amazon Quick and Amazon Bedrock AgentCore
    Manually extracting data from hundreds of vendor contracts doesn't scale, and RAG chat tools fall short on portfolio-wide questions. This post shares a contract intelligence platform on AWS that uses AI agents to extract and verify contract fields, then answers aggregate and single-contract questions through Amazon Quick analytics.  ( 121 min )
    How Condé Nast built multimodal video discovery with Amazon Bedrock
    Condé Nast's editorial teams spent an average of 250 minutes per task searching a library of more than 140,000 videos using only titles and descriptions. Working with the AWS Generative AI Innovation Center, they built a multimodal video discovery solution on Amazon Bedrock and Amazon OpenSearch Service that cut discovery time to under 2 minutes.  ( 119 min )

  • Open

    Grok 4.7 is now available on Amazon Bedrock
    xAI's Grok 4.7 is now available on Amazon Bedrock: a frontier model for coding, long-running agents, and knowledge work. It offers a 500K token context window and four configurable reasoning effort levels, reachable through the Responses, Chat Completions, and Converse APIs.  ( 120 min )
    Introducing Claude Sonnet 5.5 on AWS
    Claude Sonnet 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. It's a smarter, more efficient Sonnet model for focused coding and knowledge work, with a lower cost per task at faster speed. This post covers what's new, when to choose Sonnet, and how to get started.  ( 115 min )
    Build real-time voice applications with vLLM-Omni on SageMaker AI – Part 1
    Deploy a text-to-speech model on Amazon SageMaker AI with the AWS vLLM-Omni Deep Learning Container and stream generated speech over a persistent bidirectional connection. This Part 1 tutorial deploys Qwen3-TTS and streams speech through a Gradio application.  ( 117 min )
    Generate images and video with vLLM-Omni on SageMaker AI – Part 2
    Deploy two generative media models from one AWS vLLM-Omni Deep Learning Container on Amazon SageMaker AI. Generate an image with FLUX.2-klein through real-time inference, then animate it into video with Wan2.1-VACE through asynchronous inference, and retrieve the MP4 from Amazon S3.  ( 118 min )
    Implementing synthetic monitoring using Amazon Nova Act
    Learn an agent-driven approach to synthetic monitoring using Amazon Nova Act and Amazon Bedrock AgentCore. The post covers the architecture and patterns for resilient, managed user-journey validation that moves beyond brittle UI scripts, with a complete sample implementation.  ( 120 min )
    Automating Amazon Textract adapter lifecycle management across accounts
    Learn how to operationalize Amazon Textract Custom Queries adapters for production: infrastructure as code with AWS CloudFormation and Terraform, a cross-account adapter promotion process, a pre-classification routing pattern for multiple form versions, and production security controls such as VPC endpoints, encryption, and least-privilege IAM.  ( 126 min )
  • Open

    SRE Weekly Issue #536
    View on sreweekly.com A message from our sponsor, Planetscale: PlanetScale is headed to SREcon26 in Dublin this October. Swing by our booth to grab some merch, catch a live demo, and chat with the team behind the world’s fastest databases. We can’t wait to see you there. → If you’re not at SREcon but still […]  ( 4 min )

  • Open

    Scaling MoE reinforcement learning on Amazon EKS with EFA and DeepEP with 40% more throughput
    Learn how to scale Mixture-of-Experts (MoE) reinforcement learning on Amazon EKS using Elastic Fabric Adapter (EFA) and DeepEP. This post presents an architecture that combines Amazon EKS, EFA, and Amazon S3 and increased aggregate reinforcement learning rollout throughput by 40% for large-scale RLHF and GRPO training.  ( 125 min )
    Accelerate multimodal RL training with SkyRL on Amazon SageMaker HyperPod
    Learn how to run SkyRL, an open-source reinforcement learning framework, on Amazon SageMaker HyperPod to post-train a Qwen3-VL-8B vision-language model with GRPO. This walkthrough covers building the container image, launching a Ray cluster from SageMaker Studio, submitting and monitoring the job, and hosting the trained LoRA adapter for inference.  ( 125 min )
    NarrateAI: production-ready LLM quality assurance on Amazon Bedrock
    NarrateAI delivers production-ready LLM quality assurance on Amazon Bedrock. This post details five techniques—adaptive pipeline orchestration, cross-account multi-model failover, real-time streaming evaluation, composite evaluation, and data accuracy verification—that reach about 99% numerical accuracy while streaming responses in real time.  ( 128 min )
    Deploying real-time personalized speech with Qwen3-TTS on Amazon SageMaker AI
    Deploy the publicly available Qwen3-TTS-12Hz-1.7B-Base text-to-speech model from Amazon SageMaker JumpStart to a fully managed, real-time endpoint, and clone a voice from a short reference clip. Cross-lingual cloning preserves the speaker's identity across languages.  ( 120 min )
    How Datacor built self-service rental analytics with Amazon Quick Sight
    Learn how Datacor built a self-service rental analytics experience for gas and welding distributors by embedding Amazon Quick Sight dashboards and natural language querying into its TrackAbout platform, powered by an automated cross-cloud data pipeline and multi-tenant row-level security.  ( 122 min )
    Multi-Region training with Amazon SageMaker HyperPod and Qumulo
    Amazon SageMaker HyperPod and Cloud Native Qumulo let you place training compute in one AWS Region while keeping your dataset in another. This post shares the architecture and validation results from a cross-Region training run, where a remote cluster matched a co-located cluster's throughput after a brief NeuralCache warmup.  ( 124 min )

  • Open

    Speaker-labeled transcription with WhisperX on SageMaker AI
    The AWS WhisperX Deep Learning Container packages Whisper, wav2vec2 forced alignment, and speaker diarization into a GPU-ready image. Learn how to deploy it to Amazon SageMaker AI real-time and asynchronous endpoints for word-level, speaker-labeled transcription, plus the production details that matter: the GPU AMI pin, scaling, and cost controls.  ( 121 min )
    Build a multi-account AI agent with AgentCore Gateway and MCP
    Build a multi-account architecture that keeps each team's data in its own AWS account while giving AI agents a unified way to query across them. A central platform account runs the agent using Amazon Bedrock AgentCore Gateway and MCP, while line-of-business accounts expose their data as MCP servers with secure cross-account access and fine-grained authorization.  ( 125 min )
    Aderant builds intelligent ticket triage with Amazon Nova
    Learn how Aderant built an intelligent ticket triage system on Amazon Nova Lite through Amazon Bedrock, automating context gathering, classification, routing, and knowledge enrichment for its cloud operations team.  ( 118 min )
  • Open

    Milo cancer diary part 25 – remission again (again)
    This is very much a repeat of part 22 ‘remission again’ from earlier in the year. The fifth protocol has worked, and once again everything went smoothly until the visit to NDSR for Epirubicin where we had to delay things a little whilst his neutrophils recovered. I should probably pay more attention to the treatment […]  ( 12 min )
    Milo cancer diary part 25 – remission again (again)
    This is very much a repeat of part 22 ‘remission again’ from earlier in the year. The fifth protocol has worked, and once again everything went smoothly until the visit to NDSR for Epirubicin where we had to delay things a little whilst his neutrophils recovered. I should probably pay more attention to the treatment […]  ( 13 min )

  • Open

    From portal-hopping to instant answers: HEMA’s journey with MCP and Amazon Bedrock
    HEMA, a 100-year-old Dutch retailer, turned developer portal-hopping into instant answers by building HAL, an internal AI assistant on Amazon Bedrock AgentCore. Using Model Context Protocol (MCP), HAL delivers governed knowledge inside the tools teams already use, with no AWS credentials on the client and security anchored in Microsoft Entra ID.  ( 121 min )
    Agentic conversational video intelligence built on AWS
    Learn how to build a conversational video intelligence solution on AWS using an agentic architecture. A single Strands Agents SDK agent orchestrates Amazon Bedrock, Amazon Rekognition, and Amazon Transcribe at runtime, deciding which service to call so you can ask natural language questions about your videos and get answers in seconds.  ( 126 min )
    Use open weight models as your AI coding agent with Amazon Bedrock
    Pair OpenCode, an open-source terminal-native AI coding agent, with open weight models on Amazon Bedrock to get a secure, flexible, pay-per-use coding assistant. Learn how to configure multi-model workflows, match the right model to each task, and keep your data in your own AWS account with no infrastructure to manage.  ( 123 min )

  • Open

    Bring more intelligence to everyday work with GPT-6 Sol and GPT-6 Luna on Amazon Bedrock
    GPT-6 Sol and GPT-6 Luna are now generally available on Amazon Bedrock, giving you more options to match intelligence and efficiency to each workload.  ( 116 min )
    Claude Opus 5.5 is now available on AWS
    Claude Opus 5.5, Anthropic's most capable Opus model for agentic coding, knowledge work, and long-running tasks, is now available on Amazon Bedrock and Claude Platform on AWS. This post covers what's new in Opus 5.5, practical guidance, and how to start building with the model on Amazon Bedrock.  ( 116 min )
    Evaluate skill-equipped agents with Strands Evals and Amazon Bedrock AgentCore
    Skills let you encode domain-specific procedures as reusable, portable instructions for agents, but a fluent answer doesn't prove the agent picked the right skill or followed it. Learn how to measure skill selection and instruction following with Strands Evals and Amazon Bedrock AgentCore Evaluations.  ( 122 min )
    How Reactiv automates mobile commerce 80% faster with Amazon Bedrock AgentCore
    Reactiv used Amazon Bedrock AgentCore to build a multi-agent AI Scheduler that autonomously refreshes Shopify merchants' mobile apps on a schedule, reducing merchant configuration time by 80% and getting to production 33% faster.  ( 119 min )
    Right-size generative AI endpoints with concurrency sweeps on Amazon SageMaker AI
    Concurrency sweeps help you right-size a generative AI endpoint on Amazon SageMaker AI by systematically benchmarking it at increasing load levels. This post walks through deploying a model, running automated concurrency sweeps with the CreateAIBenchmarkJob API, and using the results to make data-driven capacity decisions about fleet size.  ( 119 min )
    How Trane gets building insights 60x faster with Amazon Bedrock AgentCore
    In about four weeks, Trane Technologies built an AI-powered agentic solution on Amazon Bedrock AgentCore that reduced a 20-minute, multi-screen building diagnostic workflow to a 20-second natural language interaction, a 60x improvement in time-to-insight. This post shares the architectural approach and key design decisions behind the solution.  ( 122 min )
    How Tata Elxsi detects industrial safety risks in seconds on AWS
    Learn how Tata Elxsi built IRIS, a real-time industrial safety platform on AWS. IRIS filters camera video at the edge, streams metadata through Amazon Kinesis, runs computer vision on Amazon SageMaker AI, and correlates detections into high-confidence alerts, detecting unsafe conditions in seconds instead of minutes.  ( 120 min )
    Extending public sector intelligence with Agentforce and AWS
    Public sector agencies process large volumes of unstructured evidence, such as body camera footage and scanned documents. This post shows how to combine Amazon Bedrock Data Automation with the Model Context Protocol (MCP) to turn that data into structured insights and surface them through natural language queries in Salesforce Agentforce.  ( 122 min )

  • Open

    xAI’s Grok 4.6 is now available in Amazon Bedrock
    xAI's Grok 4.6 is now available in Amazon Bedrock: a frontier model for long-running agents, coding, and knowledge work, with a 500K token context window and four reasoning effort levels. It runs on both the bedrock-mantle and bedrock-runtime endpoints, with Converse API and cross-Region inference support.  ( 122 min )
    How BMW Group detects cost anomalies across 14,000 cloud accounts
    BMW Group operates CLEA, a FinOps platform monitoring more than 14,000 cloud accounts. This post shows how BMW added automated daily cost anomaly detection, moving from reactive dashboards to proactive alerts using Prophet forecasting, AWS Step Functions, and a serverless pipeline that processes every account for about $50 per month.  ( 121 min )
    Run Positron on Amazon SageMaker AI for data science workflows
    Positron, Posit's IDE for data science, now runs on Amazon SageMaker AI. This post shows how a data scientist explores an Amazon Athena table, validates features in R, trains an XGBoost model in Python, deploys a real-time SageMaker AI endpoint, and reports results with Quarto, all in one governed SageMaker Studio Space.  ( 120 min )
    How Benchling secured multi-tenant AI agents with Amazon Bedrock AgentCore
    Learn how Benchling built a defense-in-depth security architecture to run untrusted, AI agent-generated scientific code across thousands of life sciences tenants using Amazon Bedrock AgentCore Code Interpreter in VPC mode, combined with Amazon Route 53 Resolver DNS Firewall and VPC endpoint policies to block data exfiltration, including through DNS.  ( 122 min )
    Reducing medical claims review time with AI on AWS: The EXL Medical IDP solution
    EXL built an AI-powered Medical intelligent document processing (IDP) solution on AWS, combining IDP with domain-specific large language models on Amazon SageMaker and Amazon Bedrock to extract, summarize, and query medical records at enterprise scale and cut claims review time from over 100 minutes per case.  ( 122 min )
  • Open

    SRE Weekly Issue #535
    View on sreweekly.com A message from our sponsor, Planetscale: Neki brings horizontal sharding to Postgres. It handles the hard parts of running Postgres at scale: Online schema changes Version upgrades with no downtime Better connection pooling Resharding Graceful planned and unplanned failovers → Neki is available today. Learn more. Interning at incident.io: Rate Limiting, Resiliently […]  ( 4 min )

  • Open

    Amazon SageMaker Inference: 2026 year-to-date launches in review
    Amazon SageMaker AI shipped 13 inference launches in year-to-date across two deployment paths: fully managed endpoints and Amazon SageMaker HyperPod Inference. This post reviews each launch, from inference recommendations and capacity-aware instance pools to tiered KV caching and disaggregated prefill and decode.  ( 125 min )
    Introducing Kimi K3 on Amazon Bedrock
    Kimi K3 from Moonshot AI is now available on Amazon Bedrock, giving you a powerful new open-weight option for coding and knowledge work. It offers native vision, a 1-million-token context window, and explicit prompt caching to reduce latency and input costs.  ( 117 min )
    Migrating multi-model AI agents to Amazon Bedrock AgentCore runtime
    Migrate a multi-model healthcare AI agent from self-managed Amazon ECS with AWS Fargate to Amazon Bedrock AgentCore runtime, preserving triple-model orchestration and vector-enhanced knowledge retrieval while reducing infrastructure management. The framework-agnostic pattern applies across healthcare, financial services, and manufacturing.  ( 119 min )
    The new AgentCore runtime: Elastic, optimized, and consistently fast starts
    Today we are announcing the new AgentCore runtime, a capability of Amazon Bedrock AgentCore built for the speed, flexibility, and cost efficiency that production agents demand. It reclaims memory as sessions release it and delivers consistent cold starts regardless of image size or concurrency.  ( 121 min )
    Deploy Hugging Face models on Amazon SageMaker AI with coding agents
    Deploy production-ready Hugging Face models on Amazon SageMaker AI using six open-source agent skills. Point a coding agent at a model and get back a real-time endpoint with the right serving container, autoscaling, Amazon CloudWatch alarms, and a verified teardown path.  ( 121 min )
    Introducing Amazon SageMaker HyperPod Inference Gateway
    Amazon SageMaker HyperPod Inference Gateway is a Kubernetes-native, GPU-aware routing add-on for Amazon EKS. It uses real-time GPU signals to send each inference request to the best-suited pod, cutting first-token latency by up to 82% with no changes to your model servers or client applications.  ( 118 min )
  • Open

    The stories we tell
    TL;DR Large Language Model (LLM) based Artificial Intelligence (AI) may be great at cranking out plausible sentences. But stringing those sentences together to form some kind of narrative doesn’t get a good story. That’s a problem, and an opportunity; because humans run on stories, and (for now at least) it seems that we still need […]  ( 14 min )
    The stories we tell
    TL;DR Large Language Model (LLM) based Artificial Intelligence (AI) may be great at cranking out plausible sentences. But stringing those sentences together to form some kind of narrative doesn’t get a good story. That’s a problem, and an opportunity; because humans run on stories, and (for now at least) it seems that we still need […]  ( 14 min )

  • Open

    Reduce time-to-hire for quality candidates with AI-powered Amazon Connect Talent
    Amazon Connect Talent is an AI hiring solution built for talent acquisition leaders managing scaled hiring. It delivers AI-led interviews, data-driven assessments, and consistent evaluation, helping recruiters identify strong candidates more efficiently while providing applicants with a flexible interview experience. Informed by decades of Amazon's hiring science, Amazon Connect Talent provides transparency for every assessment, interview, and candidate score, enabling recruiters to stay in control of final hiring decisions.  ( 116 min )
    Selecting a vector store for Amazon Bedrock Knowledge Bases
    Choosing the right vector store for your Amazon Bedrock Knowledge Bases RAG application affects performance and cost. This post compares Amazon OpenSearch Service, Amazon Aurora PostgreSQL with pgvector, and Amazon S3 Vectors across three RAG use cases, with benchmarks and a practical selection framework.  ( 129 min )
    A serverless, data-driven Git metrics dashboard using Amazon Quick Sight
    Learn how to build a fully serverless pipeline that automatically collects Git metrics from GitHub and GitLab and visualizes them in interactive Amazon Quick Sight dashboards, giving engineering teams near-real-time delivery analytics at low cost.  ( 121 min )
    A shared agentic platform for Wood Mackenzie, on Amazon Bedrock AgentCore
    Wood Mackenzie built APEX, a shared agentic AI platform on Amazon Bedrock AgentCore so every team can ship production agents without rebuilding runtime, identity, observability, and guardrails from scratch. Learn why they chose AgentCore, how APEX Studio operates it, and where multi-agent systems go next.  ( 130 min )
    How MRH Trowe enabled secure self-service AI agents in financial services
    Learn how MRH Trowe, one of Germany's leading commercial and industrial insurance brokers, gave about 400 employees secure, self-service access to AI agents in its first month of production - using Strands Agents, Amazon Bedrock AgentCore, and LibreChat to meet the security, data residency, and compliance requirements of the German financial sector.  ( 119 min )
    Implementing defense-in-depth authorization for MCP tools on Amazon Quick
    Learn how to enforce defense-in-depth authorization for Model Context Protocol (MCP) tools on Amazon Quick. This walkthrough wires Microsoft Entra ID group and claims-based JWTs through an Amazon Bedrock AgentCore Gateway interceptor to apply per-user, per-tool role-based and attribute-based access control, with a server-side check and an immutable audit trail.  ( 128 min )
    Enhancing industrial safety AI with synthetic data on Amazon SageMaker AI
    Learn how to build a synthetic data augmentation pipeline on Amazon SageMaker AI and Amazon Rekognition that generates photo-realistic, auto-labeled training images for industrial safety AI. This approach improved person detection by up to 160% without manual annotation or hazardous data collection near heavy machinery.  ( 124 min )

  • Open

    Improving HCLS AI reasoning with open-source agent skills
    AI agents on foundation models often misapply healthcare and life sciences decision frameworks, citing the right guideline but applying it incorrectly. This post shares 38 open-source agent skills across 11 HCLS domains that close this gap, with installation steps, three worked use cases, and a 410-prompt evaluation showing a 70-86% win rate.  ( 125 min )
    Fault tolerant distributed training on Amazon EKS using NVRx
    Integrate NVIDIA Resiliency Extension (NVRx) into PyTorch FSDP training on Amazon EKS to overlap checkpoint I/O with training and recover from GPU faults in seconds. This post covers async checkpointing, in-process restart, and ft_launcher in-job restart, with H100 benchmarks at 2 to 8 nodes showing 99%+ training efficiency and second-scale recovery.  ( 123 min )
    Optimizing agent system prompts with Amazon Bedrock AgentCore
    AgentCore optimization turns production traces into proposed configuration changes, then validates them before promotion. This technical companion to the launch post explains how the system prompt optimizer's reflector engine works and shares benchmark results for the Single Agent and Sub-Agent Reflectors.  ( 119 min )
    Build a serverless PII redaction pipeline with Amazon Bedrock Data Automation
    Learn how to automate end-to-end PII detection and redaction from scanned documents at scale using Amazon Bedrock Data Automation with a custom blueprint, AWS Step Functions, and AWS Lambda. A custom blueprint redacts sensitive fields with field-level precision, and a token matching quality check raises recall across degraded and handwritten documents.  ( 122 min )

  • Open

    Optimizing cost and latency with Amazon Bedrock prompt caching
    Prompt caching in Amazon Bedrock can cut input token costs by up to 90% when you repeatedly send the same context to foundation models. This post walks through six practical prompt caching scenarios using the Converse API: message content, system prompt, tool definition, mixed TTL, tenant isolation, and LangChain integration.  ( 148 min )
    Build an AI-powered product tagging system with Amazon SageMaker serverless model customization
    Manually tagging thousands of catalog products is slow and inconsistent. This walkthrough shows how to customize Qwen3-8B with supervised fine-tuning (SFT) and reinforcement learning with verifiable rewards (RLVR) on Amazon SageMaker serverless model customization, then deploy it for asynchronous inference to build a cost-efficient product tagging system.  ( 120 min )
    Announcing instance preference lists for Amazon SageMaker AI training jobs
    Amazon SageMaker AI now offers instance preference lists for training and processing jobs. Specify an ordered list of up to five instance types, and SageMaker AI automatically launches on the first type with available capacity, eliminating manual retry loops and capacity-watching scripts.  ( 120 min )

  • Open

    Abnormal AI: Amazon Bedrock AgentCore for agentic email security at scale
    Learn how Abnormal AI deployed Amazon Bedrock AgentCore Code Interpreter as an ephemeral compute scratch pad for the agents behind its real-time email threat detection at billion-message scale, plus the sandbox design decisions and practical lessons for builders deploying Code Interpreter in production.  ( 117 min )
    Manage end-user OAuth consent for AI agents with Amazon Bedrock AgentCore
    Amazon Bedrock AgentCore Identity now offers a Consent portal, a managed web experience and session binding endpoint for AgentCore Gateway. This post walks through provisioning a portal, configuring GitHub and Slack 3LO targets, and the end-user consent flow, and shows how to review activity in AWS CloudTrail.  ( 121 min )
    How Ninth Wave built AI-powered open finance onboarding on Amazon Bedrock
    Learn how Ninth Wave built Compass, a multi-agent AI onboarding assistant on Amazon Bedrock AgentCore that validates bank APIs against Financial Data Exchange (FDX) standards, scores compliance, and compresses open finance onboarding from weeks to minutes while meeting SOC 2 and PCI DSS requirements.  ( 120 min )
    The generative AI customization spectrum: From prompt engineering to custom models on AWS
    Pick the right generative AI customization approach on AWS with an 8-step decision framework, from prompt engineering and RAG to fine-tuning, continued pre-training, and Amazon Nova Forge. Start simple and escalate only when you must.  ( 125 min )
    Automate replenishment with MMF, Databricks Genie, and Amazon Quick
    Foundation models made catalog-wide demand forecasting easy; the hard part is now acting on the forecast. This post builds a closed detect-decide-act loop on Databricks and Amazon Quick that reconciles demand surges against live supplier availability and places replenishment orders unattended, escalating to a human only when no supplier can cover a surge.  ( 125 min )
  • Open

    The impersonal AI advantage
    TL;DR Open source projects that I contribute to have adopted AI bots to ease the pull request (PR) process. It might just be me (though I doubt it), but somehow it feels less like a personal attack when a bot tells you what’s wrong with your PR than when it’s a person doing it. This […]  ( 13 min )
    The impersonal AI advantage
    TL;DR Open source projects that I contribute to have adopted AI bots to ease the pull request (PR) process. It might just be me (though I doubt it), but somehow it feels less like a personal attack when a bot tells you what’s wrong with your PR than when it’s a person doing it. This […]  ( 13 min )
  • Open

    SRE Weekly Issue #534
    View on sreweekly.com A message from our sponsor, Planetscale: Neki brings horizontal sharding to Postgres. It handles the hard parts of running Postgres at scale: Online schema changes Version upgrades with no downtime Better connection pooling Resharding Graceful planned and unplanned failovers → Neki is available today. Learn more. AI handles incidents, engineers lose touch […]  ( 4 min )

  • Open

    Monitoring production agent lifecycle with AWS DevOps Agent and AgentCore Evaluations
    Multi-agent systems fail in ways traditional monitoring misses. This post presents a dual-layer approach to monitoring production agents: Amazon Bedrock AgentCore Evaluations for continuous quality scoring and AWS DevOps Agent for autonomous infrastructure investigation, shown on a four-agent airline reservation system.  ( 129 min )
    Beyond the price per token: Choosing the right OpenAI model on Amazon Bedrock for your workload
    Comparing models on dollars per million tokens misses what production workloads actually pay for: outcomes. This post shares an open-source benchmarking harness that measures cost per correct answer, agent trajectory cost, and rubric-graded deliverable quality across OpenAI models on Amazon Bedrock.  ( 120 min )
    Build interactive MCP Apps using Amazon Bedrock AgentCore
    Learn how to build and deploy an MCP App with interactive HTML widgets on Amazon Bedrock AgentCore. Because MCP Apps is a host-agnostic standard, the same server delivers the same rich experience across AI hosts like ChatGPT and Claude that support the extension.  ( 121 min )

  • Open

    Reduce LLM latency with prefix-aware routing on Amazon SageMaker Inference
    Amazon SageMaker Inference now offers prefix-aware routing, a routing strategy that sends requests sharing the same prompt prefix to the same instance so the KV cache stays warm. In benchmarks on Llama 3.1 70B, it reduced P50 time-to-first-token by up to 77% and raised KV cache hit rates from about 25% to over 80%.  ( 119 min )
    Reduce inference cold starts on Amazon SageMaker HyperPod with model caching
    Amazon SageMaker HyperPod now supports model caching for inference, which pre-loads model weights and container images onto cluster nodes so pods read from local NVMe storage instead of downloading over the network. Learn how model caching cuts cold starts from tens of minutes to seconds, how it works, and how to enable it.  ( 119 min )
    Video and image search in Amazon Bedrock Knowledge Base using Marengo 3.0
    TwelveLabs Marengo Embed 3.0 is now generally available as an embedding model in Amazon Bedrock Knowledge Bases, bringing fully managed natural language search to video, image, and audio content. This walkthrough shows how to build a knowledge base powered by Marengo 3.0 and run semantic queries against your media.  ( 117 min )
    Amazon Quick is now generally available on desktop
    Your teams get an AI assistant that handles real work while your data stays in your environment and your conversations stay private Today, the Amazon Quick desktop application is generally available on macOS and Windows. We’re also adding a new activity feed to the mobile experience on iOS and Android that consolidates email, calendar, CRM, […]  ( 117 min )
    Build an end-to-end RFI questionnaire workflow using Amazon Quick Automate
    Learn how to build an end-to-end RFI questionnaire workflow with Amazon Quick Automate. Read a multi-tab RFI workbook from Amazon S3, use natural-language prompts to extract and structure the questionnaire data, refine the workflow through conversation, and write clean CSV output back to Amazon S3 — cutting development from days to hours.  ( 119 min )
    Model-agnostic PII detection with LLMs
    A configurable, model-agnostic detector that turns any large language model on Amazon Bedrock into a PII detector. Because the entities to detect live in a prompt rather than in code, one detector adapts to new entity types without retraining, and it outperforms an off-the-shelf tool across five public corpora and nine LLM-based detectors.  ( 123 min )
    Agent Evaluation Metric for multi-turn conversations
    Multi-turn agents fail in ways single-turn evaluation misses: one early mistake corrupts every later turn. This post introduces the Agent Evaluation Metric (AEM), a decomposable, turn-level way to measure agent quality, applied to its first dimension, correctness, to pinpoint the turn that caused a failure and separate it from the turns that inherited it.  ( 124 min )
    How AvioBook builds turnaround insights from operational data with Amazon Bedrock AgentCore
    AvioBook, a Thales Group Company, prototyped Connected Analytics on Amazon Bedrock AgentCore to turn AvioBook Connect's operational data into plain-language, evidence-based answers for airline managers and dispatchers, helping them find and act on the causes of flight turnaround delays.  ( 121 min )

  • Open

    Deploying Qwen3.8-2.4T-A95B on Amazon SageMaker HyperPod with vLLM
    Learn how to deploy Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter open-weight model, on Amazon SageMaker HyperPod with vLLM. This walkthrough covers cluster provisioning, NVFP4 quantization, and an OpenAI-compatible endpoint with built-in reasoning, tool calling, and native MTP speculative decoding.  ( 126 min )
    ICYMI: What landed for AI builders in August 2026
    A recap of August 2026 launches for AI builders across Amazon Bedrock, Amazon Bedrock AgentCore, and Strands: million-token context for OpenAI models, cross-Region inference, agents that run for up to 14 days on dedicated compute, expanded AWS GovCloud availability, and Strands Robots for physical deployment.  ( 117 min )
    How Heurist Finance built an AI-native investment workbench on Amazon Bedrock AgentCore
    Learn how Heurist built Heurist Finance, a conversational AI investment workbench, on Amazon Bedrock AgentCore. This customer story shows how AgentCore payments, Identity, Memory, Code Interpreter, and Observability let a small team buy premium market data per query, isolate analysis in a sandbox, and keep every action auditable.  ( 117 min )
    Simplify and support your TorchServe workloads using Ray Serve Deep Learning Containers
    TorchServe is no longer maintained, leaving teams to own the entire GPU inference stack. The AWS Ray Serve Deep Learning Container is a supported, pre-tested container with the framework, GPU drivers, and serving layer already assembled. This post walks through deploying a vision-language model on Amazon EKS using the Ray Serve DLC on a single GPU node.  ( 117 min )
    Automate user-level custom permissions for Amazon Quick
    Amazon Quick custom permissions let you enforce least-privilege access by toggling features per user. This post walks through four patterns to automate custom permissions across the user lifecycle: a RegisterUser API parameter, account and role defaults, event-driven Amazon EventBridge and AWS Lambda automation, and a retroactive batch update script.  ( 126 min )

  • Open

    Take on your most ambitious work with GPT-6 Astra on Amazon Bedrock
    GPT-6 Astra from OpenAI is now generally available on Amazon Bedrock. It brings deeper reasoning and sharper judgment to your most demanding tasks, running on the Amazon Bedrock inference engine built for high performance, security, and scale.  ( 116 min )
    Pathway’s brain-inspired architecture development on Amazon SageMaker HyperPod
    Pathway's Baby Dragon Hatchling (BDH) is a brain-inspired, post-transformer architecture that reasons in latent space instead of emitting chain-of-thought tokens. See how Pathway develops and scales BDH on Amazon SageMaker HyperPod, and how BDH-CQ set a new cost-efficiency mark on the ARC-AGI-1 benchmark.  ( 119 min )
    Pathway’s brain-inspired architecture development on Amazon SageMaker HyperPod
    Pathway's Baby Dragon Hatchling (BDH) is a brain-inspired, post-transformer architecture that reasons in latent space instead of emitting chain-of-thought tokens. See how Pathway develops and scales BDH on Amazon SageMaker HyperPod, and how BDH-CQ set a new cost-efficiency mark on the ARC-AGI-1 benchmark.  ( 119 min )
    Amazon SageMaker Feature Store introduces UpdateRecord for feature-level writes
    Amazon SageMaker Feature Store now supports feature-level writes. With the new UpdateRecord API, you can update one or more feature values in a single call without reading or rewriting the entire record. It is available for both the Standard (Amazon DynamoDB) and In-Memory (Amazon ElastiCache) online store tiers.  ( 120 min )
    Govern models with MLflow and Amazon SageMaker AI Model Registry sync: Part 2
    Governing models across accounts is the next step after automatic model registration. This post extends managed MLflow and Amazon SageMaker AI Model Registry sync to two cross-account governance topologies: a hub-and-spoke pattern that centralizes governance with AWS RAM, and a hybrid pattern that keeps development accounts isolated.  ( 121 min )
    Govern models with MLflow and Amazon SageMaker AI Model Registry sync: Part 1
    Managed MLflow on Amazon SageMaker AI now syncs richer model metadata (training metrics, evaluation results, inference specs, and lineage) into the SageMaker AI Model Registry, with lifecycle stage promotion. Part 1 shows how to govern candidate models in a single account using IAM guardrails.  ( 122 min )
    Automated agent evaluation with Amazon Bedrock AgentCore and GitHub Actions
    Wire Amazon Bedrock AgentCore Evaluations into a GitHub Actions pipeline: deploy an AI agent and an OAuth-protected MCP server to AgentCore runtime, invoke the agent with test prompts, score the responses, and automatically block pull requests when agent behavior regresses.  ( 125 min )
    Benchmarking small LLM inference on SageMaker AI: G7 vs G5 and G6
    Benchmark two 30B Mixture-of-Experts models, Qwen3-Coder-30B and NVIDIA Nemotron-3-Nano-30B, across G5, G6, G6e, and G7 GPU instances on Amazon SageMaker AI. Compare throughput, latency, and cost-per-token, and see how G7's NVIDIA Blackwell GPUs deliver measurable price-performance gains for real-time LLM inference.  ( 123 min )
    How HPE Zerto built an agentic troubleshooting system with Amazon Bedrock
    HPE Zerto built an agentic troubleshooting system powered by Amazon Bedrock that runs on-premises inside the customer environment. This post describes the multi-agent architecture, the on-premises deployment model built with Strands Agents, and the engineering challenges of grounding agents in live disaster recovery data.  ( 121 min )
    How DiDi built intelligent contact center QA with Amazon Bedrock
    DiDi built a transparent, self-owned contact center quality assurance (QA) system on Amazon Bedrock, replacing an opaque third-party tool. Intent verification accuracy rose from 38% to 86%, compliance scoring topped 90%, and Voice of Customer trend analysis dropped from hours to minutes across Spanish and Portuguese support.  ( 122 min )
  • Open

    球场上的镜像神经元
    球场上的镜像神经元 昨天打羽毛球,我从一个刚输给我们的对手身上学到一件事:复盘不能只盯着自己的失误,也要研究对手为什么打得好。 打完两局双打,对面打得更好的那位球友径直走到我面前,问我在羽毛球群里叫什么名字。真正让我意外的,是他接下来谈的不是输赢。 一场酣畅淋漓的混双 昨晚的这球场,我和一位公司招聘团队的之前就认识的女生搭档。 球场上,她落落大方,动作干净利索,在前场左右横移,敢扑敢压,根本不在意动作是否好看。我在后场调动和下压,她在前场抽挡、拦截。我们一边打球一边说笑,第一次搭档却配合得相当默契。 对面是两个男生。其中一位打得不错,球风犀利,落点也很聪明;另一位的球路则直来直往,更喜欢硬碰硬。 我们没有跟他们比蛮力,而是不断变换落点,把他们前后左右调动起来。有时他们被拉得前后踉跄,有时又不得不左右飞奔。等空当出来,我就在后场追身重杀,或者把球吊到网前,让他们来不及上网,只能望球兴叹。 两局下来,四个人都满头大汗。我们打得非常尽兴。对面那位相对弱一些的球友,因为拖了队友后腿,偶尔会尴尬地笑一笑。另一位却既没有不服,也没有憋屈。他下场后的反应,反而让我刮目相看。 不只复盘自己,也要观察对手 他走到我面前,先问我在羽毛球群里的名字,然后直接说出了自己的观点: 球场上要交流。既要跟队友交流,也要跟对手交流。 这也是他下场后马上来认识我的原因。 他接着说,打完球不能只反思自己哪个球没打好、下次应该怎么改,也要留意对手打出了哪些好球。为对手喝彩的同时,也是在学习对手:这个球为什么打得好?他的动作、意识和球路有什么特别?以后自己能不能模仿? 在他看来,提高不能只盯着自己。观察别人怎么打,同样是一条学习路径。 这番话让我意外。刚才在场上,我们一次次把他们调动得很被动。我原以为他下来后多少会有点懊恼,没想到他首先做的却是认识对手,主动交流,并且琢磨能从刚才的对抗中学到什么。 镜像神经元:观察也是一种练习 这让我联想起前段时间读的《智能简史》提到,即使只是观察别人完成动作,大脑也会对这个动作进行某种模拟。参与这一过程的神经元被称为“镜像神经元”(Mirror Neurons),也被认为与模仿有关。 这也让我重新理解那位球友的做法:盯住一个好球,观察对手的动作、意识和线路,在脑子里重新过一遍,再设法变成自己的东西。 想到这里这里,我又联想起自己曾经教孩子运动时常犯的错。 不能只是埋头苦练 每个孩子的运动天赋和学习节奏都不一样。看到别人一学就会,而自己的孩子练了很多遍,提高却不明显,大人很容易着急,甚至脱口而出:“为什么别人会,你还不会?”这种比较除了打击信心,并不能帮他找到问题。 我以前更容易把“多练”当作答案,觉得动作做得不够好,就再多做几遍。后来才意识到,不能只让孩子埋头苦练、独自琢磨。适当地让他看教程里的动作分解,看比赛视频中高手怎样移动、发力和处理球,再回到场上尝试,也可能带来提高。 当然,看懂不等于做到,观察也代替不了亲自练习。但练习并不只发生在自己挥拍的那一刻。认真看别人怎样完成一个动作,也是在给下一次尝试做准备。 这位球友让我看到,输掉一场球和学到一点东西并不矛盾。能从对手身上看到值得喝彩、值得模仿的地方,也许正是他越打越好的原因。  ( 1 min )

  • Open

    SRE Weekly Issue #533
    View on sreweekly.com A message from our sponsor, Planetscale: Most database incidents start with one expensive query, not the database being down. PlanetScale gives SRE teams high-availability Postgres and MySQL with automated failover, query insights, and Database Traffic Control to stop runaway queries before they page you. → Explore PlanetScale Incidents start before the response […]  ( 4 min )

  • Open

    Deploy a multimodal WhatsApp ordering assistant with Amazon Bedrock AgentCore
    Learn how to deploy a multimodal WhatsApp ordering assistant that takes customer orders through text, voice notes, and real-time voice calls on a single business number, built on Amazon Bedrock AgentCore with Amazon Nova 2. The channel and ordering layers stay separate, and one shared memory recognizes each customer across all three channels.  ( 126 min )
    Designing lifecycle policies for AgentCore memory
    Long-running AI agents accumulate outdated memories that degrade quality and create compliance risk. Learn how to design memory lifecycle policies for Amazon Bedrock AgentCore: scoring, consolidating, and pruning agent memories on a nightly AWS Step Functions workflow, with a deployable AWS CDK stack.  ( 124 min )
    Build a Physical AI model factory with NVIDIA Cosmos 3 on SageMaker HyperPod
    Building a Physical AI system takes a continuous pipeline, not a single training job. This post shows how to run that model factory (synthetic data generation, post-training, and closed-loop evaluation with NVIDIA Cosmos 3) on a persistent, resilient Amazon SageMaker HyperPod cluster on Amazon EKS, with GPU goodput as the metric that matters.  ( 137 min )
    Run agent-driven Amazon SageMaker HyperPod operations with InstantStart
    HyperPod InstantStart is an open source control plane that composes Amazon EKS orchestration with the managed capabilities of Amazon SageMaker HyperPod. It drives the same guarded operations through both a web interface and an AI agent, turning cluster bootstrap, capacity, training, inference, and storage into dependable, agent-driven infrastructure.  ( 129 min )
    Customizing your knowledge base on Amazon Bedrock for large and complex documents using Amazon Textract
    Learn how to customize an Amazon Bedrock knowledge base for large, complex documents by combining the high-accuracy text extraction of Amazon Textract with the generative AI of Amazon Bedrock. This post shows how to ingest and preprocess PDFs and images, then query utility bills at scale for faster, more accurate customer interactions.  ( 118 min )
    How Intuit built an agentic disaster recovery assistant with Amazon Bedrock
    Disaster recovery at scale is hard. Learn how Intuit built EWOK Agent, an agentic disaster recovery assistant on Amazon Bedrock that lets on-call engineers run production failovers from a plain-language request while keeping every action audited, policy-compliant, and safe.  ( 127 min )

  • Open

    August 2026
    Pupdate We were back in the Lake District (more on that later), so the boys were able to get out for some nice long walks. Going a month later than usual meant that there were no grouse on the estate but there were loads of ripe blackberries, which the boys enjoyed browsing. Sadly Milo’s scan […]  ( 15 min )
    August 2026
    Pupdate We were back in the Lake District (more on that later), so the boys were able to get out for some nice long walks. Going a month later than usual meant that there were no grouse on the estate but there were loads of ripe blackberries, which the boys enjoyed browsing. Sadly Milo’s scan […]  ( 15 min )
  • Open

    AI-driven development lifecycle using Amazon Bedrock AgentCore
    Engineering teams adopting the AI-Driven Development Lifecycle (AI-DLC) often struggle to turn concepts into working code. This post walks through two reference implementations on Amazon Bedrock AgentCore, Kiro, and Claude Code: an SQL-to-ER-diagram generator and a multi-agent code security analyzer that put the AI-DLC construction phase into practice.  ( 121 min )
    Migrate agentic workloads to Amazon Bedrock AgentCore
    An agent that works in a notebook is not an agent in production. This post walks through migrating a LangGraph customer support agent to Amazon Bedrock AgentCore in two stages: onto Runtime, Gateway, and Memory, then to model-driven planning on Strands Agents, retiring operational burdens along the way.  ( 126 min )
    Integrating Outlook with Amazon Quick for AI-powered email automation
    Integrate Microsoft Outlook with Amazon Quick to automate email management, calendar scheduling, and workflow coordination. This post walks through the end-to-end setup and shows automation scenarios using Amazon Quick chat agents, Amazon Quick Flows, and Amazon Quick Automate.  ( 117 min )
    Set up OpenAI ChatGPT Codex with LiteLLM on Amazon ECS and Amazon Bedrock
    Deploy a customer-operated LiteLLM gateway on Amazon ECS with AWS Fargate, connect it to an OpenAI model on Amazon Bedrock, and configure Codex to route requests through the gateway's Responses API with scoped identities, budgets, rate limits, and telemetry. We also compare direct IAM Identity Center access and a managed Portkey deployment.  ( 124 min )
    Best practices for building agentic automations with Amazon Quick Automate
    Learn best practices for building production-grade, agent-based business process automations with Amazon Quick Automate: choosing the right process, designing focused agents, combining them with deterministic steps, applying human-in-the-loop review, and building in evaluation and observability.  ( 121 min )
    Embed Quick Sight visuals using Cognito user authentication
    Learn how to embed individual Amazon Quick Sight visuals into a React application with per-user access control. This walkthrough uses Amazon Cognito authentication and a serverless AWS Lambda backend to generate scoped embed URLs, deployed with a single AWS CloudFormation stack.  ( 123 min )

  • Open

    Accessing OpenAI models on Amazon Bedrock from Australia with global cross-Region inference
    Australian teams can now access OpenAI GPT-5.6 Sol, Terra, and Luna models on Amazon Bedrock with global cross-Region inference from the Asia Pacific (Sydney) and Asia Pacific (Melbourne) Regions. This post shows how to invoke the models, use prompt caching, set up Codex with OpenID Connect authentication, and monitor usage with Amazon CloudWatch.  ( 120 min )
    Modernizing and scaling support operations with generative AI on AWS
    Learn how to build a generative AI-based support operations platform on AWS that converts training videos into structured SOPs, applies Retrieval-Augmented Generation to guide ticket resolution, and uses machine learning to predict SLA risk and prioritize work.  ( 127 min )
    How an AWS team detects dashboard content failures at scale using Amazon Bedrock
    Business intelligence dashboards can fail silently, showing blank, stale, or wrong data even when every infrastructure monitor reports healthy. Learn how an AWS team built an automated, AI-powered content validation solution on Amazon Bedrock that scans hundreds of dashboards and alerts owners, cutting mean time to detection from days to under an hour.  ( 120 min )
    From code to diagrams: Agentic architecture documentation with Amazon Bedrock AgentCore
    Learn how a global interdealer broker built an automated architecture documentation pipeline on Amazon Bedrock AgentCore that analyzes .NET code bases, generates architecture diagrams, and maintains searchable documentation through Amazon Bedrock Knowledge Bases and AWS CodePipeline.  ( 126 min )
    Trinity: Agentic AI-powered transition planning for students with disabilities
    Learn how University Startups and its AWS partner g/d/n/a scaled Trinity, a conversational AI solution for students with disabilities, into a serverless multi-agent architecture on Amazon Bedrock that produces IDEA-aligned transition plans for school districts across the US.  ( 118 min )
2026-10-01T08:17:38.702Z osmosfeed 1.15.1