Skip to content

Latest commit

 

History

17 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Agents on Call

This repository contains the companion code for the "Agents on Call" series on ercan.ai. It documents a multi-agent platform for operational incident response, built on AWS Bedrock and the Strands Agents SDK.

The scenario: a mid-size B2B SaaS company with ~50 engineers across ~30 AWS accounts needs faster incident triage and response. Instead of paging an on-call human directly, automated agents handle the first pass: collecting context, checking runbooks, optimizing costs of affected workloads, and surfacing a decision tree to the human via Slack. The agents themselves are read-only by default; every action that mutates state sits behind an approval gate operated by humans.

This is an educational reference architecture, not a production-ready deployment as-is. The patterns for guardrails, approval gates, and read-only defaults generalize across organizations but require per-organization hardening before production use.

Architecture

AWS topology

  • Hub-and-spoke AWS Organizations. A management account holds SCPs and no workloads. A security account holds delegated admin for CloudTrail, GuardDuty, and Security Hub. Roughly 30 workload accounts are spokes, each exposing an ops-readonly role and a separate ops-mutate role to the ops-tooling account.
  • ops-tooling account is the entire platform. EventBridge and a Slack app (API Gateway) are the two entry points. AgentCore Runtime hosts four Strands agents: supervisor, incident-triage, runbook, and cost.
  • Model layer. Bedrock via cross-region inference profiles, with Bedrock Guardrails on every invocation and a Knowledge Base over the runbook corpus for the triage and runbook agents.
  • Tool layer is read-only by construction. AgentCore Gateway exposes four Lambdas: cloudwatch-read, logs-read, cost-read, and ssm-execute. The first three assume ops-readonly in the target spoke. ssm-execute is the only mutating tool, and it cannot reach a spoke directly.
  • Approval gate and observability. Mutating proposals pause in Step Functions on waitForTaskToken, post full context to Slack, and resume only on human approval; only the post-approval executor Lambda can assume ops-mutate. CloudWatch, S3 invocation logs, and AWS Budgets watch the platform itself.

The diagram above is a rendered snapshot (docs/img/topology.py, built with the diagrams library using official AWS icons). The Mermaid diagrams in docs/architecture.md are the source of truth and go account-by-account and layer-by-layer in more detail, including the Well-Architected mapping.

Blog series

All eight parts in reading order live at ercan.ai/series/agents-on-call.

# Date Title Link
1 2026-05-28 The Scenario: Why an Ops Team Hires Agents ercan.ai/agents-on-call-part-1-the-scenario
2 2026-06-04 The Foundation: Terraform Before Tokens ercan.ai/agents-on-call-part-2-terraform-before-tokens
3 2026-06-11 First Agent: Incident Triage in Strands ercan.ai/agents-on-call-part-3-incident-triage-strands
4 2026-06-18 Tools and the Gateway: MCP, Allowlists, Read-Only Default ercan.ai/agents-on-call-part-4-gateway-read-only-default
5 2026-06-25 The Team: Supervisor and Three Specialists ercan.ai/agents-on-call-part-5-supervisor-three-specialists
6 2026-07-02 Guardrails: The Part Everyone Skips ercan.ai/agents-on-call-part-6-bedrock-guardrails
7 2026-07-09 Sizing: Token Math Nobody Does Upfront ercan.ai/agents-on-call-part-7-sizing-token-math
8 2026-07-16 Production: Observability, Evals, and the Day It Lies ercan.ai/agents-on-call-part-8-observability-evals

Repository structure

terraform/
  00-foundation/       org accounts, spoke IAM roles, model access        -> Part 2
  10-agent-runtime/    AgentCore Runtime, all four agents + cost schedule -> Part 3, Part 5
  20-gateway-tools/    AgentCore Gateway, the four tool Lambdas           -> Part 4
  30-guardrails/       Bedrock Guardrails, denied topics, PII masking     -> Part 6
  40-observability/    OTEL, CloudWatch dashboards, Budgets, DLQs         -> Part 8

agents/
  supervisor/          routes to triage, runbook, cost                   -> Part 3, Part 5
  triage/               first agent, evidence-first incident triage      -> Part 3
  runbook/              runbook lookup against the Knowledge Base        -> Part 5
  cost_optimizer/      daily cost sweep agent                            -> Part 5
  tools/                shared tool implementations behind the Gateway   -> Part 4

docs/
  architecture.md       Mermaid source of truth for account and layer topology
  sizing-model.md       token math, quota ceilings, provisioned throughput breakeven -> Part 7
  img/topology.py, topology.png   this README's rendered AWS diagram

evals/                  golden-set eval harness, CI gate on agent behavior -> Part 8

Prerequisites

  • Terraform >= 1.9
  • Python >= 3.12
  • AWS account with Bedrock access (Claude 3 models)
  • Strands Agents SDK Python package

Each part is self-contained and can be deployed independently. Part 2 lays the foundation; Part 3 through 5 add agents and controls; Part 6 through 8 harden for production.

License

MIT, copyright 2026 Ercan Ermis.

About

This repository contains the companion code for the "Agents on Call" series on ercan.ai. It documents a multi-agent platform for operational incident response, built on AWS Bedrock and the Strands Agents SDK.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages