This repository contains the companion code for the "Agents on Call" series on ercan.ai. It documents a multi-agent platform for operational incident response, built on AWS Bedrock and the Strands Agents SDK.
The scenario: a mid-size B2B SaaS company with ~50 engineers across ~30 AWS accounts needs faster incident triage and response. Instead of paging an on-call human directly, automated agents handle the first pass: collecting context, checking runbooks, optimizing costs of affected workloads, and surfacing a decision tree to the human via Slack. The agents themselves are read-only by default; every action that mutates state sits behind an approval gate operated by humans.
This is an educational reference architecture, not a production-ready deployment as-is. The patterns for guardrails, approval gates, and read-only defaults generalize across organizations but require per-organization hardening before production use.
- Hub-and-spoke AWS Organizations. A management account holds SCPs and no workloads. A security account holds delegated admin for CloudTrail, GuardDuty, and Security Hub. Roughly 30 workload accounts are spokes, each exposing an
ops-readonlyrole and a separateops-mutaterole to the ops-tooling account. - ops-tooling account is the entire platform. EventBridge and a Slack app (API Gateway) are the two entry points. AgentCore Runtime hosts four Strands agents: supervisor, incident-triage, runbook, and cost.
- Model layer. Bedrock via cross-region inference profiles, with Bedrock Guardrails on every invocation and a Knowledge Base over the runbook corpus for the triage and runbook agents.
- Tool layer is read-only by construction. AgentCore Gateway exposes four Lambdas:
cloudwatch-read,logs-read,cost-read, andssm-execute. The first three assumeops-readonlyin the target spoke.ssm-executeis the only mutating tool, and it cannot reach a spoke directly. - Approval gate and observability. Mutating proposals pause in Step Functions on
waitForTaskToken, post full context to Slack, and resume only on human approval; only the post-approval executor Lambda can assumeops-mutate. CloudWatch, S3 invocation logs, and AWS Budgets watch the platform itself.
The diagram above is a rendered snapshot (docs/img/topology.py, built with the diagrams library using official AWS icons). The Mermaid diagrams in docs/architecture.md are the source of truth and go account-by-account and layer-by-layer in more detail, including the Well-Architected mapping.
All eight parts in reading order live at ercan.ai/series/agents-on-call.
| # | Date | Title | Link |
|---|---|---|---|
| 1 | 2026-05-28 | The Scenario: Why an Ops Team Hires Agents | ercan.ai/agents-on-call-part-1-the-scenario |
| 2 | 2026-06-04 | The Foundation: Terraform Before Tokens | ercan.ai/agents-on-call-part-2-terraform-before-tokens |
| 3 | 2026-06-11 | First Agent: Incident Triage in Strands | ercan.ai/agents-on-call-part-3-incident-triage-strands |
| 4 | 2026-06-18 | Tools and the Gateway: MCP, Allowlists, Read-Only Default | ercan.ai/agents-on-call-part-4-gateway-read-only-default |
| 5 | 2026-06-25 | The Team: Supervisor and Three Specialists | ercan.ai/agents-on-call-part-5-supervisor-three-specialists |
| 6 | 2026-07-02 | Guardrails: The Part Everyone Skips | ercan.ai/agents-on-call-part-6-bedrock-guardrails |
| 7 | 2026-07-09 | Sizing: Token Math Nobody Does Upfront | ercan.ai/agents-on-call-part-7-sizing-token-math |
| 8 | 2026-07-16 | Production: Observability, Evals, and the Day It Lies | ercan.ai/agents-on-call-part-8-observability-evals |
terraform/
00-foundation/ org accounts, spoke IAM roles, model access -> Part 2
10-agent-runtime/ AgentCore Runtime, all four agents + cost schedule -> Part 3, Part 5
20-gateway-tools/ AgentCore Gateway, the four tool Lambdas -> Part 4
30-guardrails/ Bedrock Guardrails, denied topics, PII masking -> Part 6
40-observability/ OTEL, CloudWatch dashboards, Budgets, DLQs -> Part 8
agents/
supervisor/ routes to triage, runbook, cost -> Part 3, Part 5
triage/ first agent, evidence-first incident triage -> Part 3
runbook/ runbook lookup against the Knowledge Base -> Part 5
cost_optimizer/ daily cost sweep agent -> Part 5
tools/ shared tool implementations behind the Gateway -> Part 4
docs/
architecture.md Mermaid source of truth for account and layer topology
sizing-model.md token math, quota ceilings, provisioned throughput breakeven -> Part 7
img/topology.py, topology.png this README's rendered AWS diagram
evals/ golden-set eval harness, CI gate on agent behavior -> Part 8
- Terraform >= 1.9
- Python >= 3.12
- AWS account with Bedrock access (Claude 3 models)
- Strands Agents SDK Python package
Each part is self-contained and can be deployed independently. Part 2 lays the foundation; Part 3 through 5 add agents and controls; Part 6 through 8 harden for production.
MIT, copyright 2026 Ercan Ermis.
