When a coding agent works in your AWS account, CloudTrail records its calls as yours, and nobody can easily say what it did, whether any of it was risky, or what access it actually needed. Chaperone answers those three questions from records AWS already writes. It reads CloudTrail, groups every AWS API call under the MCP tool call that made it, flags risky calls with fixed rules, and measures the least-privilege policy from what the agent really used.

I built it in four days for the AWS Zero to Shipped hackathon, with Claude Code doing the building through the AWS MCP Server. So the first thing Chaperone ever recorded was its own construction. Every session on the live site is real.

Chaperone's flight path for the Day 1 build session: one lane each for AWS MCP server tool calls, Terraform, AWS CLI, SDK, AWS services and people, with six amber key icons marking iam:PutRolePolicy identity escalations over the Terraform lane
Figure 1: The Day 1 build session on Chaperone. 3,365 AWS calls, 27 MCP tool calls, six identity escalations (the key icons).

Links: live Chaperone · 3-minute tour · code on GitHub


Why does a coding agent need a flight recorder?

Because access to AWS is being handed to agents faster than anyone reviews what they do with it. Most teams give the agent the developer's own credentials, so in CloudTrail the agent and the human are the same person.

In May, AWS made the AWS MCP Server generally available, which means any MCP-capable agent can call AWS APIs directly. That's a good thing. But it moves the hard question from "can the agent do this?" to "what did the agent do?"

That question isn't hypothetical. In July 2025, a malicious prompt shipped inside version 1.84.0 of the Amazon Q Developer extension for VS Code, telling the agent to delete cloud resources. AWS says it failed to run and no customers were affected. Every team that had it installed still had to ask: did our agent run any of it? The answer is sitting in CloudTrail. It's just not in a shape a person can read.


How do you tell the agent apart from the developer?

Give the agent its own identity. That one decision is what makes everything else in Chaperone possible.

On Day 0 I created an IAM Identity Center user called chaperone-agent with a permission set that expires every four hours. The agent runs on a Linux VPS with no browser, so it signs in with a device code that I approve from my laptop:

aws sso login --profile chaperone-agent --use-device-code --no-browser
aws sts get-caller-identity --profile chaperone-agent
# → arn:aws:sts::111122223333:assumed-role/AWSReservedSSO_ChaperoneAgent_.../chaperone-agent

Then Claude Code gets the managed AWS MCP Server through the SigV4 proxy, locked to that profile:

claude mcp add-json aws-mcp '{"type":"stdio","command":"uvx",
  "args":["mcp-proxy-for-aws-cli@latest","https://aws-mcp.us-east-1.api.aws/mcp",
          "--metadata","AWS_REGION=us-east-1"],
  "env":{"AWS_PROFILE":"chaperone-agent"}}'
Claude Code's /mcp list showing aws-mcp Connected
Figure 2: aws-mcp connected in Claude Code, signed in as the agent's own identity.

Two rules came out of this. The human and the agent never share an identity, and the agent never holds a long-lived key. The four-hour expiry turned out to matter more than I expected: every working day started with the agent's session dead and me approving it again. That's annoying, and it's also exactly the point.

The identity paid off on Day 1. The first live run filed 558 of the agent's calls as generic SDK traffic, because Terraform's user agent doesn't mention Claude Code. With its own identity, that didn't matter: anything done by chaperone-agent is the agent's, whatever tool made the call. The user agent only tells you the channel (MCP, Terraform, CLI).


The discovery: one tool call, three levels in CloudTrail

The whole product rests on something I found on Day 0, about an hour in. I asked the agent to query CloudTrail for its own events, and one agent action showed up at three levels.

  1. The MCP tool call. An event from aws-mcp.amazonaws.com, named CallReadWriteTool, whose user agent names Claude Code and its version. It lists every AWS API call the tool made, with request IDs.
  2. Each AWS API call under it, marked invokedBy: aws-mcp.amazonaws.com, with the matching request ID.
  3. The agent's plain CLI calls, separately, with their own user agent.

Here's the shape of the first level, trimmed, simplified, and with identifiers replaced:

{
  "eventSource": "aws-mcp.amazonaws.com",
  "eventName": "CallReadWriteTool",
  "eventType": "AwsMcpEvent",
  "userAgent": "python-httpx/0.28.1 mcp-proxy-for-aws/1.7.0 claude-code/2.1.280",
  "requestParameters": {
    "method": "tools/call",
    "params": {
      "name": "aws___run_script",
      "arguments": "[HIDDEN_DUE_TO_SECURITY_REASONS]"
    }
  },
  "additionalEventData": {
    "downstreamRequests": [
      {
        "apiName": "s3:CreateBucket",
        "requestId": "EXAMPLE1",
        "awsRegion": "us-east-1"
      },
      {
        "apiName": "s3:PutObject",
        "requestId": "EXAMPLE2",
        "awsRegion": "us-east-1"
      },
      {
        "apiName": "cloudfront:CreateOriginAccessControl",
        "requestId": "EXAMPLE3",
        "awsRegion": "us-east-1"
      }
    ]
  }
}

Join that list to the API events on request ID and you get a tree: session → MCP tool call → AWS API call. CloudTrail holds every link. No AWS console shows the tree.

Three details in that event shaped the design:

  • Every tool call is readOnly: false, even GetCallerIdentity. So Chaperone judges risk from the API calls underneath, never from the tool call.
  • AWS hides the script the agent sent. Arguments come back as [HIDDEN_DUE_TO_SECURITY_REASONS]. Chaperone can show what the agent called, not the code it wrote.
  • MCP reveals calls the trail doesn't log. s3:PutObject is a data event, which a management-events trail never records. It appears only in the MCP event's downstream list. The agent's own record is more complete than the trail.

The first thing the agent built that night was a private S3 bucket behind CloudFront: two MCP run_script calls, 16 seconds of AWS API calls, live at chaperone.fullstackfusions.com before midnight. That hello-world became the first session Chaperone ever replayed.


Four minutes that changed the architecture

The plan was simple: an EventBridge rule catches CloudTrail events in near real time and a Lambda stores them. Before writing any of it, I gave the question a 30-minute timebox: does EventBridge deliver the MCP tool-call events at all?

The agent built a probe through MCP: an SQS queue and a rule matching every event from :chaperone-agent, read-only events included. Then one read and one write. Four minutes later the queue held six ordinary API events and zero MCP tool-call events, although LookupEvents showed both tool calls sitting in CloudTrail.

So ingest became a hybrid:

  • EventBridge forwards API calls from every enabled region to us-east-1 within seconds. (Every region came later, on Day 1, after the agent changed the trail in us-east-2 and the audit-tampering rule never saw it. CloudTrail delivers to EventBridge only in the region where the call happened.)
  • A poller runs every minute and fetches the MCP tool-call events with LookupEvents, which is free.

The two paths arrive in any order, so I stopped trying to assign events to sessions at write time. DynamoDB stores events per actor (PK = ACTOR#<actor>, SK = <time>#<KIND>#<eventID>). Sessions (a 30-minute idle gap), the tool-call join and the risk roll-up are pure functions that run when someone reads. Writes are idempotent, so EventBridge, the poller and a backfill can overlap safely, and every rule is testable on recorded events.

The MCP side of the model is small. From the backend:

def _tool_fields(r: dict) -> dict:
    """An MCP event. Tool arguments and results are hidden by AWS; what's left is the tool
    name and the AWS APIs it called."""
    req = r.get("requestParameters") or {}
    params = req.get("params") if isinstance(req.get("params"), dict) else {}
    downstream, verdicts = [], []
    for d in (r.get("additionalEventData") or {}).get("downstreamRequests") or []:
        service, _, action = (d.get("apiName") or ":").partition(":")
        v = rules.classify(service, action)
        verdicts.append(v)
        downstream.append({"api": d.get("apiName"), "request_id": d.get("requestId"),
                           "region": d.get("awsRegion"), "risk": v.risk})
    top = rules.worst(verdicts)
    return {"tool": params.get("name") or r.get("eventName"),
            "downstream": downstream, "risk": top.risk, "reasons": top.reasons}

The query layer later re-judges each call with its full request parameters once the API events are joined in. A bucket policy is only "public exposure" if it really grants Principal: * without a condition, not because the API is called PutBucketPolicy.

Chaperone architecture on AWS: CloudTrail in every region feeds EventBridge and a poller Lambda into an ingest Lambda with risk rules, stored in DynamoDB; an API Lambda answers using Cloud Control, IAM Access Analyzer and Amazon Bedrock; CloudFront serves the console from S3 and the API to visitors; IDE agents reach the API through the Chaperone MCP server
Figure 3: Chaperone as built. Capture, process, store, answer, serve. Everything is in Terraform.

What did the four days look like?

Each day had one job. The agent did the building inside AWS; I approved its sign-ins, made the decisions, handled DNS in Cloudflare, and pressed apply when the agent's guardrails said a human should.

Day What shipped The moment worth remembering
0 (Thu) Account hardened, CloudTrail on, agent identity, AWS MCP Server, hello-world on CloudFront with a custom domain The three-level discovery, and the EventBridge probe that came back without MCP events
1 (Fri) Terraform state in S3; Day 0 imported; ingest pipeline, risk rules, every region, review API, MCP server terraform plan: 12 to import, 0 to add, 0 to change. Terraform described what the agent had built by hand, exactly
2 (Sat) Public view of the API, CloudFront /api/*, React console with the flight path Live at the domain a day before the "live or drop" checkpoint
3 (Sun) Bedrock explanations per session, the judges' tour, prerendered pages The agent reviewing its own session on camera

Two things from those days changed how I think about working with an agent.

The guardrail stopped the agent at the right place. On Day 1, Claude Code's auto mode refused a terraform apply -auto-approve. From then on it was always plan -out=x.tfplan, read the plan, apply the file. On Day 2 it refused even the reviewed apply, as a protected infrastructure change. So the agent wrote and reviewed the plan, and I ran the apply. For a product about supervising agents, that's the right shape.

Real events made the best test fixtures. I pulled all 90 of Day 0's agent events and sanitized them into a fixture file. Doing that turned up two surprises: CloudTrail had logged 40 full session tokens in AssumeRoleWithSAML responses, and the Identity Center SAML issuer URL carried the account ID base64-encoded, so a plain-text search for the account ID missed it. Every rule since then is tested against recorded calls, not invented ones. The suite ended Day 3 at 116 tests.


What did the recorder find in its own build?

Its first finding was about itself. On Day 1, Chaperone flagged its own build session for identity escalation: the agent had put inline policies on the Lambda roles it just created. That flag is correct.

The Day 1 build session in Figure 1 is the clearest example. In 1 hour 16 minutes, the agent made 3,365 AWS calls and 27 MCP tool calls. Six calls were iam:PutRolePolicy from Terraform: three on roles that session had just created, three on roles that already existed. Those last three are the ones a reviewer should open. Chaperone marks risky calls on resources created in the same session (on_resource_created_this_session), because cleanup isn't damage and a policy on a brand-new role isn't the same as one on a role someone else depends on. It's context for the reviewer, not a lower risk class.

Least privilege, measured against AWS's own tool

The agent was allowed "Action": "*". That session used 65 actions. I wanted to know how Chaperone's measurement compared with the policy IAM Access Analyzer generates from CloudTrail, so I ran both on the same window.

Access Analyzer took 3 minutes 27 seconds; Chaperone's preview is instant from the same recorded events. The first diff had 14 differences. Tracing each one to its CloudTrail events found four Chaperone bugs:

  • Not found still means allowed. Terraform probes optional bucket settings and gets NoSuchCORSConfiguration on every plan. Dropping failed calls loses permissions Terraform needs. Now only access-denied calls are dropped.
  • Event names aren't IAM actions. ListBuckets is s3:ListAllMyBuckets; GetBucketReplication is s3:GetReplicationConfiguration.
  • AWS services act under your identity. Lambda encrypting environment variables logs kms:Decrypt as the agent, with invokedBy: lambda.amazonaws.com. The agent never called KMS, so those don't belong in its policy.

After the fixes, 72 of 75 actions matched. The rest is where each tool sees something the other can't:

  • Access Analyzer missed logs:FilterLogEvents: 18 successful calls inside the window. A policy built only from its output would break the agent's log tailing.
  • Both missed iam:PassRole. CloudTrail doesn't record it as an event, but both CreateFunction calls needed it. The role ARN is in the request, so Chaperone reads it from there and scopes PassRole to exactly those roles. Access Analyzer can't: it reads permissions, not requests.

So Chaperone shows both, side by side, and says what each one missed.


Can the agent review its own work?

Yes, and that's the part I'm proudest of. Chaperone ships an MCP server with five tools: list_sessions, what_did_the_agent_do, risky_calls, review_session and least_privilege. Pass "me" as the session and it resolves, from the IAM identity that signed the request, to the caller's own latest session. Before an agent says "done", it can ask what it actually did, from a record AWS wrote and the agent can't edit.

In the clip below, the agent creates and deletes a queue through the AWS MCP Server, then asks Chaperone about its own session:

Figure 4: The agent asks Chaperone about itself (30 seconds, muted).

The answer it reads back, trimmed:

{
  "tool": "aws___run_script",
  "apis": ["sqs:CreateQueue", "sqs:DeleteQueue"],
  "risk": "destructive",
  "risky": [
    {
      "action": "sqs:DeleteQueue",
      "target": ".../111122223333/chaperone-demo-queue",
      "via": "mcp",
      "on_resource_created_this_session": true
    }
  ]
}

One destructive call, on a resource the same session created. That's a clean bill of health, and the agent can say so with evidence.

The MCP server is one of three ways in. The same API serves a web console for people and signed HTTP calls for scripts and CI, so an agent, a reviewer and a pipeline always see the same answer.

Three ways into one Chaperone API: an HTTP API for scripts and CI signed with SigV4, an MCP server for the agent signed and unmasked, and a web console for people through CloudFront masked and read-only; the API uses Cloud Control, IAM Access Analyzer and Bedrock, and reads sessions built from CloudTrail via EventBridge, a poller, an ingest Lambda and DynamoDB
Figure 5: One API, three ways in. Only the masking differs.

How do you put real account data on a public website?

Mask the finished response, not the fields you remember. The public console shows real sessions from my AWS account, so the API needed a public mode that can't leak.

  • CloudFront marks every website request with an origin header (x-chaperone-view: public) that visitors can't override. A request that sends its own x-chaperone-view: private still gets masked data, because CloudFront's header wins.
  • Masking runs on the final JSON text, after everything else. It catches the account ID (plain and base64), the Identity Center role suffix, directory and identity store IDs, IPs, access keys, tokens, emails, and CloudFront and function URL IDs. A new field added next month can't slip past it.
  • The public view can't start work. id=me, Access Analyzer jobs and Bedrock generation all return 403. Visitors never cause a model call.
  • It serves a slim replay. The longest session's full timeline was 4.1 MB; the public replay is 873 KB, and about 102 KB on the wire with brotli.

On launch day I leak-scanned 46 live responses (1.6 MB) for every identifier shape. Nothing came back. That's why the account on the site reads 111122223333.


What surprised me along the way?

I kept a list of 41 gotchas. These are the ones most likely to save you an afternoon:

  1. The Bedrock catalog lists models your account can't call. Newer Anthropic and OpenAI models appeared in ListFoundationModels and then failed with "not available for this account". Service Quotas showed why: their tokens-per-minute quota was 0 on this account. Check quotas before designing around a model. I used Claude Sonnet 4.6, after submitting Anthropic's use-case form.
  2. A tiny probe can pass where the real call fails. A 5-token "hi" to Sonnet 4.6 answered; the real call with a system prompt was refused until the use-case form was on file. Probe with the real request shape.
  3. EventBridge can silently skip CloudTrail events. Twelve kms:Decrypt calls made for the agent never reached the rule. Don't treat EventBridge as complete: Chaperone reconciles with LookupEvents every few minutes. A fully paginated check across 17 regions then read agent 651/651.
  4. Widening an EventBridge pattern can build a feedback loop. Forwarding AWS-service events includes EventBridge assuming its own forwarder role. Exclude your own roles in the pattern and test it with TestEventPattern first.
  5. agent is a DynamoDB reserved word. SET agent = :a fails. It only showed up on the first agent event, because sign-in events skip that clause.
  6. CloudFront to a Lambda function URL needs two permissions: lambda:InvokeFunctionUrl and lambda:InvokeFunction, both scoped to the distribution. Older examples show only the first.
  7. Cloud Control's IAM prefix is cloudformation:, not cloudcontrol:. A policy with cloudcontrol:GetResource applies without complaint and denies every call. Access Analyzer's ValidatePolicy catches it.
  8. The MCP Python SDK 2.x renamed FastMCP to MCPServer (mcp.server.mcpserver). And only ToolError messages reach the model; any other exception shows up as "Error executing tool", so the agent can't tell an expired SSO session from a 404.
  9. A single-page app is invisible without JavaScript. curl of the live site returned an empty <div id="root">. I prerendered the overview and the tour so crawlers and AI assistants can read them.

What does Chaperone deliberately not do?

  • It never blocks, deletes or reverts anything. A recorder that acts would be one more agent to watch.
  • A model never decides risk. Fixed rules do, tested on real recorded events. Bedrock writes only the plain-English summary, once per session.
  • It shows what was called, not the code. AWS doesn't record the scripts an agent sends through MCP.
  • It's review, not prevention. API calls arrive within seconds; MCP tool calls within a few minutes.

If you want prevention, AWS gives you the pieces: the AWS MCP Server's IAM condition keys, permission boundaries, SCPs. Chaperone's least-privilege output is meant to feed exactly those.


By the numbers

15
sessions recorded, 9 by the agent
~6,800
AWS calls, 21 risky
65
actions used, with "Action": "*" allowed
$0.07
whole account, Sept 1–27

Numbers are as of Sep 27, 2026, from the live system and Cost Explorer. The month's bill stayed under a dollar because everything is serverless and on-demand: Lambda on arm64, DynamoDB on-demand with a throughput cap, CloudFront in front, and Bedrock called once per session, never per visitor.


How can you try it?

  1. Open the 3-minute tour.
  2. Open the session marked Start here: read the explanation, then the flight path, then click a key icon to see the call, the rule that flagged it and the MCP tool call above it.
  3. Check the masking: /api/replay?id=me returns 403 on the public site.
  4. Run it in your own account: the repository's README has the steps. You need a CloudTrail trail, Terraform, and an agent with its own identity.

The next step is a read-only pilot on one non-production account: the agent-activity review, and a least-privilege policy per agent role built from every session that role ran, not just one. If your team is letting coding agents work in AWS and you'd like to try it, I'd like to hear from you.


Further Reading

Built for the AWS Zero to Shipped hackathon, with Claude Code and the AWS MCP Server.