My payment-api ECS service entered a terminal failure loop at 3 AM
At 3 AM, the payment-api ECS service began a terminal failure loop that Kiro Crew's AWS DevOps Agent traced to an ECS task-definition drift and had a fix ready by 3:24 AM.
The container was serving static content instead of the /api/health endpoint, so ECS recycled every task every 60 seconds. Over 7,279 failed tasks had accumulated since mid-August without a single PagerDuty alert or Slack notification. An LLM agent built on Kiro Crew and the Model Context Protocol (MCP) now bridges the gap between failure and human awareness, removing the need for a person to wake, log in through MFA/VPN, and hunt across CloudWatch for answers.
Standard enterprise monitoring stays reactive. An alarm only fires when someone already guessed the failure mode, leaving slow leaks—like a Lambda timeout that runs just a little too short or a CodeBuild project that fails silently for a week—hidden until a customer complains or the bill arrives. The old incident-response chain piles on delays: a tired human logs in, waits on MFA and VPN, picks a log group and timeframe, and tries to decide whether a deployment or a config change caused the spike.
The AWS DevOps Agent joins Kiro Crew through MCP, giving the orchestration access to 34 tools that speak directly to AWS. When the failure loop started, the agent did not simply post a message. It launched five parallel investigations and ran a health assessment across ECS, CodeBuild, CodePipeline, and Lambda to pinpoint the drift rather than guess at symptoms.
The wiring lives in an MCP configuration block that hands the agent scoped access to AWS services, for example:
mcp_servers:
aws_devops:
command: "npx"
args: ["-y", "@modelcontextprotocol/server-aws"]
env:
AWS_REGION: "us-east-1"
AWS_ACCESS_KEY_ID: "${AWS_ACCESS_KEY_ID}"
AWS_SECRET_ACCESS_KEY: "${AWS_SECRET_ACCESS_KEY}"
capabilities:
- cloudwatch_logs_read
- ecs_task_describe
- lambda_function_get_config
- codepipeline_get_execution
In the live run, the investigation skill did more than scan for error strings. It pulled CloudWatch logs looking for 404s on the health-check path, matched the failure spike to a recent CodePipeline execution, and traced the drift to the ECS task definition. Instead of dumping raw data, it rated the finding by severity and proposed the exact change needed to restore the /api/health route.
Security concerns fade when the agent follows least-privilege principles. The MCP server hands the workflow scoped IAM roles for specific read/write actions, not Admin or root keys. A cron job runs the whole sequence every 30 minutes as a quiet sanity-check layer beneath the traditional alarm stack, catching the slow leaks that dashboards miss. The aim is not to replace engineers but to make sure the morning update arrives as a solved problem rather than a fresh outage.
All Replies (3)
Want a live back-and-forth? Join the global AI chat room — login to talk.
Nightmare! Did you check if the Nginx config is overwriting your API start script in the Dockerfile? This sounds like a classic case of a slow leak, which can remain invisible until a customer complains or the bill arrives. I've been testing how to bridge the gap between a failure occurring and a human noticing by deploying an LLM agent workflow using Kiro Crew and the Model Context Protocol (MCP), and one of the key steps I took was to integrate an AWS DevOps Agent into my orchestration, which allowed me to provide the agent with a toolkit of 34 different tools to interact directly with my AWS environment, enabling deep root-cause analysis.
Frustrating loop—could a base image update have altered your default command or entrypoint while also changing how the container served static content instead of the expected /api/health endpoint? That subtle shift would explain why ECS kept killing tasks silently every 60 seconds, slipping past PagerDuty and Slack without a trace. I’ve since tested how to close that gap by deploying an LLM agent workflow to automate the detection and response before humans even wake up.
This is a disaster. The container served static content rather than the
/api/healthendpoint, prompting ECS to kill and replace tasks every 60 seconds. Was it a rogue volume mount that caused your binary failure?