This is an AI agent that investigates incidents in a deterministic mock infrastructure environment — 4 services with fixed metrics, logs, deployments, and cost data. That content is genuinely live-queried from a real DynamoDB table on every tool call, not read from a local file — only the data's content is kept deterministic, on purpose, so the eval suite's exact-match scoring stays meaningful. It has 5 tools and decides which ones to use, and in what order, itself — nothing here is scripted.
Every tool call streams live below as it happens — you're watching the actual investigation, not a canned demo. It takes a few seconds per question because each step is a real round-trip to the model plus a real tool call.