02 · Operational Intelligence

Ask "How is production doing right now?" and watch an agent hierarchy investigate through the Grafana MCP server — metrics, logs, and alert rules queried at runtime — diagnose, predict, recommend, act under a human's authority, and re-query to prove the fix worked.

The locked operational loop

OBSERVE → CORRELATE → DIAGNOSE → PREDICT → RECOMMEND → AUTHORIZE → ACT → VERIFY
   │           │           │          │          │           │        │       │
Grafana MCP  hypothesis  competing  time-to-  bounded    Studio   studio   Grafana MCP
(metrics,    validated   root-cause threshold remediation  Head   control  re-queried:
logs,        against     hypotheses (computed  plan       decides  plane   before/after
alerts)      evidence    w/ contra-  from real                             must improve
                         dictions    slopes)

Agent hierarchy, locked: Operational Executive → Metrics / Log / Trace Analysts and an Alert Scanner → Correlation Agent → Incident Diagnosis Agent → Risk/Prediction Agent → Remediation Planning Agent — each with scoped permissions.

What the track asked for, and where it lives

RequirementWhere
Grafana at runtime via the official Grafana MCP server (query_prometheus, query_loki_logs, list_alert_rules, list_datasources) app/tools/grafana/mcp_client.py; the hosted pod runs an mcp-grafana sidecar against Grafana Cloud with a service account
Gemini at runtimeapp/tools/google/gemini.py
Closed-loop actuation + verificationapp/agents/executive/executive.py — a recommendation is not done until the same signal is re-queried and has improved
Human authority boundary and permission tiersapp/governance/authority.py

Why it is built this way

The brief's thesis was "build an agent that uses observability data, and observe the agent you build." So the time-to-threshold prediction is computed from real slopes, not guessed; competing root-cause hypotheses are carried with their contradictions; and every investigation is itself observable. The Grafana annotation at the end is the human's review point, not the agent's victory lap.

Run it

uv venv .venv && uv pip install -p .venv -e ".[dev]"
.venv/bin/uvicorn app.main:app --port 8010     # MOCK mode: a deterministic render-pipeline incident, offline
cd ops && docker compose up -d && cd ..        # LIVE: Grafana + Prometheus + Loki + mcp/grafana + render-farm simulator