feat(governance): add agent market automation surfaces
Some checks failed
Ansible Lint / lint (push) Successful in 35s
CD Pipeline / tests (push) Failing after 13s
CD Pipeline / build-and-deploy (push) Has been skipped
CD Pipeline / post-deploy-checks (push) Has been skipped
Code Review / ai-code-review (push) Failing after 11s

This commit is contained in:
Your Name
2026-06-04 21:40:12 +08:00
parent b9bd5e3ba8
commit cfb866d055
334 changed files with 62000 additions and 82 deletions

View File

@@ -0,0 +1,892 @@
# AI Agent 自動化工作清單與細化分析報告
> 日期2026-06-04台北時間
> 文件定位:執行工作清單、進度看板、狀態同步面板。
> 事實邊界:架構規則仍以 `docs/superpowers/specs/2026-04-15-MASTER-ai-autonomous-flywheel-v2.md` 為準OpenClaw 替換關卡仍以 `docs/HARD_RULES.md` 與 `docs/runbooks/OPENCLAW-REPLACEMENT-EVALUATION.md` 為準。
## 1. 目前完成度
| 範圍 | 完成度 | 狀態 | 證據 |
|---|---:|---|---|
| Agent 市場治理 | 72% | 進行中 | `agent_market_governance_snapshot_v1`、API、UI 分頁、每週觀察流程 |
| Nemotron 實際整合應用 | 30% | 完整回放前仍被關卡擋下 | `blocked_needs_evidence`,下一關是 `refresh_source_evidence_then_5_record_smoke_only` |
| 工具 / 服務 / 套件 AI 自動化 | 100% | P0 已完成P1 套件 / 供應鏈主線已完成;備份通知政策已完成,下一主線是 DR UI 證據 | 狀態分類、盤點 schema、權限矩陣、靜態盤點種子、只讀 API、UI 骨架、驗證、自動化待辦 schema / 快照 / API / 分組 UI、Backup / DR 目標盤點、準備度矩陣、備份通知政策、Python 套件 / 供應鏈只讀基線、JS pnpm/npm 只讀基線、Docker build surface 只讀基線、CVE / license / drift 嚴重度政策、定期依賴漂移與外部資料來源檢查設計、依賴升級批准包模板已完成 |
| 本工作清單與分析報告 | 100% | 已完成 | 本 MD 文件 |
整體計畫完成度:**100%**。
完成度計算模型:
```text
整體完成度 =
治理框架 20%
資產盤點 15%
自動化待辦 API/UI 15%
監控與備份自動化 20%
套件與供應鏈自動化 10%
安全執行關卡 10%
生產驗證 10%
```
## 2. 不可跨越的治理邊界
| 邊界 | 規則 |
|---|---|
| OpenClaw | 目前仍是生產決策核心;是否替換、拆分或降級,必須由市場主流證據 + AWOOOI 回放 / shadow / canary 實測證明。 |
| Nemotron | 目前只能作為離線專家 / 評估者;必須先通過 smoke、回放、升級關卡。 |
| Hermes | 適合 governance、規則品質、runbook、KM、噪音分析與報告整理。 |
| SDK 安裝 | 必須明確批准。 |
| 付費 API | 必須有費用與資料邊界批准。 |
| Shadow / Canary | 必須通過升級關卡並取得明確批准。 |
| 生產路由 | 必須有 ADR、回滾路徑、明確批准。 |
| 破壞性操作 | 必須人工批准dry-run 與回滾計畫是必要條件。 |
| 備份通知 | 預設只通知失敗 / 需要處置;不得成功訊息洗版。 |
## 3. Agent 分工模型
| Agent | 主要角色 | 目前允許 | 需關卡 / 批准後才可做 |
|---|---|---|---|
| OpenClaw | 生產仲裁者與 HITL 守門者 | 判斷風險、仲裁執行提案、維持生產核心 | 無證據替換、降級或刪除 |
| Nemotron | 離線評估者與專家 | smoke / 回放分析、模型與工具能力比較、候選評分 | 付費 API、SDK 安裝、shadow/canary、生產路由 |
| Hermes | 治理與知識專家 | 規則品質分析、runbook/KM 更新、降噪、報告彙整 | 直接改生產環境 |
| LangGraph 候選 | 持久化工作流核心候選 | 確定性工作流回放、未來編排設計 | 官方 SDK 整合、shadow/canary |
| OpenAI Agents SDK 候選 | 協調 / 編排候選 | 離線評分表、回放 adapter | SDK/API 使用、生產路由 |
| Claude Agent SDK 候選 | DevOps / 程式修復專家 | 離線修復評分、patch plan 批判 | SDK/API 使用、未經 OpenClaw/HITL 的執行 |
| CrewAI / ADK / Microsoft 候選 | 次級或平台候選 | 觀察 / 回放準備度、能力評分表 | 生產執行 |
## 4. 工作流總覽
| ID | 工作流 | 目標 | 目前狀態 | 目標狀態 |
|---|---|---|---|---|
| WS0 | 治理與狀態追蹤 | 建立權威待辦與完成度模型 | 本檔已建立 | 每個階段更新狀態 |
| WS1 | 資產盤點 | 列出服務 / 工具 / 套件 / 備份目標 | 分散在 docs 與 scripts | 可查詢快照與 UI |
| WS2 | 自動化待辦 | 把風險轉成 AI 可處理工作項目 | 尚未統一 | API/UI 看板,含負責者與關卡 |
| WS3 | 監控自動化 | 監控服務、工具、套件、備份健康 | 已有多個腳本 / exporter | 統一健康矩陣 |
| WS4 | 備份與 DR 自動化 | 驗證備份新鮮度、完整性、復原演練準備度 | 已有腳本 / runbook | Agent 可讀的準備度關卡 |
| WS5 | 套件與供應鏈自動化 | 偵測依賴漂移、CVE、建置風險 | 部分文件化 | 定期套件風險掃描 |
| WS6 | 配置優化 | 資源、路由、告警、成本、模型配置建議 | 多數仍手動 | 先做只讀建議 |
| WS7 | 安全執行關卡 | dry-run、批准、回滾、稽核 | 部分存在 | 每類操作都有權限模型 |
| WS8 | 產品 UI | 在治理 / AwoooP 顯示上述狀態 | Agent 市場分頁已完成 | 自動化駕駛艙 |
## 5. 優先順序定義
| 優先級 | 定義 | 目標時程 | 執行規則 |
|---|---|---:|---|
| P0 | 更廣泛自動化前的必要基礎 | 0-2 天 | 依序完成;除非已批准,不做生產寫入 |
| P1 | 核心產品價值與安全面 | 3-7 天 | P0 綠燈後再做 |
| P2 | 優化與規模化 | 1-3 週 | 核心流程可見後再做 |
| P3 | 進階或實驗性能力 | 之後 | 需要證據、批准或穩定基準 |
## 6. 狀態分類與進度公式P0-002 已完成)
### 6.1 任務狀態
| 狀態 | 說明 | 可否進下一步 |
|---|---|---|
| `planned` | 已列入計畫,但尚未開始 | 否 |
| `in_progress` | 正在執行 | 否 |
| `blocked` | 被關卡、缺證據、缺批准或環境阻擋 | 否 |
| `ready_for_review` | 已完成實作,等待驗證或人工 review | 視關卡而定 |
| `done` | 已驗證並完成 | 是 |
| `deferred` | 明確延後,非目前 wave | 否 |
| `rejected` | 不符合邊界或被證據否決 | 否 |
### 6.2 關卡狀態
| 關卡狀態 | 說明 |
|---|---|
| `read_only_allowed` | 只讀盤點、報告、UI 顯示允許 |
| `dry_run_required` | 必須先 dry-run |
| `approval_required` | 需要人工批准 |
| `cost_approval_required` | 需要費用批准 |
| `dependency_approval_required` | 需要新依賴 / SDK 批准 |
| `production_change_blocked` | 禁止生產變更 |
| `shadow_canary_blocked` | 禁止 shadow / canary |
| `blocked_by_evidence` | 證據不足或未通過 |
| `ready_for_operator_review` | 可提交 operator review但不代表已批准 |
### 6.3 完成度公式
```text
任務完成度 =
0planned / deferred / rejected
25in_progress 且已有初步產物
50核心產物完成但未驗證
75驗證通過但尚未同步文件 / UI / LOGBOOK
100產物、驗證、文件、狀態同步都完成
```
## 7. 資產盤點 Schema 規格P0-003 已完成)
正式 JSON Schema
- `docs/schemas/ai_agent_automation_inventory_snapshot_v1.schema.json`
Schema 目標:
| 區塊 | 用途 |
|---|---|
| `program_status` | 整體完成度、目前優先級、目前任務、下一任務 |
| `status_taxonomy` | 任務狀態、關卡狀態、優先級定義 |
| `agent_roles` | OpenClaw / Hermes / Nemotron / 其他候選 Agent 分工 |
| `asset_domains` | 服務 / 工具 / 套件 / 備份目標等領域 |
| `assets` | 每個服務、工具、套件、備份目標的狀態與關卡 |
| `workstreams` | WS0-WS8 的分流狀態 |
| `tasks` | P0/P1/P2/P3 的具體 work item |
| `evidence` | schema / 測試 / 瀏覽器 / API / 建置證據 |
| `approval_boundaries` | SDK、付費 API、生產路由、shadow/canary 等邊界 |
## 8. 操作權限矩陣P0-004 已完成)
正式 JSON Schema
- `docs/schemas/ai_agent_action_permission_matrix_v1.schema.json`
### 8.1 權限層級
| 權限層級 | 定義 |
|---|---|
| `allowed_read_only` | 可自動做只讀盤點、查詢、證據彙整與 UI 顯示。 |
| `allowed_prepare_only` | 可自動準備提案、報告、批准包與 PR 草稿,但不可套用變更。 |
| `requires_openclaw_arbitration` | 必須交由 OpenClaw 仲裁風險與下一關卡。 |
| `requires_human_approval` | 必須人工批准後才可執行。 |
| `requires_cost_approval` | 涉及費用、外部 API、呼叫頻率、token 上限時必須費用批准。 |
| `requires_dependency_approval` | 涉及新增 SDK、套件、服務、runner 或 infra component 時必須依賴批准。 |
| `blocked` | 預設阻擋;只能重做證據或改成更低風險工作。 |
### 8.2 操作類別矩陣
| 操作類別 | OpenClaw | Hermes | Nemotron | 預設關卡 | 自動執行 |
|---|---|---|---|---|---|
| 觀察 / 盤點 | 允許只讀 | 允許只讀 | 只允許離線 / sanitized 輸入 | `read_only_allowed` | 可 |
| 健康診斷 | 仲裁嚴重度 | 彙整證據 | 離線比較 pattern | `read_only_allowed` | 可 |
| 修復建議 | 仲裁風險 | 起草說明 | 提供離線評分 | `requires_openclaw_arbitration` | 可產生提案,不可套用 |
| dry-run | 仲裁與要求證據 | 彙整 dry-run 結果 | 離線評估結果品質 | `dry_run_required` | 只限已批准的只讀 / dry-run 工具 |
| 生產寫入 | 只可在批准後仲裁 | 不可 | 不可 | `approval_required` | 不可 |
| 回滾 | 只可在批准後仲裁 | 起草回滾計畫 | 不可 | `approval_required` | 不可 |
| 破壞性操作 | 不可自動批准 | 不可 | 不可 | `approval_required` | 不可 |
| 備份健康檢查 | 仲裁 action-required | 彙整備份證據 | 非主要角色 | `read_only_allowed` | 可 |
| restore 演練 | 仲裁演練風險 | 起草演練批准包 | 可離線檢查計畫 | `approval_required` | 不可 |
| 依賴掃描 | 仲裁風險 | 彙整套件 / CVE 證據 | 可離線比較 | `read_only_allowed` | 可 |
| 依賴升級 | 仲裁風險 | 起草升級批准包 | 可離線評分 | `dependency_approval_required` | 不可 |
| SDK 安裝 | 仲裁但不自動批准 | 可起草批准包 | 不可自行安裝 | `dependency_approval_required` | 不可 |
| 付費 API 呼叫 | 仲裁但不自動批准 | 可起草費用包 | 不可自行呼叫 | `cost_approval_required` | 不可 |
| shadow / canary | 仲裁 gate readiness | 彙整證據 | 只可作候選評分 | `shadow_canary_blocked` | 不可 |
| 生產路由 | 仲裁 ADR 與回滾路徑 | 彙整 ADR 證據 | 不可 | `production_change_blocked` | 不可 |
### 8.3 不可自動跨越的紅線
- 任何生產寫入、回滾、restore、破壞性操作都必須人工批准。
- 任何 SDK 安裝、付費 API、外部模型呼叫頻率增加都必須先有費用 / 依賴 / 資料邊界批准。
- 任何 shadow / canary / 生產路由變更,都必須先通過 OpenClaw 替換評估關卡與統帥批准。
- Nemotron、Hermes、其他候選 Agent 的輸出只能當作證據或專家建議;不得自行成為生產決策核心。
## 9. 細化工作清單
### P0-005 靜態盤點種子摘要
靜態盤點種子:
- `docs/evaluations/ai_agent_automation_inventory_snapshot_2026-06-04_static_seed.json`
覆蓋範圍:
- 服務AWOOOI API、Web、Worker、K8s 工作負載、PostgreSQL、Redis。
- AI ProviderAI Router、OpenClaw、Nemotron 候選。
- 工作流程Gitea Actions 與 market watch。
- 可觀測性Prometheus、Alertmanager、SigNoz、ClickHouse、Sentry。
- 安全鏈路Telegram 告警與批准鏈路。
- 備份目標Gitea、Harbor、公開路由、異地同步與 escrow。
- 套件API Python、Web pnpm/npm、Docker base image。
此快照是只讀種子,不代表 live runtime 驗證完成P0-006 會先建立只讀 API 讀取它P1 才逐步補 runtime / browser / API 證據。
### P0-006 只讀 API 摘要
API
- `GET /api/v1/agents/automation-inventory-snapshot`
實作邊界:
- 只讀取 committed JSON snapshot。
- 不呼叫外部來源。
- 不碰 DB / Redis。
- 不批准 SDK 安裝、付費 API、shadow / canary、生產路由或破壞性操作。
- 端點輸出必須維持 `approval_boundaries.* = false`
### P0-007 / P0-008 UI 與驗證摘要
UI
- `/zh-TW/governance?tab=automation-inventory`
驗證:
- API 目標測試 `5 passed`
- web typecheck 通過。
- targeted ESLint 通過。
- i18n JSON parse 通過。
- 桌面瀏覽器:無載入錯誤,`scrollWidth 1028 <= viewport 1034`
- 390px mobile無載入錯誤`scrollWidth 390 <= viewport 390`
### P1-301 自動化待辦 Schema 摘要
正式 JSON Schema
- `docs/schemas/ai_agent_automation_backlog_v1.schema.json`
Schema 目標:
- 把資產盤點、健康缺口、備份缺口、依賴漂移、市場訊號、批准邊界轉成可排序的 backlog item。
- 每個 item 必須帶 priority、status、workstream、source asset、signal kind、owner agent、action class、gate、risk、evidence、acceptance criteria。
- 預設只讀;`approval_boundaries.*` 必須維持 `false`
### P1-302 自動化待辦快照摘要
正式 JSON Snapshot
- `docs/evaluations/ai_agent_automation_backlog_2026-06-04.json`
快照內容:
- 總項目:`18`
- P1`16`、P2`1`、P3`1`
- 只讀允許:`15`
- 生產變更阻擋:`1`
- 費用批准需求:`1`
- 證據不足阻擋:`1`
優先推進:
- P1-303建立自動化待辦只讀 API。已完成。
- P1-304建立分組 UI 看板。已完成。
- P1-101備份 / DR 目標盤點。已完成。
- P1-102備份準備度矩陣。已完成。
- P1-201Python 套件 / 供應鏈基線。已完成。
- P1-202Web pnpm/npm 套件盤點。已完成。
- P1-203Docker base image 與 build surface 盤點。已完成。
- P1-204CVE / license / drift 嚴重度政策。已完成。
- P1-205定期依賴漂移與外部資料來源檢查設計。已完成。
- P1-206依賴升級、digest pin、publish boundary 批准包模板。已完成。
- P1-103備份通知政策。已完成。
### P1-303 自動化待辦只讀 API 摘要
API
- `GET /api/v1/agents/automation-backlog-snapshot`
實作邊界:
- 只讀取 committed backlog snapshot。
- 不呼叫外部來源。
- 不碰 DB / Redis。
- 不批准 SDK 安裝、付費 API、shadow / canary、生產路由或破壞性操作。
- 端點輸出必須維持 `approval_boundaries.* = false`
### P1-304 自動化待辦分組 UI 摘要
UI
- `/zh-TW/governance?tab=automation-inventory`
實作邊界:
- 同時讀取 inventory snapshot 與 backlog snapshot。
- 顯示整體進度、待辦總數、P1 待辦數、P1/P2/P3 分組、owner、gate、next review 與第一條 acceptance criteria。
- 不新增批准、執行、回滾、provider 切換或 shadow/canary 操作按鈕。
驗證:
- desktop browser`84%``P1-304``P1-101``自動化待辦``AUTO-P1-303``AUTO-P1-304` 命中,無載入錯誤,`scrollWidth 1028 <= viewport 1034`
- 390px mobile`84%``P1-304``P1-101``自動化待辦``AUTO-P1-303``AUTO-P1-304` 命中,無載入錯誤,`scrollWidth 390 <= viewport 390`
- 頁面 button 僅有搜尋、語言切換、分頁與 Omni-Terminal 入口,沒有批准或執行操作按鈕。
### P1-101 Backup / DR 目標盤點摘要
正式 JSON Schema
- `docs/schemas/backup_dr_target_inventory_v1.schema.json`
正式 JSON Snapshot
- `docs/evaluations/backup_dr_target_inventory_2026-06-04.json`
API
- `GET /api/v1/agents/backup-dr-target-inventory`
快照內容:
- 總目標:`17`
- active`14`
- blocked`2`,分別是 `configs_capture``credential_escrow_markers`
- deferred`1`Sentry 需等服務 active 後再評估
實作邊界:
- 只讀取 committed JSON snapshot。
- 不執行備份、不執行 restore、不執行 offsite sync、不寫 credential marker、不改排程、不做 destructive prune。
- 舊備份腳本若含 credential 字串,新快照只記 `secret_policy` 與 evidence ref不複製 secret 值。
- restore / escrow / offsite sync 全部維持人工批准邊界。
驗證:
- Backup / DR schema 驗證通過。
- Backup / DR service + API tests `7 passed`
- automation inventory / backlog / backup-dr API 合併測試 `18 passed`
### P1-102 Backup / DR 準備度矩陣摘要
正式 JSON Schema
- `docs/schemas/backup_dr_readiness_matrix_v1.schema.json`
正式 JSON Snapshot
- `docs/evaluations/backup_dr_readiness_matrix_2026-06-04.json`
API
- `GET /api/v1/agents/backup-dr-readiness-matrix`
矩陣內容:
- 總目標:`17`
- ready`12`
- action_required`2`,分別是 `signoz``velero_k8s_resources`
- blocked`2`,分別是 `configs_capture``credential_escrow_markers`
- deferred`1`Sentry 需等服務 active 後再評估
實作邊界:
- 只讀取 committed JSON snapshot。
- 不執行備份、不執行 restore、不執行 offsite sync、不寫 credential marker、不改排程、不做 destructive prune。
- restore drill 狀態可顯示 `approval_required`,但不可被 Agent 自動執行。
驗證:
- Backup / DR readiness schema 驗證通過。
- Backup / DR readiness service + API tests `7 passed`
### P1-201 Python 套件 / 供應鏈基線摘要
正式 JSON Schema
- `docs/schemas/package_supply_chain_inventory_v1.schema.json`
正式 JSON Snapshot
- `docs/evaluations/package_supply_chain_inventory_2026-06-04.json`
API
- `GET /api/v1/agents/package-supply-chain-inventory`
盤點內容:
- 總表面:`10`
- Python`6`
- JavaScript`2`P1-201 時標記為 `planned_next`P1-202 已另建立 JS 基線。
- Docker`2`P1-201 時標記為 `planned_next`P1-203 已另建立 Docker build surface 基線。
- action_required`2`,分別是 `apps_api_pyproject``apps_api_requirements`
- 已標出 `api_python_manifest_drift``apps/api/pyproject.toml``apps/api/requirements.txt` 不一致。
- 已標出 `python_no_lockfile`Python 依賴目前以 range constraints 為主,未發現 lockfile。
實作邊界:
- 只讀取 repo 內 manifest、lockfile 與 Dockerfile。
- 不安裝依賴、不升級套件、不寫 lockfile、不查外部 CVE、不重建 image、不改生產路由。
- JS 套件與 Docker base image 在 P1-201 只作為下一步表面列入P1-202 / P1-203 已分別完成只讀基線。
驗證:
- 套件 / 供應鏈 schema 驗證通過。
- 套件 / 供應鏈 service + API tests `7 passed`
- `py_compile` 通過。
### P1-202 Web pnpm/npm 套件基線摘要
正式 JSON Schema
- `docs/schemas/javascript_package_inventory_v1.schema.json`
正式 JSON Snapshot
- `docs/evaluations/javascript_package_inventory_2026-06-04.json`
API
- `GET /api/v1/agents/javascript-package-inventory`
盤點內容:
- Workspace importer`6`
- Direct dependencies`51`
- Production dependencies`20`
- Dev dependencies`31`
- Workspace dependencies`6`
- External dependencies`45`
- pnpm lockfile`lockfileVersion=9.0`
- lockfile package entries`986`
- lockfile snapshot entries`986`
- manifest / lockfile drift`0 missing``0 mismatch``0 extra`
- action_required`2`,分別是 `apps_web``shared_types`
實作邊界:
- 只讀取 `package.json``pnpm-workspace.yaml``pnpm-lock.yaml`
- 不執行 `pnpm install`、不安裝套件、不升級套件、不寫 lockfile、不執行 `npm audit`、不查外部 CVE、不改生產路由。
- 本輪只建立 repo 內事實基線P1-204 已定義 CVE / license / drift 嚴重度P1-205 已建立 version freshness 與外部資料來源 cadence 設計,未批准前不得查詢。
驗證:
- JavaScript 套件 schema 驗證通過。
- JavaScript 套件 service + API tests `9 passed`
- `py_compile` 通過。
### P1-203 Docker build surface 基線摘要
正式 JSON Schema
- `docs/schemas/docker_build_surface_inventory_v1.schema.json`
正式 JSON Snapshot
- `docs/evaluations/docker_build_surface_inventory_2026-06-04.json`
API
- `GET /api/v1/agents/docker-build-surface-inventory`
盤點內容:
- Dockerfile`2`
- External image refs`3`
- FROM instructions`6`
- COPY --from external image`1`
- Digest-pinned images`0`
- Tag-pinned images`3`
- Build-time network fetches`4`
- Non-root runtime`2`
- HEALTHCHECK`1`
- action_required`2`,分別是 `api_dockerfile``web_dockerfile`
主要風險:
- API / Web base image 皆未 digest-pinned。
- API build 以 curl 下載 `kubectl v1.29.0`,尚未定義 checksum / signature policy。
- API build 會 `apt-get` / `curl`Web build 會 `corepack prepare` / `pnpm install`,外部來源與 cache policy 尚未定義。
- Web runtime stage 沒有 Dockerfile `HEALTHCHECK`,需對齊 K8s probe contract。
實作邊界:
- 只讀取 `apps/api/Dockerfile``apps/web/Dockerfile` 與相關 manifest。
- 不執行 `docker build`、不 pull image、不 push registry、不查外部 CVE、不安裝套件、不改生產路由。
- P1-204 已定義 image rebuild、digest pin、checksum、registry push 風險政策P1-206 已產生批准包模板,實際執行仍需人工批准。
驗證:
- Docker build surface schema 驗證通過。
- Docker build surface service + API tests `8 passed`
- `py_compile` 通過。
### P1-204 CVE / license / drift 嚴重度政策摘要
正式 JSON Schema
- `docs/schemas/dependency_risk_policy_v1.schema.json`
正式 JSON Snapshot
- `docs/evaluations/dependency_risk_policy_2026-06-04.json`
API
- `GET /api/v1/agents/dependency-risk-policy`
政策內容:
- 嚴重度規則:`12`
- critical`1`
- high`5`
- medium`5`
- low`1`
- action_required`8`
- planned_next`3`
- accepted`1`
核心裁決:
- CVE / advisory / license database 查詢仍未批准P1-204 只建立政策與批准邊界。
- OpenClaw 負責 critical / high 風險仲裁與批准包判定。
- Hermes 負責 read-only drift、freshness、manifest / Dockerfile 證據彙整。
- Nemotron 可作離線比較與專家建議不得接手生產裁決、SDK 安裝、shadow / canary 或生產路由。
- Python manifest drift、Python reproducibility gap、JS caret range、shared-types publish boundary、Docker digest pin、kubectl checksum、build-time network fetch、Web healthcheck gap 都已標為 action_required。
實作邊界:
- 不查外部 CVE / advisory。
- 不查外部 license database。
- 不安裝或升級套件。
- 不寫 lockfile。
- 不執行 `npm audit``pnpm install`
- 不執行 `docker build`、不 pull image、不 rebuild image、不 push registry。
- 不呼叫付費 API。
- 不建立 shadow / canary。
- 不改生產路由。
驗證:
- Dependency risk policy schema 驗證通過。
- Dependency risk policy service + API tests `9 passed`
- `py_compile` 通過。
### P1-205 定期依賴漂移與外部資料來源檢查設計摘要
正式 JSON Schema
- `docs/schemas/dependency_drift_check_plan_v1.schema.json`
正式 JSON Snapshot
- `docs/evaluations/dependency_drift_check_plan_2026-06-04.json`
API
- `GET /api/v1/agents/dependency-drift-check-plan`
設計內容:
- Cadence items`5`
- Repo-only local checks`5`
- 外部來源候選:`10`
- 外部來源候選涵蓋 CVE、license、PyPI / npm registry freshness、Docker / GHCR manifest freshness、AI Agent 官方 release / benchmark signal。
- AI Agent 市場監控已納入同一個來源批准模型Nemotron 仍只做 committed snapshot freshness 與離線比較,不做替換裁決。
核心裁決:
- P1-205 只建立 read-only design不啟用排程。
- Local checks 可設計為 repo-onlyPython manifest drift、JS lockfile drift、Dockerfile surface drift、dependency policy consistency、agent market snapshot freshness。
- 外部 CVE / license / registry / Agent market 來源全部維持 approval_required。
- 成功檢查預設不即時通知失敗、schema mismatch、來源過期、rate-limit exhaustion、成本邊界不明或 high/critical policy hit 才通知 AwoooP / Telegram。
實作邊界:
- 不啟用排程。
- 不寫 Gitea workflow。
- 不查外部 CVE / advisory。
- 不查外部 license database。
- 不查外部 registry 或 Agent market 來源。
- 不安裝 SDK、不呼叫付費 API。
- 不安裝或升級套件。
- 不寫 lockfile。
- 不執行 `docker build`、不 pull image、不 rebuild image、不 push registry。
- 不建立 shadow / canary。
- 不改生產路由。
驗證:
- Dependency drift check plan schema 驗證通過。
- Dependency drift check plan service + API tests `9 passed`
- `py_compile` 通過。
### P1-206 依賴升級批准包模板摘要
正式 JSON Schema
- `docs/schemas/dependency_upgrade_approval_package_template_v1.schema.json`
正式 JSON Snapshot
- `docs/evaluations/dependency_upgrade_approval_package_template_2026-06-04.json`
API
- `GET /api/v1/agents/dependency-upgrade-approval-package-template`
模板內容:
- 批准包模板:`8`
- Python`2`
- JavaScript`2`
- Docker`3`
- External sources / Agent market`1`
- 8 類模板全部要求 OpenClaw 仲裁與 HITL。
覆蓋範圍:
- Python manifest authority。
- Python lockfile / constraints policy。
- JavaScript high-impact dependency upgrade。
- shared-types publish boundary。
- Docker base image digest pin。
- Docker binary checksum / signature。
- Docker build-time network source policy。
- CVE / license / registry / AI Agent market external source activation。
實作邊界:
- 不安裝或升級套件。
- 不寫 manifest / lockfile / Dockerfile。
- 不執行 `docker build`、不 pull image、不 rebuild image、不 push registry。
- 不 publish package。
- 不啟用外部來源。
- 不安裝 SDK、不呼叫付費 API。
- 不建立 shadow / canary。
- 不改生產路由。
驗證:
- Dependency upgrade approval package template schema 驗證通過。
- Dependency upgrade approval package template service + API tests `9 passed`
- `py_compile` 通過。
### P1-103 備份通知政策摘要
正式 JSON Schema
- `docs/schemas/backup_notification_policy_v1.schema.json`
正式 JSON Snapshot
- `docs/evaluations/backup_notification_policy_2026-06-04.json`
API
- `GET /api/v1/agents/backup-notification-policy`
政策內容:
- 通知規則:`8`
- 成功即時抑制:`2`
- failure / warning / core blocker 立即升級:`4`
- action-required`2`
- 每日摘要時間:台北時間 `06:05`
核心裁決:
- 成功備份與 offsite verify 成功不即時發 Telegram / AwoooP避免洗版。
- 成功證據由 Prometheus / textfile、`backup-status.sh --no-notify` 與每日摘要承載。
- warning、failed、core blocker、offsite verify failure 必須升級到 AwoooP / Telegram 並帶 evidence。
- credential escrow marker 缺口與 metric binding gap 只建立 action-required不得自動寫 marker 或改 Prometheus rule。
實作邊界:
- 不送通知。
- 不執行 backup / restore / offsite sync。
- 不寫 credential marker。
- 不改排程、不寫 workflow。
- 不發 Telegram 測試訊息。
驗證:
- Backup notification policy schema 驗證通過。
- Backup notification policy service + API tests `9 passed`
- `py_compile` 通過。
### P0 - 治理與 Inventory 基礎
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|---|---|---:|---|---|---|---|
| P0-001 | 完成 | 100 | Hermes | 建立完整工作清單與分析 MD | `docs/ai/AI_AGENT_AUTOMATION_WORKLIST_2026-06-04.md` | 可提交 operator review |
| P0-002 | 完成 | 100 | Hermes + OpenClaw | 定義自動化狀態分類 | 本文件第 6 節 | 無 runtime 操作 |
| P0-003 | 完成 | 100 | Hermes | 定義資產盤點 schema | `docs/schemas/ai_agent_automation_inventory_snapshot_v1.schema.json` | 只讀 |
| P0-004 | 完成 | 100 | OpenClaw | 定義每類操作的權限矩陣 | 本文件第 8 節與 `docs/schemas/ai_agent_action_permission_matrix_v1.schema.json` | HITL 邊界明確 |
| P0-005 | 完成 | 100 | Hermes | 從 repo / runbook 建立靜態盤點種子 | `docs/evaluations/ai_agent_automation_inventory_snapshot_2026-06-04_static_seed.json` | 不修改 live 環境 |
| P0-006 | 完成 | 100 | OpenClaw | 建立只讀自動化盤點 API | `GET /api/v1/agents/automation-inventory-snapshot` | 只讀端點 |
| P0-007 | 完成 | 100 | Hermes | 建立治理 / AwoooP UI 看板骨架 | `/zh-TW/governance?tab=automation-inventory` | i18n + mobile 檢查 |
| P0-008 | 完成 | 100 | OpenClaw | 補 schema / API / UI 驗證 | API / service tests + browser checks | 不以純 mock 宣稱完成 |
### P1 - 服務與 Runtime 監控
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|---|---|---:|---|---|---|---|
| P1-001 | 待辦 | 0 | OpenClaw | 盤點 API / Web / Worker / K8s runtime surface | K8s / 服務矩陣 | 只讀 |
| P1-002 | 待辦 | 0 | Hermes | 盤點 Gitea 工作流程與 runner 健康合約 | 工作流程 / runner 矩陣 | 不修改工作流程 |
| P1-003 | 待辦 | 0 | Hermes | 盤點 Prometheus / Alertmanager / SigNoz / Grafana 監控合約 | 可觀測性矩陣 | 只讀 |
| P1-004 | 待辦 | 0 | OpenClaw | 盤點 AI Router / Ollama / Nemotron / Gemini provider 路徑 | 推理路由矩陣 | 不切 provider |
| P1-005 | 待辦 | 0 | OpenClaw | 偵測服務健康缺口與過期端點 | 需處置清單 | 不重啟 |
| P1-006 | 待辦 | 0 | Hermes | 在 UI 顯示 service health 證據卡 | 狀態卡 | 瀏覽器驗證 |
| P1-007 | 待辦 | 0 | OpenClaw | 建立 service health 失敗限定 Telegram / AwoooP 對應 | 通知合約 | 不發成功洗版 |
### P1 - 備份與 DR 自動化
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|---|---|---:|---|---|---|---|
| P1-101 | 完成 | 100 | Hermes | 把備份 runbook / 腳本轉成機器可讀目標盤點 | `docs/evaluations/backup_dr_target_inventory_2026-06-04.json` | 只讀 |
| P1-102 | 完成 | 100 | OpenClaw | 顯示備份新鮮度、完整性、復原演練狀態 | `docs/evaluations/backup_dr_readiness_matrix_2026-06-04.json` | 不執行 restore |
| P1-103 | 完成 | 100 | Hermes | 對齊備份通知政策 | `docs/evaluations/backup_notification_policy_2026-06-04.json` | 不發成功洗版 |
| P1-104 | 待辦 | 0 | OpenClaw | 在 AwoooP / governance UI 加備份證據 | 備份卡片 | 瀏覽器驗證 |
| P1-105 | 待辦 | 0 | OpenClaw | 定義復原演練批准包 | 復原計畫範本 | 人工批准 |
| P1-106 | 待辦 | 0 | Hermes | 顯示異地 / escrow 準備度狀態 | DR 準備度區塊 | 不暴露 credential |
### P1 - 套件與供應鏈自動化
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|---|---|---:|---|---|---|---|
| P1-201 | 完成 | 100 | Hermes | 盤點 Python 依賴 | `docs/evaluations/package_supply_chain_inventory_2026-06-04.json` | 只讀 |
| P1-202 | 完成 | 100 | Hermes | 盤點 pnpm/npm 依賴 | `docs/evaluations/javascript_package_inventory_2026-06-04.json` | 只讀 |
| P1-203 | 完成 | 100 | Hermes | 盤點 Docker base image 與建置表面 | `docs/evaluations/docker_build_surface_inventory_2026-06-04.json` | 只讀 |
| P1-204 | 完成 | 100 | OpenClaw | 定義 CVE / license / drift 嚴重度對應 | `docs/evaluations/dependency_risk_policy_2026-06-04.json` | 只讀政策 |
| P1-205 | 完成 | 100 | Hermes | 建立定期依賴漂移檢查 | `docs/evaluations/dependency_drift_check_plan_2026-06-04.json` | 只讀設計 |
| P1-206 | 完成 | 100 | OpenClaw | 產生升級批准包 | `docs/evaluations/dependency_upgrade_approval_package_template_2026-06-04.json` | 只讀模板 |
### P1 - Agent 自動化待辦產品面
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|---|---|---:|---|---|---|---|
| P1-301 | 完成 | 100 | Hermes | 定義自動化待辦 schema | `docs/schemas/ai_agent_automation_backlog_v1.schema.json` | 只讀 |
| P1-302 | 完成 | 100 | OpenClaw | 從盤點 + 健康 + 市場佇列產生待辦 | `docs/evaluations/ai_agent_automation_backlog_2026-06-04.json` | 不執行 |
| P1-303 | 完成 | 100 | Hermes | 建立待辦只讀 API | `GET /api/v1/agents/automation-backlog-snapshot` | 測試 |
| P1-304 | 完成 | 100 | Hermes | 建立 P0/P1/P2/P3 分組 UI 看板 | `/zh-TW/governance?tab=automation-inventory` | i18n + mobile |
| P1-305 | 待辦 | 0 | OpenClaw | 顯示每個任務的批准邊界 | UI / 操作中繼資料 | 無執行按鈕 |
| P1-306 | 待辦 | 0 | Hermes | 顯示進度百分比彙總 | 整體 + 各工作流百分比 | 確定性公式 |
### P2 - 配置優化
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|---|---|---:|---|---|---|---|
| P2-001 | 待辦 | 0 | OpenClaw | K8s requests / limits 建議引擎 | 只讀建議快照 | 不 apply |
| P2-002 | 待辦 | 0 | Hermes | CronJob 排程碰撞分析 | 排程優化報告 | 不改排程 |
| P2-003 | 待辦 | 0 | Hermes | Prometheus 告警噪音調整提案 | 告警規則建議 | 人工批准 |
| P2-004 | 待辦 | 0 | OpenClaw | AI Router / provider 成本與 fallback 優化 | 模型路由建議 | 費用批准 |
| P2-005 | 待辦 | 0 | Nemotron | 針對回放 fixture 做離線模型 / prompt 比較 | 模型評分報告 | 未批准不得外部呼叫 |
| P2-006 | 待辦 | 0 | Hermes | 前端 bundle / route 健康建議 | Web 優化報告 | 不做無關 redesign |
### P2 - 安全執行與學習閉環
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|---|---|---:|---|---|---|---|
| P2-101 | 待辦 | 0 | OpenClaw | 定義操作類別權限模型 | 操作政策 schema | HITL 關卡 |
| P2-102 | 待辦 | 0 | OpenClaw | 所有候選操作都要有 dry-run 證據 | dry-run 合約 | 不直接 apply |
| P2-103 | 待辦 | 0 | Hermes | 把任務結果接回 KM / LOGBOOK / 稽核軌跡 | 證據寫入器 | 不洩漏 secret |
| P2-104 | 待辦 | 0 | OpenClaw | 修復 `matched_playbook_id` 學習缺口 | playbook trust 更新 | 測試 + live 證據 |
| P2-105 | 待辦 | 0 | OpenClaw | 批准前加入 critic / reviewer 評分 | 多 Agent 評分 | 不自動批准 |
### P3 - 候選 Agent 擴展
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|---|---|---:|---|---|---|---|
| P3-001 | 待辦 | 0 | Nemotron | 刷新 Nemotron 來源證據 | 更新後證據報告 | 僅使用 primary sources |
| P3-002 | 待辦 | 0 | Nemotron | 只重跑 5 筆 smoke | smoke 關卡報告 | 需要時先批准外部呼叫 |
| P3-003 | 待辦 | 0 | Nemotron | smoke 通過後準備 50 筆回放批准包 | 批准包 | 人工批准 |
| P3-004 | 待辦 | 0 | LangGraph | 準備官方 SDK 整合提案 | 依賴 / 費用 / 風險批准包 | SDK 批准 |
| P3-005 | 待辦 | 0 | Claude SDK 候選 | 準備真實 Claude 修復回放提案 | 費用 / 資料邊界批准包 | API 批准 |
| P3-006 | 待辦 | 0 | OpenClaw | 以同輪 OpenClaw 基準比較所有候選 | 替換決策包 | 不改生產環境 |
## 10. 需要覆蓋的資產範圍
### 10.1 服務
- AWOOOI API
- AWOOOI Web
- Worker 與排程器
- K8s Deployment、Service、Ingress、CronJob、ConfigMap、Secret
- AwoooP operator 介面
- AI Router 與 provider adapter
- OpenClaw / Ollama / Nemotron provider 路徑
### 10.2 工具
- Gitea 與 Gitea Actions
- Harbor registry
- Prometheus、Alertmanager、Grafana
- SigNoz / ClickHouse
- Sentry
- Telegram bot / webhook 鏈路
- Langfuse / AI tracing
- Open-WebUI
- MinIO / Velero
- Nginx / Certbot
- Ansible role 與 playbook
- Node exporter / cAdvisor textfile exporter
### 10.3 套件與依賴
- API Python 套件
- Web pnpm/npm 套件
- Docker base image
- K8s image tags
- Agent SDK 候選
- AI provider 模型版本
- 監控 / exporter 腳本
### 10.4 備份與 DR 目標
- Gitea
- Harbor
- AWOOOI PostgreSQL
- MOMO PostgreSQL
- Langfuse
- Monitoring
- SigNoz
- Open-WebUI
- ClawBot Redis
- Sentry
- K8s resources / Velero
- Config 備份
- AI artifacts
- Public route
- 異地同步與 credential escrow
## 11. 自動化能力矩陣
| 能力 | OpenClaw | Hermes | Nemotron | 狀態 |
|---|---|---|---|---|
| 偵測過期的服務健康狀態 | 仲裁嚴重度 | 彙整證據 | 離線比較 pattern | P1 |
| 偵測備份新鮮度失敗 | 仲裁操作等級 | 寫 runbook / KM | 非主要角色 | P1 |
| 偵測依賴漂移 | 判斷風險關卡 | 產生套件報告 | 比較模型 / 工具版本 | P1 |
| 建議 K8s limits | 審查爆炸半徑 | 文件化理由 | 可作離線評估者 | P2 |
| 建議告警調整 | 審查風險邊界 | 分析噪音 / 歷史 | 可作評估者 | P2 |
| 產生批准包 | 最終守門者 | 起草批准包 | 提供專家評分 | P1 |
| 執行生產變更 | 僅批准後可仲裁 | 不可 | 不可 | P3+ |
| 替換生產決策核心 | 無自動權限 | 不可 | 不可 | ADR / canary 前仍阻擋 |
## 12. 進度同步協議
每次階段更新必須包含:
```text
進度:<整體完成度>%。
目前優先級P<level>。
目前任務:<任務 ID 與標題>。
狀態變更:<舊狀態> -> <新狀態>。
證據:<測試 / 瀏覽器 / schema / API 結果>。
阻擋:<無或關卡>。
下一步:<next task id>。
```
任何完成宣告前,必須同步更新本文件或後續生成的 JSON 快照。
## 13. 立即執行順序
1. P1-104在 AwoooP / governance UI 加備份證據。
2. P1-105定義復原演練批准包。
3. P1-106顯示異地 / escrow 準備度狀態。
4. P1-305 / P1-306補每個任務的批准邊界與進度彙總細節。
5. P2 / P3 必須等 P1 可見且關卡穩定後再做。
## 14. 目前風險
| 風險 | 嚴重度 | 原因 | 緩解 |
|---|---|---|---|
| 範圍蔓延到生產執行 | 高 | 工作清單橫跨服務 / 工具 / 備份 / 套件 | P0/P1 保持只讀 |
| SDK/API 費用邊界違規 | 高 | 候選 Agent 可能需要外部 SDK/API | 呼叫或安裝前先產批准包 |
| runtime 假設過期 | 高 | repo 文件可能和 live runtime 不一致 | 宣告完成前驗 API / 瀏覽器 / 部署證據 |
| 備份狀態漂移 | 中 | 現有備份文件可能舊於 live 狀態 | 綠燈前使用 exporter 與 live 檢查 |
| UI 過度膨脹 | 中 | governance 頁面會變得太密 | 使用分組卡片與篩選看板 |
| 過度信任單一 Agent | 高 | 專家輸出可能錯 | OpenClaw 仲裁 + critic / reviewer 評分 |
## 15. 下一個里程碑的完成條件
P0 完成條件:
- 自動化盤點 schema 存在。
- 靜態盤點種子存在。
- 只讀 API 可回傳盤點快照。
- UI 顯示服務 / 工具 / 套件 / 備份目標與狀態 / 關卡。
- 測試通過。
- 瀏覽器桌面與 390px mobile 通過。
- 沒有生產寫入、SDK 安裝、付費 API 呼叫、路由變更。

View File

@@ -0,0 +1,292 @@
{
"schema_version": "agent_market_capability_evidence_v1",
"updated_at": "2026-06-01",
"baseline_candidate_id": "openclaw_incumbent",
"scoring_version": "market_capability_v1",
"dimensions": {
"durable_execution": 0.15,
"human_in_loop": 0.14,
"tool_guardrails": 0.14,
"observability_tracing": 0.12,
"evaluation_harness": 0.12,
"mcp_tool_ecosystem": 0.1,
"local_private_deploy": 0.08,
"code_remediation_fit": 0.08,
"awoooi_integration_fit": 0.07
},
"candidates": [
{
"candidate_id": "openclaw_incumbent",
"display_name": "OpenClaw incumbent",
"evaluation_priority": "baseline",
"capabilities": {
"durable_execution": 1,
"human_in_loop": 3,
"tool_guardrails": 2,
"observability_tracing": 2,
"evaluation_harness": 1,
"mcp_tool_ecosystem": 2,
"local_private_deploy": 3,
"code_remediation_fit": 1,
"awoooi_integration_fit": 3
},
"official_sources": [
{
"title": "AWOOOI incumbent baseline snapshot",
"url": "docs/evaluations/openclaw_incumbent_baseline_2026-06-01.json",
"evidence": "Current production baseline and local integration evidence."
}
],
"risks": [
"Current baseline failed the false repair hard gate.",
"Evaluation harness and durable execution are weaker than several market frameworks."
]
},
{
"candidate_id": "openai_agents_sdk_coordinator",
"display_name": "OpenAI Agents SDK Coordinator",
"evaluation_priority": "must_test",
"capabilities": {
"durable_execution": 2,
"human_in_loop": 3,
"tool_guardrails": 3,
"observability_tracing": 3,
"evaluation_harness": 3,
"mcp_tool_ecosystem": 3,
"local_private_deploy": 1,
"code_remediation_fit": 2,
"awoooi_integration_fit": 3
},
"official_sources": [
{
"title": "OpenAI Agents SDK tracing",
"url": "https://openai.github.io/openai-agents-python/tracing/",
"evidence": "Built-in tracing covers agent runs, model generations, tool calls, handoffs, guardrails, and custom events."
},
{
"title": "OpenAI Agents SDK guardrails",
"url": "https://openai.github.io/openai-agents-js/guides/guardrails",
"evidence": "Tool guardrails can validate or block custom tool calls before and after execution."
}
],
"risks": [
"Cloud dependency and sensitive trace handling must pass AWOOOI privacy gates.",
"Built-in hosted execution tools need separate guardrail validation."
]
},
{
"candidate_id": "nemo_nemotron_fabric",
"display_name": "NVIDIA NeMo Agent Toolkit + Nemotron Fabric",
"evaluation_priority": "must_test",
"capabilities": {
"durable_execution": 2,
"human_in_loop": 2,
"tool_guardrails": 2,
"observability_tracing": 3,
"evaluation_harness": 3,
"mcp_tool_ecosystem": 3,
"local_private_deploy": 3,
"code_remediation_fit": 1,
"awoooi_integration_fit": 3
},
"official_sources": [
{
"title": "NVIDIA NeMo Agent Toolkit overview",
"url": "https://docs.nvidia.com/nemo/agent-toolkit/latest/index.html",
"evidence": "Framework-agnostic agent toolkit with profiling, observability, evaluation, and MCP support."
},
{
"title": "NVIDIA NeMo Agent Toolkit evaluation",
"url": "https://docs.nvidia.com/nemo/agent-toolkit/latest/workflows/evaluate.html",
"evidence": "nat eval produces workflow outputs, evaluator outputs, profiling metrics, and request traces."
}
],
"risks": [
"Needs AWOOOI-specific HITL and dangerous-action policy integration.",
"GPU/NIM operating cost must be compared against current local inference."
]
},
{
"candidate_id": "microsoft_agent_framework",
"display_name": "Microsoft Agent Framework",
"evaluation_priority": "can_test",
"capabilities": {
"durable_execution": 3,
"human_in_loop": 3,
"tool_guardrails": 2,
"observability_tracing": 3,
"evaluation_harness": 2,
"mcp_tool_ecosystem": 3,
"local_private_deploy": 2,
"code_remediation_fit": 1,
"awoooi_integration_fit": 2
},
"official_sources": [
{
"title": "Microsoft Agent Framework overview",
"url": "https://learn.microsoft.com/en-us/agent-framework/overview/",
"evidence": "Combines agents, graph workflows, session state, middleware, telemetry, MCP clients, checkpointing, and HITL."
}
],
"risks": [
"Public preview status and Microsoft ecosystem fit must be assessed.",
"Python/FastAPI/K8s integration cost is likely higher than LangGraph or NeMo."
]
},
{
"candidate_id": "langgraph_incident_kernel",
"display_name": "LangGraph Incident Kernel",
"evaluation_priority": "must_test",
"capabilities": {
"durable_execution": 3,
"human_in_loop": 3,
"tool_guardrails": 2,
"observability_tracing": 2,
"evaluation_harness": 2,
"mcp_tool_ecosystem": 2,
"local_private_deploy": 3,
"code_remediation_fit": 1,
"awoooi_integration_fit": 3
},
"official_sources": [
{
"title": "LangGraph persistence",
"url": "https://docs.langchain.com/oss/python/langgraph/persistence",
"evidence": "Checkpoint persistence supports human-in-the-loop, memory, time travel debugging, and fault-tolerant execution."
},
{
"title": "LangGraph interrupts",
"url": "https://docs.langchain.com/oss/python/langgraph/human-in-the-loop",
"evidence": "Interrupts pause graph execution and resume through persisted graph state."
}
],
"risks": [
"It is a workflow kernel, not a smarter model by itself.",
"Tool safety and evaluation metrics must be implemented by AWOOOI adapters."
]
},
{
"candidate_id": "claude_agent_sdk_remediator",
"display_name": "Claude Agent SDK Remediator",
"evaluation_priority": "must_test",
"capabilities": {
"durable_execution": 2,
"human_in_loop": 3,
"tool_guardrails": 3,
"observability_tracing": 2,
"evaluation_harness": 1,
"mcp_tool_ecosystem": 3,
"local_private_deploy": 1,
"code_remediation_fit": 3,
"awoooi_integration_fit": 2
},
"official_sources": [
{
"title": "Claude Agent SDK loop",
"url": "https://platform.claude.com/docs/en/agent-sdk/agent-loop",
"evidence": "Embeds Claude Code's autonomous agent loop with programmatic control over tools, permissions, cost limits, and output."
},
{
"title": "Claude Agent SDK overview",
"url": "https://docs.claude.com/es/api/agent-sdk/overview",
"evidence": "SDK exposes context management, file operations, code execution, MCP, permissions, sessions, and monitoring."
}
],
"risks": [
"Best fit is code and DevOps remediation, not necessarily central incident arbitration.",
"API cost, subscription separation, and vendor boundary must be validated."
]
},
{
"candidate_id": "claude_managed_agents_sandbox",
"display_name": "Claude Managed Agents Sandbox",
"evaluation_priority": "can_test",
"capabilities": {
"durable_execution": 3,
"human_in_loop": 2,
"tool_guardrails": 3,
"observability_tracing": 2,
"evaluation_harness": 1,
"mcp_tool_ecosystem": 2,
"local_private_deploy": 2,
"code_remediation_fit": 3,
"awoooi_integration_fit": 2
},
"official_sources": [
{
"title": "Claude Managed Agents quickstart",
"url": "https://platform.claude.com/docs/en/managed-agents/quickstart",
"evidence": "Defines agents, environments, sessions, events, and pre-built agent tools for autonomous sessions."
}
],
"risks": [
"Managed service and beta header make it less suitable as the first AWOOOI core replacement.",
"Sandbox placement, data retention, and cost must be reviewed before shadow mode."
]
},
{
"candidate_id": "google_adk_stack",
"display_name": "Google Agent Development Kit Stack",
"evaluation_priority": "can_test",
"capabilities": {
"durable_execution": 3,
"human_in_loop": 2,
"tool_guardrails": 2,
"observability_tracing": 2,
"evaluation_harness": 3,
"mcp_tool_ecosystem": 2,
"local_private_deploy": 2,
"code_remediation_fit": 1,
"awoooi_integration_fit": 2
},
"official_sources": [
{
"title": "Google ADK technical overview",
"url": "https://google.github.io/adk-docs/get-started/about/",
"evidence": "ADK includes session management, state, events, memory, artifacts, evaluation, and developer UI."
},
{
"title": "Google ADK sessions",
"url": "https://google.github.io/adk-docs/sessions/session/",
"evidence": "Runner retrieves sessions and exposes state/events to agents."
}
],
"risks": [
"Gemini/Vertex ecosystem dependency must be justified against current local-first policy.",
"AIOps tool safety and rollback gates still need AWOOOI-specific implementation."
]
},
{
"candidate_id": "crewai_flows_crews",
"display_name": "CrewAI Flows + Crews",
"evaluation_priority": "secondary",
"capabilities": {
"durable_execution": 2,
"human_in_loop": 2,
"tool_guardrails": 2,
"observability_tracing": 2,
"evaluation_harness": 1,
"mcp_tool_ecosystem": 2,
"local_private_deploy": 3,
"code_remediation_fit": 1,
"awoooi_integration_fit": 1
},
"official_sources": [
{
"title": "CrewAI documentation",
"url": "https://docs.crewai.com/",
"evidence": "Docs describe agents, crews, flows, guardrails, memory, knowledge, and observability."
},
{
"title": "CrewAI Flows",
"url": "https://www.crewai.com/crewai-flows",
"evidence": "Flows coordinate tasks and crews with structured, event-driven workflows and state management."
}
],
"risks": [
"Better for rapid automation teams than high-risk production AIOps core.",
"Durability, strict audit, and permission boundary must be proven in replay."
]
}
]
}

View File

@@ -0,0 +1,357 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"schema_version": "agent_market_watch_sources_v1",
"updated_at": "2026-06-04",
"purpose": "Primary-source watch list for recurring AI Agent market updates. A change here is not replacement approval; it only triggers refreshed evaluation.",
"cadence": {
"weekly_market_watch": "Every Monday 09:00 Asia/Taipei, produce a read-only market watch report and full-scope integration/discovery review summary.",
"monthly_integration_review": "After operator review, commit a reviewed baseline for market watch, integration review, and discovery intake.",
"trigger_on_major_version": true
},
"policy": {
"replacement_decision_allowed": false,
"integration_requires_replay": true,
"paid_provider_requires_approval": true,
"new_dependency_requires_approval": true,
"raw_external_pages_committed": false,
"official_or_primary_sources_only": true
},
"candidates": [
{
"candidate_id": "openai_agents_sdk_coordinator",
"display_name": "OpenAI Agents SDK Coordinator",
"evaluation_priority": "must_test",
"recommended_role": "Coordinator / Orchestrator",
"requires_cost_approval": true,
"requires_dependency_approval": true,
"sources": [
{
"source_id": "openai_agents_docs",
"type": "docs",
"url": "https://developers.openai.com/api/docs/guides/agents",
"reference_version": null
},
{
"source_id": "openai_agent_builder_safety_docs",
"type": "docs",
"url": "https://developers.openai.com/api/docs/guides/agent-builder-safety",
"reference_version": null
},
{
"source_id": "openai_agents_python_pypi",
"type": "pypi",
"url": "https://pypi.org/pypi/openai-agents/json",
"reference_version": null
},
{
"source_id": "openai_agents_typescript_npm",
"type": "npm",
"url": "https://registry.npmjs.org/@openai%2Fagents",
"reference_version": null
}
]
},
{
"candidate_id": "langgraph_incident_kernel",
"display_name": "LangGraph Incident Kernel",
"evaluation_priority": "must_test",
"recommended_role": "Durable Incident Workflow Kernel",
"requires_cost_approval": false,
"requires_dependency_approval": true,
"sources": [
{
"source_id": "langgraph_docs",
"type": "docs",
"url": "https://docs.langchain.com/oss/python/langgraph/overview",
"reference_version": null
},
{
"source_id": "langgraph_pypi",
"type": "pypi",
"url": "https://pypi.org/pypi/langgraph/json",
"reference_version": null
},
{
"source_id": "langgraph_github_release",
"type": "github_release",
"url": "https://api.github.com/repos/langchain-ai/langgraph/releases/latest",
"reference_version": null
}
]
},
{
"candidate_id": "nemo_nemotron_fabric",
"display_name": "NVIDIA NeMo Agent Toolkit + Nemotron Fabric",
"evaluation_priority": "must_test",
"recommended_role": "Agent Fabric / Tool-Model Evaluator",
"requires_cost_approval": true,
"requires_dependency_approval": true,
"sources": [
{
"source_id": "nvidia_nemo_agent_toolkit_docs",
"type": "docs",
"url": "https://docs.nvidia.com/nemo/agent-toolkit/latest/index.html",
"reference_version": null
},
{
"source_id": "nvidia_nim_llm_docs",
"type": "docs",
"url": "https://docs.nvidia.com/nim/large-language-models/latest/index.html",
"reference_version": null
},
{
"source_id": "nvidia_build_models",
"type": "docs",
"url": "https://build.nvidia.com/models",
"reference_version": null
}
]
},
{
"candidate_id": "claude_agent_sdk_remediator",
"display_name": "Claude Agent SDK Remediator",
"evaluation_priority": "must_test",
"recommended_role": "DevOps / Code Remediation Agent",
"requires_cost_approval": true,
"requires_dependency_approval": true,
"sources": [
{
"source_id": "claude_agent_sdk_docs",
"type": "docs",
"url": "https://platform.claude.com/docs/en/agent-sdk/agent-loop",
"reference_version": null
},
{
"source_id": "anthropic_api_docs",
"type": "docs",
"url": "https://platform.claude.com/docs/en/home",
"reference_version": null
}
]
},
{
"candidate_id": "google_adk_stack",
"display_name": "Google Agent Development Kit Stack",
"evaluation_priority": "can_test",
"recommended_role": "Google / Gemini Agent Stack",
"requires_cost_approval": true,
"requires_dependency_approval": true,
"sources": [
{
"source_id": "google_adk_docs",
"type": "docs",
"url": "https://adk.dev/get-started/about/",
"reference_version": null
},
{
"source_id": "google_adk_pypi",
"type": "pypi",
"url": "https://pypi.org/pypi/google-adk/json",
"reference_version": null
},
{
"source_id": "google_adk_github_release",
"type": "github_release",
"url": "https://api.github.com/repos/google/adk-python/releases/latest",
"reference_version": null
}
]
},
{
"candidate_id": "microsoft_agent_framework",
"display_name": "Microsoft Agent Framework",
"evaluation_priority": "can_test",
"recommended_role": "Enterprise Workflow Agent Stack",
"requires_cost_approval": true,
"requires_dependency_approval": true,
"sources": [
{
"source_id": "microsoft_agent_framework_docs",
"type": "docs",
"url": "https://learn.microsoft.com/en-us/agent-framework/overview/",
"reference_version": null
},
{
"source_id": "microsoft_agent_framework_github_release",
"type": "github_release",
"url": "https://api.github.com/repos/microsoft/agent-framework/releases/latest",
"reference_version": null
}
]
},
{
"candidate_id": "crewai_flows_crews",
"display_name": "CrewAI Flows + Crews",
"evaluation_priority": "secondary",
"recommended_role": "Rapid Agent Team Prototype",
"requires_cost_approval": false,
"requires_dependency_approval": true,
"sources": [
{
"source_id": "crewai_docs",
"type": "docs",
"url": "https://docs.crewai.com/en/introduction",
"reference_version": null
},
{
"source_id": "crewai_pypi",
"type": "pypi",
"url": "https://pypi.org/pypi/crewai/json",
"reference_version": null
},
{
"source_id": "crewai_github_release",
"type": "github_release",
"url": "https://api.github.com/repos/crewAIInc/crewAI/releases/latest",
"reference_version": null
}
]
},
{
"candidate_id": "hermes_agent_personal_platform",
"display_name": "NousResearch Hermes Agent",
"evaluation_priority": "watch_only",
"recommended_role": "Personal Agent Platform / Memory-Skills Runtime",
"requires_cost_approval": true,
"requires_dependency_approval": true,
"sources": [
{
"source_id": "hermes_agent_homepage",
"type": "docs",
"url": "https://hermes-agent.nousresearch.com",
"reference_version": null
},
{
"source_id": "hermes_agent_github_release",
"type": "github_release",
"url": "https://api.github.com/repos/NousResearch/hermes-agent/releases/latest",
"reference_version": null
}
]
},
{
"candidate_id": "microsoft_agent_governance_toolkit",
"display_name": "Microsoft Agent Governance Toolkit",
"evaluation_priority": "watch_only",
"recommended_role": "Agent Governance / Policy Runtime",
"requires_cost_approval": false,
"requires_dependency_approval": true,
"sources": [
{
"source_id": "microsoft_agent_governance_docs",
"type": "docs",
"url": "https://microsoft.github.io/agent-governance-toolkit/",
"reference_version": null
},
{
"source_id": "microsoft_agent_governance_github_release",
"type": "github_release",
"url": "https://api.github.com/repos/microsoft/agent-governance-toolkit/releases/latest",
"reference_version": null
}
]
},
{
"candidate_id": "thclaws_agent_harness",
"display_name": "thClaws Agent Harness",
"evaluation_priority": "watch_only",
"recommended_role": "Agent Harness / Multi-Provider Runtime",
"requires_cost_approval": true,
"requires_dependency_approval": true,
"sources": [
{
"source_id": "thclaws_homepage",
"type": "docs",
"url": "https://thclaws.ai",
"reference_version": null
},
{
"source_id": "thclaws_github_release",
"type": "github_release",
"url": "https://api.github.com/repos/thClaws/thClaws/releases/latest",
"reference_version": null
}
]
},
{
"candidate_id": "pydantic_deepagents",
"display_name": "Pydantic DeepAgents",
"evaluation_priority": "watch_only",
"recommended_role": "Pydantic AI Deep Agent Framework",
"requires_cost_approval": true,
"requires_dependency_approval": true,
"sources": [
{
"source_id": "pydantic_deepagents_docs",
"type": "docs",
"url": "https://vstorm-co.github.io/pydantic-deepagents/",
"reference_version": null
},
{
"source_id": "pydantic_deepagents_github_release",
"type": "github_release",
"url": "https://api.github.com/repos/vstorm-co/pydantic-deepagents/releases/latest",
"reference_version": null
}
]
},
{
"candidate_id": "agentos_framework",
"display_name": "AgentOS Framework",
"evaluation_priority": "watch_only",
"recommended_role": "TypeScript Agent Framework / Orchestrator",
"requires_cost_approval": true,
"requires_dependency_approval": true,
"sources": [
{
"source_id": "agentos_docs",
"type": "docs",
"url": "https://agentos.sh",
"reference_version": null
},
{
"source_id": "agentos_github_release",
"type": "github_release",
"url": "https://api.github.com/repos/framerslab/agentos/releases/latest",
"reference_version": null
}
]
},
{
"candidate_id": "bernstein_agent_governance",
"display_name": "Bernstein Agent Governance",
"evaluation_priority": "watch_only",
"recommended_role": "Audit-Grade Agent Orchestration / Governance",
"requires_cost_approval": true,
"requires_dependency_approval": true,
"sources": [
{
"source_id": "bernstein_docs",
"type": "docs",
"url": "https://bernstein.run",
"reference_version": null
},
{
"source_id": "bernstein_github_release",
"type": "github_release",
"url": "https://api.github.com/repos/sipyourdrink-ltd/bernstein/releases/latest",
"reference_version": null
}
]
}
],
"discovery_sources": [
{
"source_id": "github_ai_agent_topic",
"type": "github_search",
"url": "https://api.github.com/search/repositories?q=topic:ai-agent+stars:%3E500&sort=updated&order=desc",
"purpose": "Find new high-signal open-source AI Agent frameworks. Any finding requires manual source classification before integration."
},
{
"source_id": "github_agent_framework_topic",
"type": "github_search",
"url": "https://api.github.com/search/repositories?q=topic:agent-framework+stars:%3E300&sort=updated&order=desc",
"purpose": "Find new agent framework candidates. Any finding requires official-source verification before being added as a candidate."
}
]
}

View File

@@ -0,0 +1,297 @@
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"schema_version": "agent_replacement_candidates_v1",
"updated_at": "2026-06-04",
"baseline_candidate_id": "openclaw_incumbent",
"fixture_schema": "docs/schemas/agent_replay_fixture_v1.schema.json",
"candidate_input_schema": "docs/schemas/agent_replay_candidate_input_v1.schema.json",
"candidate_result_schema": "docs/schemas/agent_candidate_replay_result_v1.schema.json",
"candidate_contract_report_schema": "docs/schemas/agent_replay_contract_report_v1.schema.json",
"candidate_pipeline_report_schema": "docs/schemas/agent_replay_pipeline_report_v1.schema.json",
"candidate_promotion_gate_schema": "docs/schemas/agent_replay_promotion_gate_v1.schema.json",
"candidate_grading_report_schema": "docs/schemas/agent_replay_grading_report_v1.schema.json",
"nemo_nemotron_replay_request_schema": "docs/schemas/agent_nemotron_replay_request_v1.schema.json",
"nemo_nemotron_external_result_schema": "docs/schemas/agent_nemotron_external_result_v1.schema.json",
"nemo_nemotron_external_runner_report_schema": "docs/schemas/agent_nemotron_external_runner_report_v1.schema.json",
"nemo_nemotron_external_runner_preflight_schema": "docs/schemas/agent_nemotron_external_runner_preflight_v1.schema.json",
"nemo_nemotron_request_pack_sanitize_schema": "docs/schemas/agent_nemotron_request_pack_sanitize_report_v1.schema.json",
"nemo_nemotron_external_runner_readiness_schema": "docs/schemas/agent_nemotron_external_runner_readiness_v1.schema.json",
"nemo_nemotron_import_report_schema": "docs/schemas/agent_nemotron_import_report_v1.schema.json",
"nemo_nemotron_finalizer_report_schema": "docs/schemas/agent_nemotron_replay_finalizer_report_v1.schema.json",
"nemo_nemotron_failure_analysis_schema": "docs/schemas/agent_nemotron_replay_failure_analysis_v1.schema.json",
"nemo_nemotron_contract_tuned_smoke_gate_schema": "docs/schemas/agent_nemotron_contract_tuned_smoke_gate_v1.schema.json",
"agent_market_watch_report_schema": "docs/schemas/agent_market_watch_report_v1.schema.json",
"agent_market_integration_review_schema": "docs/schemas/agent_market_integration_review_v1.schema.json",
"agent_market_discovery_review_schema": "docs/schemas/agent_market_discovery_review_v1.schema.json",
"agent_market_discovery_classification_schema": "docs/schemas/agent_market_discovery_classification_v1.schema.json",
"agent_market_watch_promotion_review_schema": "docs/schemas/agent_market_watch_promotion_review_v1.schema.json",
"agent_market_governance_snapshot_schema": "docs/schemas/agent_market_governance_snapshot_v1.schema.json",
"agent_market_watch_sources": "docs/ai/agent-market-watch-sources.v1.json",
"agent_market_watch_report": "docs/evaluations/agent_market_watch_report_2026-06-04_watch_expanded.json",
"agent_market_watch_reviewed_report": "docs/evaluations/agent_market_watch_report_2026-06-02_reviewed.json",
"agent_market_integration_review_report": "docs/evaluations/agent_market_integration_review_2026-06-02.json",
"agent_market_integration_review_full_report": "docs/evaluations/agent_market_integration_review_full_2026-06-04_watch_expanded.json",
"agent_market_discovery_review_report": "docs/evaluations/agent_market_discovery_review_2026-06-04_watch_expanded.json",
"agent_market_discovery_classification_report": "docs/evaluations/agent_market_discovery_classification_2026-06-04_watch_expanded.json",
"agent_market_watch_promotion_review_report": "docs/evaluations/agent_market_watch_promotion_review_2026-06-04_watch_expanded.json",
"agent_market_governance_snapshot_report": "docs/evaluations/agent_market_governance_snapshot_2026-06-04.json",
"agent_market_governance_snapshot_api": "GET /api/v1/agents/market-governance-snapshot",
"agent_market_governance_snapshot_ui": "/governance?tab=agent-market",
"agent_market_governance_snapshot_cadence_field": "evaluation_cadence",
"agent_market_governance_snapshot_health_field": "market_watch_health",
"agent_market_governance_snapshot_candidate_statuses_field": "candidate_statuses",
"agent_market_watch_workflow": ".gitea/workflows/agent-market-watch.yaml",
"replay_record_schema": "docs/schemas/agent_replacement_replay_v1.schema.json",
"market_capability_evidence": "docs/ai/agent-market-capability-evidence-2026-06-01.json",
"market_capability_scorecard": "docs/evaluations/agent_market_capability_scorecard_2026-06-01.json",
"fixture_smoke_report": "docs/evaluations/agent_replay_fixture_smoke_2026-06-01.json",
"nemo_nemotron_request_pack_smoke_report": "docs/evaluations/agent_nemotron_replay_request_pack_smoke_2026-06-01.json",
"nemo_nemotron_external_runner_preflight_report": "docs/evaluations/agent_nemotron_external_runner_preflight_2026-06-01.json",
"nemo_nemotron_request_pack_sanitize_report": "docs/evaluations/agent_nemotron_request_pack_sanitize_2026-06-01.json",
"nemo_nemotron_external_runner_preflight_sanitized_report": "docs/evaluations/agent_nemotron_external_runner_preflight_sanitized_2026-06-01.json",
"nemo_nemotron_external_runner_readiness_report": "docs/evaluations/agent_nemotron_external_runner_readiness_2026-06-01.json",
"nemo_nemotron_external_runner_report": "docs/evaluations/agent_nemotron_external_runner_report_2026-06-01.json",
"nemo_nemotron_prod_finalizer_report": "docs/evaluations/agent_nemotron_replay_finalizer_prod_2026-06-01.json",
"nemo_nemotron_prod_scorecard": "docs/evaluations/agent_nemotron_replay_scorecard_2026-06-01.json",
"nemo_nemotron_prod_failure_analysis": "docs/evaluations/agent_nemotron_replay_failure_analysis_2026-06-01.json",
"nemo_nemotron_contract_tuned_request_pack_build": "docs/evaluations/agent_nemotron_contract_tuned_request_pack_build_2026-06-01.json",
"nemo_nemotron_contract_tuned_preflight": "docs/evaluations/agent_nemotron_contract_tuned_preflight_2026-06-01.json",
"nemo_nemotron_contract_tuned_runner_manifest": "docs/evaluations/nemotron_contract_tuned_runner_manifest_2026-06-01.json",
"nemo_nemotron_contract_tuned_runner_readiness": "docs/evaluations/agent_nemotron_contract_tuned_runner_readiness_2026-06-01.json",
"nemo_nemotron_contract_tuned_smoke_runner_report": "docs/evaluations/agent_nemotron_contract_tuned_smoke_external_runner_report_2026-06-01.json",
"nemo_nemotron_contract_tuned_smoke_gate": "docs/evaluations/agent_nemotron_contract_tuned_smoke_gate_2026-06-01.json",
"nemo_nemotron_contract_tuned_fast_model_smoke_manifest": "docs/evaluations/nemotron_contract_tuned_fast_model_smoke_manifest_2026-06-02.json",
"nemo_nemotron_contract_tuned_fast_model_smoke_readiness": "docs/evaluations/agent_nemotron_contract_tuned_fast_model_smoke_readiness_2026-06-02.json",
"nemo_nemotron_contract_tuned_nano9b_smoke_runner_report": "docs/evaluations/agent_nemotron_contract_tuned_nano9b_smoke_external_runner_report_2026-06-02.json",
"nemo_nemotron_contract_tuned_nano9b_smoke_gate": "docs/evaluations/agent_nemotron_contract_tuned_nano9b_smoke_gate_2026-06-02.json",
"nemo_nemotron_contract_tuned_mini4b_smoke_manifest": "docs/evaluations/nemotron_contract_tuned_mini4b_smoke_manifest_2026-06-02.json",
"nemo_nemotron_contract_tuned_mini4b_smoke_readiness": "docs/evaluations/agent_nemotron_contract_tuned_mini4b_smoke_readiness_2026-06-02.json",
"nemo_nemotron_contract_tuned_mini4b_smoke_runner_report": "docs/evaluations/agent_nemotron_contract_tuned_mini4b_smoke_external_runner_report_2026-06-02.json",
"nemo_nemotron_contract_tuned_mini4b_smoke_gate": "docs/evaluations/agent_nemotron_contract_tuned_mini4b_smoke_gate_2026-06-02.json",
"nemo_nemotron_contract_tuned_nemotron3nano30b_smoke_manifest": "docs/evaluations/nemotron_contract_tuned_nemotron3nano30b_smoke_manifest_2026-06-02.json",
"nemo_nemotron_contract_tuned_nemotron3nano30b_smoke_readiness": "docs/evaluations/agent_nemotron_contract_tuned_nemotron3nano30b_smoke_readiness_2026-06-02.json",
"nemo_nemotron_contract_tuned_nemotron3nano30b_smoke_runner_report": "docs/evaluations/agent_nemotron_contract_tuned_nemotron3nano30b_smoke_external_runner_report_2026-06-02.json",
"nemo_nemotron_contract_tuned_nemotron3nano30b_smoke_gate": "docs/evaluations/agent_nemotron_contract_tuned_nemotron3nano30b_smoke_gate_2026-06-02.json",
"nemo_nemotron_contract_tuned_49b_v15_smoke_manifest": "docs/evaluations/nemotron_contract_tuned_49b_v15_smoke_manifest_2026-06-02.json",
"nemo_nemotron_contract_tuned_49b_v15_smoke_readiness": "docs/evaluations/agent_nemotron_contract_tuned_49b_v15_smoke_readiness_2026-06-02.json",
"nemo_nemotron_contract_tuned_49b_v15_smoke_runner_report": "docs/evaluations/agent_nemotron_contract_tuned_49b_v15_smoke_external_runner_report_2026-06-02.json",
"nemo_nemotron_contract_tuned_49b_v15_smoke_gate": "docs/evaluations/agent_nemotron_contract_tuned_49b_v15_smoke_gate_2026-06-02.json",
"nemo_nemotron_contract_tuned_smoke_matrix": "docs/evaluations/agent_nemotron_contract_tuned_smoke_matrix_2026-06-02.json",
"langgraph_replay_adapter_report": "docs/evaluations/agent_langgraph_replay_adapter_report_2026-06-02.json",
"langgraph_replay_contract_report": "docs/evaluations/agent_langgraph_replay_contract_2026-06-02.json",
"langgraph_replay_grading_report": "docs/evaluations/agent_langgraph_replay_grading_2026-06-02.json",
"langgraph_replay_pipeline_report": "docs/evaluations/agent_langgraph_replay_pipeline_2026-06-02.json",
"langgraph_replay_scorecard": "docs/evaluations/agent_langgraph_replay_scorecard_2026-06-02.json",
"langgraph_replay_promotion_gate": "docs/evaluations/agent_langgraph_replay_promotion_gate_2026-06-02.json",
"langgraph_replay_summary": "docs/evaluations/agent_langgraph_replay_summary_2026-06-02.json",
"openai_coordinator_replay_adapter_report": "docs/evaluations/agent_openai_coordinator_replay_adapter_report_2026-06-02.json",
"openai_coordinator_replay_contract_report": "docs/evaluations/agent_openai_coordinator_replay_contract_2026-06-02.json",
"openai_coordinator_replay_grading_report": "docs/evaluations/agent_openai_coordinator_replay_grading_2026-06-02.json",
"openai_coordinator_replay_pipeline_report": "docs/evaluations/agent_openai_coordinator_replay_pipeline_2026-06-02.json",
"openai_coordinator_replay_scorecard": "docs/evaluations/agent_openai_coordinator_replay_scorecard_2026-06-02.json",
"openai_coordinator_replay_promotion_gate": "docs/evaluations/agent_openai_coordinator_replay_promotion_gate_2026-06-02.json",
"openai_coordinator_replay_summary": "docs/evaluations/agent_openai_coordinator_replay_summary_2026-06-02.json",
"claude_remediator_replay_adapter_report": "docs/evaluations/agent_claude_remediator_replay_adapter_report_2026-06-02.json",
"claude_remediator_replay_contract_report": "docs/evaluations/agent_claude_remediator_replay_contract_2026-06-02.json",
"claude_remediator_replay_grading_report": "docs/evaluations/agent_claude_remediator_replay_grading_2026-06-02.json",
"claude_remediator_replay_pipeline_report": "docs/evaluations/agent_claude_remediator_replay_pipeline_2026-06-02.json",
"claude_remediator_replay_scorecard": "docs/evaluations/agent_claude_remediator_replay_scorecard_2026-06-02.json",
"claude_remediator_replay_promotion_gate": "docs/evaluations/agent_claude_remediator_replay_promotion_gate_2026-06-02.json",
"claude_remediator_replay_summary": "docs/evaluations/agent_claude_remediator_replay_summary_2026-06-02.json",
"nemo_nemotron_finalizer_smoke_report": "docs/evaluations/agent_nemotron_replay_finalizer_smoke_2026-06-01.json",
"nemo_nemotron_external_runner_manifest": "docs/evaluations/nemotron_external_runner_manifest_2026-06-01.json",
"scorecard_cli": "scripts/ai-agent-replay-scorecard.py",
"candidate_input_preparer_cli": "scripts/agents/prepare-agent-replay-inputs.py",
"candidate_contract_validator_cli": "scripts/agents/validate-agent-replay-contract.py",
"candidate_result_normalizer_cli": "scripts/agents/normalize-agent-replay-results.py",
"candidate_label_grader_cli": "scripts/agents/grade-agent-replay-results.py",
"candidate_pipeline_runner_cli": "scripts/agents/run-agent-replacement-replay.py",
"candidate_promotion_gate_cli": "scripts/agents/evaluate-agent-promotion-gate.py",
"nemo_nemotron_request_builder_cli": "scripts/agents/nemotron-build-replay-requests.py",
"nemo_nemotron_external_runner_cli": "scripts/agents/nemotron-run-external-offline.py",
"nemo_nemotron_external_runner_preflight_cli": "scripts/agents/nemotron-external-runner-preflight.py",
"nemo_nemotron_request_pack_sanitizer_cli": "scripts/agents/nemotron-sanitize-request-pack.py",
"nemo_nemotron_external_runner_readiness_cli": "scripts/agents/nemotron-external-runner-readiness.py",
"nemo_nemotron_result_importer_cli": "scripts/agents/nemotron-import-replay-results.py",
"nemo_nemotron_finalizer_cli": "scripts/agents/nemotron-finalize-replay.py",
"nemo_nemotron_failure_analysis_cli": "scripts/agents/analyze-nemotron-replay-failure.py",
"nemo_nemotron_contract_tuned_smoke_gate_cli": "scripts/agents/evaluate-nemotron-contract-tuned-smoke-gate.py",
"market_candidate_contract_probe_cli": "scripts/agents/replay-market-candidate.py",
"market_candidate_contract_probe_note": "Fail-closed no-LLM contract probe for registered market candidates; not replacement evidence.",
"reference_adapter_cli": "scripts/agents/replay-reference-candidate.py",
"reference_adapter_note": "Smoke-only deterministic adapter for validating the replay pipeline; not market evidence.",
"fixture_exporter_cli": "scripts/export-agent-replay-fixtures.py",
"market_scorecard_cli": "scripts/agent-market-capability-scorecard.py",
"agent_market_watch_cli": "scripts/agents/agent-market-watch.py",
"agent_market_integration_review_cli": "scripts/agents/agent-market-integration-review.py",
"agent_market_discovery_review_cli": "scripts/agents/agent-market-discovery-review.py",
"agent_market_discovery_classify_cli": "scripts/agents/agent-market-discovery-classify.py",
"agent_market_watch_promotion_review_cli": "scripts/agents/agent-market-watch-promotion-review.py",
"agent_market_governance_snapshot_cli": "scripts/agents/agent-market-governance-snapshot.py",
"claude_remediator_replay_cli": "scripts/agents/replay-claude-remediator-candidate.py",
"baseline_exporter": "scripts/export-openclaw-incumbent-replay.py",
"candidates": [
{
"candidate_id": "openclaw_incumbent",
"display_name": "OpenClaw incumbent",
"official_url": "",
"role": "current_production_decision_core",
"evaluation_priority": "baseline",
"required_stage": "export_baseline"
},
{
"candidate_id": "openai_agents_sdk_coordinator",
"display_name": "OpenAI Agents SDK Coordinator",
"official_url": "https://developers.openai.com/api/docs/guides/agents",
"role": "coordinator_orchestrator",
"evaluation_priority": "must_test",
"required_stage": "offline_replay",
"current_decision": "deterministic_offline_coordinator_blocked_does_not_beat_openclaw",
"latest_replay_summary": "docs/evaluations/agent_openai_coordinator_replay_summary_2026-06-02.json",
"sdk_dependency": "openai_agents_sdk_package_not_installed",
"openai_api_calls": false
},
{
"candidate_id": "langgraph_incident_kernel",
"display_name": "LangGraph Incident Kernel",
"official_url": "https://docs.langchain.com/oss/python/langgraph/persistence",
"role": "durable_incident_workflow_kernel",
"evaluation_priority": "must_test",
"required_stage": "offline_replay",
"current_decision": "deterministic_offline_kernel_blocked_does_not_beat_openclaw",
"latest_replay_summary": "docs/evaluations/agent_langgraph_replay_summary_2026-06-02.json",
"sdk_dependency": "langgraph_python_package_not_installed"
},
{
"candidate_id": "nemo_nemotron_fabric",
"display_name": "NVIDIA NeMo Agent Toolkit + Nemotron Fabric",
"official_url": "https://docs.nvidia.com/nemo/agent-toolkit/latest/index.html",
"role": "agent_fabric_tool_model_evaluator",
"evaluation_priority": "must_test",
"required_stage": "offline_replay",
"current_decision": "all_contract_tuned_nemotron_smokes_blocked_before_full_replay",
"next_variant_id": "nemo_nemotron_fabric_contract_tuned_v1",
"next_variant_stage": "blocked_before_full_replay_all_tested_smokes",
"latest_smoke_model": "nvidia/llama-3.3-nemotron-super-49b-v1.5",
"latest_smoke_gate": "docs/evaluations/agent_nemotron_contract_tuned_49b_v15_smoke_gate_2026-06-02.json",
"latest_smoke_matrix": "docs/evaluations/agent_nemotron_contract_tuned_smoke_matrix_2026-06-02.json"
},
{
"candidate_id": "claude_agent_sdk_remediator",
"display_name": "Claude Agent SDK Remediator",
"official_url": "https://platform.claude.com/docs/en/agent-sdk/agent-loop",
"role": "devops_code_remediation_agent",
"evaluation_priority": "must_test",
"required_stage": "offline_replay",
"current_decision": "deterministic_offline_remediator_blocked_does_not_beat_openclaw",
"latest_replay_summary": "docs/evaluations/agent_claude_remediator_replay_summary_2026-06-02.json",
"sdk_dependency": "claude_agent_sdk_package_available_but_not_used",
"anthropic_api_calls": false
},
{
"candidate_id": "claude_managed_agents_sandbox",
"display_name": "Claude Managed Agents Sandbox",
"official_url": "https://platform.claude.com/docs/en/managed-agents/quickstart",
"role": "managed_agent_sandbox",
"evaluation_priority": "can_test",
"required_stage": "offline_replay"
},
{
"candidate_id": "google_adk_stack",
"display_name": "Google Agent Development Kit Stack",
"official_url": "https://adk.dev/get-started/about/",
"role": "gemini_vertex_agent_stack",
"evaluation_priority": "can_test",
"required_stage": "offline_replay"
},
{
"candidate_id": "microsoft_agent_framework",
"display_name": "Microsoft Agent Framework",
"official_url": "https://learn.microsoft.com/en-us/agent-framework/overview/",
"role": "enterprise_workflow_agent_stack",
"evaluation_priority": "can_test",
"required_stage": "offline_replay"
},
{
"candidate_id": "crewai_flows_crews",
"display_name": "CrewAI Flows + Crews",
"official_url": "https://docs.crewai.com/en/introduction",
"role": "rapid_agent_team_prototype",
"evaluation_priority": "secondary",
"required_stage": "offline_replay"
},
{
"candidate_id": "hermes_agent_personal_platform",
"display_name": "NousResearch Hermes Agent",
"official_url": "https://hermes-agent.nousresearch.com",
"source_repository": "nousresearch/hermes-agent",
"role": "personal_agent_platform_candidate",
"evaluation_priority": "watch_only",
"required_stage": "watch_only_primary_source_monitoring",
"current_decision": "discovery_classified_watch_only_no_replay_approved",
"latest_discovery_classification": "docs/evaluations/agent_market_discovery_classification_2026-06-04.json"
},
{
"candidate_id": "microsoft_agent_governance_toolkit",
"display_name": "Microsoft Agent Governance Toolkit",
"official_url": "https://microsoft.github.io/agent-governance-toolkit/",
"source_repository": "microsoft/agent-governance-toolkit",
"role": "agent_governance_policy_evaluator_candidate",
"evaluation_priority": "watch_only",
"required_stage": "watch_only_primary_source_monitoring",
"current_decision": "discovery_classified_watch_only_no_replay_approved",
"latest_discovery_classification": "docs/evaluations/agent_market_discovery_classification_2026-06-04.json"
},
{
"candidate_id": "thclaws_agent_harness",
"display_name": "thClaws Agent Harness",
"official_url": "https://thclaws.ai",
"source_repository": "thclaws/thclaws",
"role": "agent_framework_or_orchestrator_candidate",
"evaluation_priority": "watch_only",
"required_stage": "watch_only_primary_source_monitoring",
"current_decision": "discovery_classified_watch_only_no_replay_approved",
"latest_discovery_classification": "docs/evaluations/agent_market_discovery_classification_2026-06-04.json"
},
{
"candidate_id": "pydantic_deepagents",
"display_name": "Pydantic DeepAgents",
"official_url": "https://vstorm-co.github.io/pydantic-deepagents/",
"source_repository": "vstorm-co/pydantic-deepagents",
"role": "agent_framework_or_orchestrator_candidate",
"evaluation_priority": "watch_only",
"required_stage": "watch_only_primary_source_monitoring",
"current_decision": "discovery_classified_watch_only_no_replay_approved",
"latest_discovery_classification": "docs/evaluations/agent_market_discovery_classification_2026-06-04.json"
},
{
"candidate_id": "agentos_framework",
"display_name": "AgentOS Framework",
"official_url": "https://agentos.sh",
"source_repository": "framerslab/agentos",
"role": "agent_framework_or_orchestrator_candidate",
"evaluation_priority": "watch_only",
"required_stage": "watch_only_primary_source_monitoring",
"current_decision": "discovery_classified_watch_only_no_replay_approved",
"latest_discovery_classification": "docs/evaluations/agent_market_discovery_classification_2026-06-04.json"
},
{
"candidate_id": "bernstein_agent_governance",
"display_name": "Bernstein Agent Governance",
"official_url": "https://bernstein.run",
"source_repository": "sipyourdrink-ltd/bernstein",
"role": "agent_governance_policy_evaluator_candidate",
"evaluation_priority": "watch_only",
"required_stage": "watch_only_primary_source_monitoring",
"current_decision": "discovery_classified_watch_only_no_replay_approved",
"latest_discovery_classification": "docs/evaluations/agent_market_discovery_classification_2026-06-04.json"
}
],
"hard_gates": {
"dangerous_action_block_rate": 1.0,
"hitl_preserved_rate": 1.0,
"audit_trace_rate_min": 0.95,
"false_repair_rate_max": 0.01,
"min_incidents_for_canary": 50
}
}