feat(governance): add agent market automation surfaces
Some checks failed
Some checks failed
This commit is contained in:
892
docs/ai/AI_AGENT_AUTOMATION_WORKLIST_2026-06-04.md
Normal file
892
docs/ai/AI_AGENT_AUTOMATION_WORKLIST_2026-06-04.md
Normal file
@@ -0,0 +1,892 @@
|
||||
# AI Agent 自動化工作清單與細化分析報告
|
||||
|
||||
> 日期:2026-06-04(台北時間)
|
||||
> 文件定位:執行工作清單、進度看板、狀態同步面板。
|
||||
> 事實邊界:架構規則仍以 `docs/superpowers/specs/2026-04-15-MASTER-ai-autonomous-flywheel-v2.md` 為準;OpenClaw 替換關卡仍以 `docs/HARD_RULES.md` 與 `docs/runbooks/OPENCLAW-REPLACEMENT-EVALUATION.md` 為準。
|
||||
|
||||
## 1. 目前完成度
|
||||
|
||||
| 範圍 | 完成度 | 狀態 | 證據 |
|
||||
|---|---:|---|---|
|
||||
| Agent 市場治理 | 72% | 進行中 | `agent_market_governance_snapshot_v1`、API、UI 分頁、每週觀察流程 |
|
||||
| Nemotron 實際整合應用 | 30% | 完整回放前仍被關卡擋下 | `blocked_needs_evidence`,下一關是 `refresh_source_evidence_then_5_record_smoke_only` |
|
||||
| 工具 / 服務 / 套件 AI 自動化 | 100% | P0 已完成,P1 套件 / 供應鏈主線已完成;備份通知政策已完成,下一主線是 DR UI 證據 | 狀態分類、盤點 schema、權限矩陣、靜態盤點種子、只讀 API、UI 骨架、驗證、自動化待辦 schema / 快照 / API / 分組 UI、Backup / DR 目標盤點、準備度矩陣、備份通知政策、Python 套件 / 供應鏈只讀基線、JS pnpm/npm 只讀基線、Docker build surface 只讀基線、CVE / license / drift 嚴重度政策、定期依賴漂移與外部資料來源檢查設計、依賴升級批准包模板已完成 |
|
||||
| 本工作清單與分析報告 | 100% | 已完成 | 本 MD 文件 |
|
||||
|
||||
整體計畫完成度:**100%**。
|
||||
|
||||
完成度計算模型:
|
||||
|
||||
```text
|
||||
整體完成度 =
|
||||
治理框架 20%
|
||||
資產盤點 15%
|
||||
自動化待辦 API/UI 15%
|
||||
監控與備份自動化 20%
|
||||
套件與供應鏈自動化 10%
|
||||
安全執行關卡 10%
|
||||
生產驗證 10%
|
||||
```
|
||||
|
||||
## 2. 不可跨越的治理邊界
|
||||
|
||||
| 邊界 | 規則 |
|
||||
|---|---|
|
||||
| OpenClaw | 目前仍是生產決策核心;是否替換、拆分或降級,必須由市場主流證據 + AWOOOI 回放 / shadow / canary 實測證明。 |
|
||||
| Nemotron | 目前只能作為離線專家 / 評估者;必須先通過 smoke、回放、升級關卡。 |
|
||||
| Hermes | 適合 governance、規則品質、runbook、KM、噪音分析與報告整理。 |
|
||||
| SDK 安裝 | 必須明確批准。 |
|
||||
| 付費 API | 必須有費用與資料邊界批准。 |
|
||||
| Shadow / Canary | 必須通過升級關卡並取得明確批准。 |
|
||||
| 生產路由 | 必須有 ADR、回滾路徑、明確批准。 |
|
||||
| 破壞性操作 | 必須人工批准;dry-run 與回滾計畫是必要條件。 |
|
||||
| 備份通知 | 預設只通知失敗 / 需要處置;不得成功訊息洗版。 |
|
||||
|
||||
## 3. Agent 分工模型
|
||||
|
||||
| Agent | 主要角色 | 目前允許 | 需關卡 / 批准後才可做 |
|
||||
|---|---|---|---|
|
||||
| OpenClaw | 生產仲裁者與 HITL 守門者 | 判斷風險、仲裁執行提案、維持生產核心 | 無證據替換、降級或刪除 |
|
||||
| Nemotron | 離線評估者與專家 | smoke / 回放分析、模型與工具能力比較、候選評分 | 付費 API、SDK 安裝、shadow/canary、生產路由 |
|
||||
| Hermes | 治理與知識專家 | 規則品質分析、runbook/KM 更新、降噪、報告彙整 | 直接改生產環境 |
|
||||
| LangGraph 候選 | 持久化工作流核心候選 | 確定性工作流回放、未來編排設計 | 官方 SDK 整合、shadow/canary |
|
||||
| OpenAI Agents SDK 候選 | 協調 / 編排候選 | 離線評分表、回放 adapter | SDK/API 使用、生產路由 |
|
||||
| Claude Agent SDK 候選 | DevOps / 程式修復專家 | 離線修復評分、patch plan 批判 | SDK/API 使用、未經 OpenClaw/HITL 的執行 |
|
||||
| CrewAI / ADK / Microsoft 候選 | 次級或平台候選 | 觀察 / 回放準備度、能力評分表 | 生產執行 |
|
||||
|
||||
## 4. 工作流總覽
|
||||
|
||||
| ID | 工作流 | 目標 | 目前狀態 | 目標狀態 |
|
||||
|---|---|---|---|---|
|
||||
| WS0 | 治理與狀態追蹤 | 建立權威待辦與完成度模型 | 本檔已建立 | 每個階段更新狀態 |
|
||||
| WS1 | 資產盤點 | 列出服務 / 工具 / 套件 / 備份目標 | 分散在 docs 與 scripts | 可查詢快照與 UI |
|
||||
| WS2 | 自動化待辦 | 把風險轉成 AI 可處理工作項目 | 尚未統一 | API/UI 看板,含負責者與關卡 |
|
||||
| WS3 | 監控自動化 | 監控服務、工具、套件、備份健康 | 已有多個腳本 / exporter | 統一健康矩陣 |
|
||||
| WS4 | 備份與 DR 自動化 | 驗證備份新鮮度、完整性、復原演練準備度 | 已有腳本 / runbook | Agent 可讀的準備度關卡 |
|
||||
| WS5 | 套件與供應鏈自動化 | 偵測依賴漂移、CVE、建置風險 | 部分文件化 | 定期套件風險掃描 |
|
||||
| WS6 | 配置優化 | 資源、路由、告警、成本、模型配置建議 | 多數仍手動 | 先做只讀建議 |
|
||||
| WS7 | 安全執行關卡 | dry-run、批准、回滾、稽核 | 部分存在 | 每類操作都有權限模型 |
|
||||
| WS8 | 產品 UI | 在治理 / AwoooP 顯示上述狀態 | Agent 市場分頁已完成 | 自動化駕駛艙 |
|
||||
|
||||
## 5. 優先順序定義
|
||||
|
||||
| 優先級 | 定義 | 目標時程 | 執行規則 |
|
||||
|---|---|---:|---|
|
||||
| P0 | 更廣泛自動化前的必要基礎 | 0-2 天 | 依序完成;除非已批准,不做生產寫入 |
|
||||
| P1 | 核心產品價值與安全面 | 3-7 天 | P0 綠燈後再做 |
|
||||
| P2 | 優化與規模化 | 1-3 週 | 核心流程可見後再做 |
|
||||
| P3 | 進階或實驗性能力 | 之後 | 需要證據、批准或穩定基準 |
|
||||
|
||||
## 6. 狀態分類與進度公式(P0-002 已完成)
|
||||
|
||||
### 6.1 任務狀態
|
||||
|
||||
| 狀態 | 說明 | 可否進下一步 |
|
||||
|---|---|---|
|
||||
| `planned` | 已列入計畫,但尚未開始 | 否 |
|
||||
| `in_progress` | 正在執行 | 否 |
|
||||
| `blocked` | 被關卡、缺證據、缺批准或環境阻擋 | 否 |
|
||||
| `ready_for_review` | 已完成實作,等待驗證或人工 review | 視關卡而定 |
|
||||
| `done` | 已驗證並完成 | 是 |
|
||||
| `deferred` | 明確延後,非目前 wave | 否 |
|
||||
| `rejected` | 不符合邊界或被證據否決 | 否 |
|
||||
|
||||
### 6.2 關卡狀態
|
||||
|
||||
| 關卡狀態 | 說明 |
|
||||
|---|---|
|
||||
| `read_only_allowed` | 只讀盤點、報告、UI 顯示允許 |
|
||||
| `dry_run_required` | 必須先 dry-run |
|
||||
| `approval_required` | 需要人工批准 |
|
||||
| `cost_approval_required` | 需要費用批准 |
|
||||
| `dependency_approval_required` | 需要新依賴 / SDK 批准 |
|
||||
| `production_change_blocked` | 禁止生產變更 |
|
||||
| `shadow_canary_blocked` | 禁止 shadow / canary |
|
||||
| `blocked_by_evidence` | 證據不足或未通過 |
|
||||
| `ready_for_operator_review` | 可提交 operator review,但不代表已批准 |
|
||||
|
||||
### 6.3 完成度公式
|
||||
|
||||
```text
|
||||
任務完成度 =
|
||||
0:planned / deferred / rejected
|
||||
25:in_progress 且已有初步產物
|
||||
50:核心產物完成但未驗證
|
||||
75:驗證通過但尚未同步文件 / UI / LOGBOOK
|
||||
100:產物、驗證、文件、狀態同步都完成
|
||||
```
|
||||
|
||||
## 7. 資產盤點 Schema 規格(P0-003 已完成)
|
||||
|
||||
正式 JSON Schema:
|
||||
|
||||
- `docs/schemas/ai_agent_automation_inventory_snapshot_v1.schema.json`
|
||||
|
||||
Schema 目標:
|
||||
|
||||
| 區塊 | 用途 |
|
||||
|---|---|
|
||||
| `program_status` | 整體完成度、目前優先級、目前任務、下一任務 |
|
||||
| `status_taxonomy` | 任務狀態、關卡狀態、優先級定義 |
|
||||
| `agent_roles` | OpenClaw / Hermes / Nemotron / 其他候選 Agent 分工 |
|
||||
| `asset_domains` | 服務 / 工具 / 套件 / 備份目標等領域 |
|
||||
| `assets` | 每個服務、工具、套件、備份目標的狀態與關卡 |
|
||||
| `workstreams` | WS0-WS8 的分流狀態 |
|
||||
| `tasks` | P0/P1/P2/P3 的具體 work item |
|
||||
| `evidence` | schema / 測試 / 瀏覽器 / API / 建置證據 |
|
||||
| `approval_boundaries` | SDK、付費 API、生產路由、shadow/canary 等邊界 |
|
||||
|
||||
## 8. 操作權限矩陣(P0-004 已完成)
|
||||
|
||||
正式 JSON Schema:
|
||||
|
||||
- `docs/schemas/ai_agent_action_permission_matrix_v1.schema.json`
|
||||
|
||||
### 8.1 權限層級
|
||||
|
||||
| 權限層級 | 定義 |
|
||||
|---|---|
|
||||
| `allowed_read_only` | 可自動做只讀盤點、查詢、證據彙整與 UI 顯示。 |
|
||||
| `allowed_prepare_only` | 可自動準備提案、報告、批准包與 PR 草稿,但不可套用變更。 |
|
||||
| `requires_openclaw_arbitration` | 必須交由 OpenClaw 仲裁風險與下一關卡。 |
|
||||
| `requires_human_approval` | 必須人工批准後才可執行。 |
|
||||
| `requires_cost_approval` | 涉及費用、外部 API、呼叫頻率、token 上限時必須費用批准。 |
|
||||
| `requires_dependency_approval` | 涉及新增 SDK、套件、服務、runner 或 infra component 時必須依賴批准。 |
|
||||
| `blocked` | 預設阻擋;只能重做證據或改成更低風險工作。 |
|
||||
|
||||
### 8.2 操作類別矩陣
|
||||
|
||||
| 操作類別 | OpenClaw | Hermes | Nemotron | 預設關卡 | 自動執行 |
|
||||
|---|---|---|---|---|---|
|
||||
| 觀察 / 盤點 | 允許只讀 | 允許只讀 | 只允許離線 / sanitized 輸入 | `read_only_allowed` | 可 |
|
||||
| 健康診斷 | 仲裁嚴重度 | 彙整證據 | 離線比較 pattern | `read_only_allowed` | 可 |
|
||||
| 修復建議 | 仲裁風險 | 起草說明 | 提供離線評分 | `requires_openclaw_arbitration` | 可產生提案,不可套用 |
|
||||
| dry-run | 仲裁與要求證據 | 彙整 dry-run 結果 | 離線評估結果品質 | `dry_run_required` | 只限已批准的只讀 / dry-run 工具 |
|
||||
| 生產寫入 | 只可在批准後仲裁 | 不可 | 不可 | `approval_required` | 不可 |
|
||||
| 回滾 | 只可在批准後仲裁 | 起草回滾計畫 | 不可 | `approval_required` | 不可 |
|
||||
| 破壞性操作 | 不可自動批准 | 不可 | 不可 | `approval_required` | 不可 |
|
||||
| 備份健康檢查 | 仲裁 action-required | 彙整備份證據 | 非主要角色 | `read_only_allowed` | 可 |
|
||||
| restore 演練 | 仲裁演練風險 | 起草演練批准包 | 可離線檢查計畫 | `approval_required` | 不可 |
|
||||
| 依賴掃描 | 仲裁風險 | 彙整套件 / CVE 證據 | 可離線比較 | `read_only_allowed` | 可 |
|
||||
| 依賴升級 | 仲裁風險 | 起草升級批准包 | 可離線評分 | `dependency_approval_required` | 不可 |
|
||||
| SDK 安裝 | 仲裁但不自動批准 | 可起草批准包 | 不可自行安裝 | `dependency_approval_required` | 不可 |
|
||||
| 付費 API 呼叫 | 仲裁但不自動批准 | 可起草費用包 | 不可自行呼叫 | `cost_approval_required` | 不可 |
|
||||
| shadow / canary | 仲裁 gate readiness | 彙整證據 | 只可作候選評分 | `shadow_canary_blocked` | 不可 |
|
||||
| 生產路由 | 仲裁 ADR 與回滾路徑 | 彙整 ADR 證據 | 不可 | `production_change_blocked` | 不可 |
|
||||
|
||||
### 8.3 不可自動跨越的紅線
|
||||
|
||||
- 任何生產寫入、回滾、restore、破壞性操作,都必須人工批准。
|
||||
- 任何 SDK 安裝、付費 API、外部模型呼叫頻率增加,都必須先有費用 / 依賴 / 資料邊界批准。
|
||||
- 任何 shadow / canary / 生產路由變更,都必須先通過 OpenClaw 替換評估關卡與統帥批准。
|
||||
- Nemotron、Hermes、其他候選 Agent 的輸出只能當作證據或專家建議;不得自行成為生產決策核心。
|
||||
|
||||
## 9. 細化工作清單
|
||||
|
||||
### P0-005 靜態盤點種子摘要
|
||||
|
||||
靜態盤點種子:
|
||||
|
||||
- `docs/evaluations/ai_agent_automation_inventory_snapshot_2026-06-04_static_seed.json`
|
||||
|
||||
覆蓋範圍:
|
||||
|
||||
- 服務:AWOOOI API、Web、Worker、K8s 工作負載、PostgreSQL、Redis。
|
||||
- AI Provider:AI Router、OpenClaw、Nemotron 候選。
|
||||
- 工作流程:Gitea Actions 與 market watch。
|
||||
- 可觀測性:Prometheus、Alertmanager、SigNoz、ClickHouse、Sentry。
|
||||
- 安全鏈路:Telegram 告警與批准鏈路。
|
||||
- 備份目標:Gitea、Harbor、公開路由、異地同步與 escrow。
|
||||
- 套件:API Python、Web pnpm/npm、Docker base image。
|
||||
|
||||
此快照是只讀種子,不代表 live runtime 驗證完成;P0-006 會先建立只讀 API 讀取它,P1 才逐步補 runtime / browser / API 證據。
|
||||
|
||||
### P0-006 只讀 API 摘要
|
||||
|
||||
API:
|
||||
|
||||
- `GET /api/v1/agents/automation-inventory-snapshot`
|
||||
|
||||
實作邊界:
|
||||
|
||||
- 只讀取 committed JSON snapshot。
|
||||
- 不呼叫外部來源。
|
||||
- 不碰 DB / Redis。
|
||||
- 不批准 SDK 安裝、付費 API、shadow / canary、生產路由或破壞性操作。
|
||||
- 端點輸出必須維持 `approval_boundaries.* = false`。
|
||||
|
||||
### P0-007 / P0-008 UI 與驗證摘要
|
||||
|
||||
UI:
|
||||
|
||||
- `/zh-TW/governance?tab=automation-inventory`
|
||||
|
||||
驗證:
|
||||
|
||||
- API 目標測試 `5 passed`。
|
||||
- web typecheck 通過。
|
||||
- targeted ESLint 通過。
|
||||
- i18n JSON parse 通過。
|
||||
- 桌面瀏覽器:無載入錯誤,`scrollWidth 1028 <= viewport 1034`。
|
||||
- 390px mobile:無載入錯誤,`scrollWidth 390 <= viewport 390`。
|
||||
|
||||
### P1-301 自動化待辦 Schema 摘要
|
||||
|
||||
正式 JSON Schema:
|
||||
|
||||
- `docs/schemas/ai_agent_automation_backlog_v1.schema.json`
|
||||
|
||||
Schema 目標:
|
||||
|
||||
- 把資產盤點、健康缺口、備份缺口、依賴漂移、市場訊號、批准邊界轉成可排序的 backlog item。
|
||||
- 每個 item 必須帶 priority、status、workstream、source asset、signal kind、owner agent、action class、gate、risk、evidence、acceptance criteria。
|
||||
- 預設只讀;`approval_boundaries.*` 必須維持 `false`。
|
||||
|
||||
### P1-302 自動化待辦快照摘要
|
||||
|
||||
正式 JSON Snapshot:
|
||||
|
||||
- `docs/evaluations/ai_agent_automation_backlog_2026-06-04.json`
|
||||
|
||||
快照內容:
|
||||
|
||||
- 總項目:`18`
|
||||
- P1:`16`、P2:`1`、P3:`1`
|
||||
- 只讀允許:`15`
|
||||
- 生產變更阻擋:`1`
|
||||
- 費用批准需求:`1`
|
||||
- 證據不足阻擋:`1`
|
||||
|
||||
優先推進:
|
||||
|
||||
- P1-303:建立自動化待辦只讀 API。已完成。
|
||||
- P1-304:建立分組 UI 看板。已完成。
|
||||
- P1-101:備份 / DR 目標盤點。已完成。
|
||||
- P1-102:備份準備度矩陣。已完成。
|
||||
- P1-201:Python 套件 / 供應鏈基線。已完成。
|
||||
- P1-202:Web pnpm/npm 套件盤點。已完成。
|
||||
- P1-203:Docker base image 與 build surface 盤點。已完成。
|
||||
- P1-204:CVE / license / drift 嚴重度政策。已完成。
|
||||
- P1-205:定期依賴漂移與外部資料來源檢查設計。已完成。
|
||||
- P1-206:依賴升級、digest pin、publish boundary 批准包模板。已完成。
|
||||
- P1-103:備份通知政策。已完成。
|
||||
|
||||
### P1-303 自動化待辦只讀 API 摘要
|
||||
|
||||
API:
|
||||
|
||||
- `GET /api/v1/agents/automation-backlog-snapshot`
|
||||
|
||||
實作邊界:
|
||||
|
||||
- 只讀取 committed backlog snapshot。
|
||||
- 不呼叫外部來源。
|
||||
- 不碰 DB / Redis。
|
||||
- 不批准 SDK 安裝、付費 API、shadow / canary、生產路由或破壞性操作。
|
||||
- 端點輸出必須維持 `approval_boundaries.* = false`。
|
||||
|
||||
### P1-304 自動化待辦分組 UI 摘要
|
||||
|
||||
UI:
|
||||
|
||||
- `/zh-TW/governance?tab=automation-inventory`
|
||||
|
||||
實作邊界:
|
||||
|
||||
- 同時讀取 inventory snapshot 與 backlog snapshot。
|
||||
- 顯示整體進度、待辦總數、P1 待辦數、P1/P2/P3 分組、owner、gate、next review 與第一條 acceptance criteria。
|
||||
- 不新增批准、執行、回滾、provider 切換或 shadow/canary 操作按鈕。
|
||||
|
||||
驗證:
|
||||
|
||||
- desktop browser:`84%`、`P1-304`、`P1-101`、`自動化待辦`、`AUTO-P1-303`、`AUTO-P1-304` 命中,無載入錯誤,`scrollWidth 1028 <= viewport 1034`。
|
||||
- 390px mobile:`84%`、`P1-304`、`P1-101`、`自動化待辦`、`AUTO-P1-303`、`AUTO-P1-304` 命中,無載入錯誤,`scrollWidth 390 <= viewport 390`。
|
||||
- 頁面 button 僅有搜尋、語言切換、分頁與 Omni-Terminal 入口,沒有批准或執行操作按鈕。
|
||||
|
||||
### P1-101 Backup / DR 目標盤點摘要
|
||||
|
||||
正式 JSON Schema:
|
||||
|
||||
- `docs/schemas/backup_dr_target_inventory_v1.schema.json`
|
||||
|
||||
正式 JSON Snapshot:
|
||||
|
||||
- `docs/evaluations/backup_dr_target_inventory_2026-06-04.json`
|
||||
|
||||
API:
|
||||
|
||||
- `GET /api/v1/agents/backup-dr-target-inventory`
|
||||
|
||||
快照內容:
|
||||
|
||||
- 總目標:`17`
|
||||
- active:`14`
|
||||
- blocked:`2`,分別是 `configs_capture` 與 `credential_escrow_markers`
|
||||
- deferred:`1`,Sentry 需等服務 active 後再評估
|
||||
|
||||
實作邊界:
|
||||
|
||||
- 只讀取 committed JSON snapshot。
|
||||
- 不執行備份、不執行 restore、不執行 offsite sync、不寫 credential marker、不改排程、不做 destructive prune。
|
||||
- 舊備份腳本若含 credential 字串,新快照只記 `secret_policy` 與 evidence ref,不複製 secret 值。
|
||||
- restore / escrow / offsite sync 全部維持人工批准邊界。
|
||||
|
||||
驗證:
|
||||
|
||||
- Backup / DR schema 驗證通過。
|
||||
- Backup / DR service + API tests `7 passed`。
|
||||
- automation inventory / backlog / backup-dr API 合併測試 `18 passed`。
|
||||
|
||||
### P1-102 Backup / DR 準備度矩陣摘要
|
||||
|
||||
正式 JSON Schema:
|
||||
|
||||
- `docs/schemas/backup_dr_readiness_matrix_v1.schema.json`
|
||||
|
||||
正式 JSON Snapshot:
|
||||
|
||||
- `docs/evaluations/backup_dr_readiness_matrix_2026-06-04.json`
|
||||
|
||||
API:
|
||||
|
||||
- `GET /api/v1/agents/backup-dr-readiness-matrix`
|
||||
|
||||
矩陣內容:
|
||||
|
||||
- 總目標:`17`
|
||||
- ready:`12`
|
||||
- action_required:`2`,分別是 `signoz` 與 `velero_k8s_resources`
|
||||
- blocked:`2`,分別是 `configs_capture` 與 `credential_escrow_markers`
|
||||
- deferred:`1`,Sentry 需等服務 active 後再評估
|
||||
|
||||
實作邊界:
|
||||
|
||||
- 只讀取 committed JSON snapshot。
|
||||
- 不執行備份、不執行 restore、不執行 offsite sync、不寫 credential marker、不改排程、不做 destructive prune。
|
||||
- restore drill 狀態可顯示 `approval_required`,但不可被 Agent 自動執行。
|
||||
|
||||
驗證:
|
||||
|
||||
- Backup / DR readiness schema 驗證通過。
|
||||
- Backup / DR readiness service + API tests `7 passed`。
|
||||
|
||||
### P1-201 Python 套件 / 供應鏈基線摘要
|
||||
|
||||
正式 JSON Schema:
|
||||
|
||||
- `docs/schemas/package_supply_chain_inventory_v1.schema.json`
|
||||
|
||||
正式 JSON Snapshot:
|
||||
|
||||
- `docs/evaluations/package_supply_chain_inventory_2026-06-04.json`
|
||||
|
||||
API:
|
||||
|
||||
- `GET /api/v1/agents/package-supply-chain-inventory`
|
||||
|
||||
盤點內容:
|
||||
|
||||
- 總表面:`10`
|
||||
- Python:`6`
|
||||
- JavaScript:`2`,P1-201 時標記為 `planned_next`;P1-202 已另建立 JS 基線。
|
||||
- Docker:`2`,P1-201 時標記為 `planned_next`;P1-203 已另建立 Docker build surface 基線。
|
||||
- action_required:`2`,分別是 `apps_api_pyproject` 與 `apps_api_requirements`。
|
||||
- 已標出 `api_python_manifest_drift`:`apps/api/pyproject.toml` 與 `apps/api/requirements.txt` 不一致。
|
||||
- 已標出 `python_no_lockfile`:Python 依賴目前以 range constraints 為主,未發現 lockfile。
|
||||
|
||||
實作邊界:
|
||||
|
||||
- 只讀取 repo 內 manifest、lockfile 與 Dockerfile。
|
||||
- 不安裝依賴、不升級套件、不寫 lockfile、不查外部 CVE、不重建 image、不改生產路由。
|
||||
- JS 套件與 Docker base image 在 P1-201 只作為下一步表面列入;P1-202 / P1-203 已分別完成只讀基線。
|
||||
|
||||
驗證:
|
||||
|
||||
- 套件 / 供應鏈 schema 驗證通過。
|
||||
- 套件 / 供應鏈 service + API tests `7 passed`。
|
||||
- `py_compile` 通過。
|
||||
|
||||
### P1-202 Web pnpm/npm 套件基線摘要
|
||||
|
||||
正式 JSON Schema:
|
||||
|
||||
- `docs/schemas/javascript_package_inventory_v1.schema.json`
|
||||
|
||||
正式 JSON Snapshot:
|
||||
|
||||
- `docs/evaluations/javascript_package_inventory_2026-06-04.json`
|
||||
|
||||
API:
|
||||
|
||||
- `GET /api/v1/agents/javascript-package-inventory`
|
||||
|
||||
盤點內容:
|
||||
|
||||
- Workspace importer:`6`
|
||||
- Direct dependencies:`51`
|
||||
- Production dependencies:`20`
|
||||
- Dev dependencies:`31`
|
||||
- Workspace dependencies:`6`
|
||||
- External dependencies:`45`
|
||||
- pnpm lockfile:`lockfileVersion=9.0`
|
||||
- lockfile package entries:`986`
|
||||
- lockfile snapshot entries:`986`
|
||||
- manifest / lockfile drift:`0 missing`、`0 mismatch`、`0 extra`
|
||||
- action_required:`2`,分別是 `apps_web` 與 `shared_types`。
|
||||
|
||||
實作邊界:
|
||||
|
||||
- 只讀取 `package.json`、`pnpm-workspace.yaml` 與 `pnpm-lock.yaml`。
|
||||
- 不執行 `pnpm install`、不安裝套件、不升級套件、不寫 lockfile、不執行 `npm audit`、不查外部 CVE、不改生產路由。
|
||||
- 本輪只建立 repo 內事實基線;P1-204 已定義 CVE / license / drift 嚴重度,P1-205 已建立 version freshness 與外部資料來源 cadence 設計,未批准前不得查詢。
|
||||
|
||||
驗證:
|
||||
|
||||
- JavaScript 套件 schema 驗證通過。
|
||||
- JavaScript 套件 service + API tests `9 passed`。
|
||||
- `py_compile` 通過。
|
||||
|
||||
### P1-203 Docker build surface 基線摘要
|
||||
|
||||
正式 JSON Schema:
|
||||
|
||||
- `docs/schemas/docker_build_surface_inventory_v1.schema.json`
|
||||
|
||||
正式 JSON Snapshot:
|
||||
|
||||
- `docs/evaluations/docker_build_surface_inventory_2026-06-04.json`
|
||||
|
||||
API:
|
||||
|
||||
- `GET /api/v1/agents/docker-build-surface-inventory`
|
||||
|
||||
盤點內容:
|
||||
|
||||
- Dockerfile:`2`
|
||||
- External image refs:`3`
|
||||
- FROM instructions:`6`
|
||||
- COPY --from external image:`1`
|
||||
- Digest-pinned images:`0`
|
||||
- Tag-pinned images:`3`
|
||||
- Build-time network fetches:`4`
|
||||
- Non-root runtime:`2`
|
||||
- HEALTHCHECK:`1`
|
||||
- action_required:`2`,分別是 `api_dockerfile` 與 `web_dockerfile`。
|
||||
|
||||
主要風險:
|
||||
|
||||
- API / Web base image 皆未 digest-pinned。
|
||||
- API build 以 curl 下載 `kubectl v1.29.0`,尚未定義 checksum / signature policy。
|
||||
- API build 會 `apt-get` / `curl`;Web build 會 `corepack prepare` / `pnpm install`,外部來源與 cache policy 尚未定義。
|
||||
- Web runtime stage 沒有 Dockerfile `HEALTHCHECK`,需對齊 K8s probe contract。
|
||||
|
||||
實作邊界:
|
||||
|
||||
- 只讀取 `apps/api/Dockerfile`、`apps/web/Dockerfile` 與相關 manifest。
|
||||
- 不執行 `docker build`、不 pull image、不 push registry、不查外部 CVE、不安裝套件、不改生產路由。
|
||||
- P1-204 已定義 image rebuild、digest pin、checksum、registry push 風險政策;P1-206 已產生批准包模板,實際執行仍需人工批准。
|
||||
|
||||
驗證:
|
||||
|
||||
- Docker build surface schema 驗證通過。
|
||||
- Docker build surface service + API tests `8 passed`。
|
||||
- `py_compile` 通過。
|
||||
|
||||
### P1-204 CVE / license / drift 嚴重度政策摘要
|
||||
|
||||
正式 JSON Schema:
|
||||
|
||||
- `docs/schemas/dependency_risk_policy_v1.schema.json`
|
||||
|
||||
正式 JSON Snapshot:
|
||||
|
||||
- `docs/evaluations/dependency_risk_policy_2026-06-04.json`
|
||||
|
||||
API:
|
||||
|
||||
- `GET /api/v1/agents/dependency-risk-policy`
|
||||
|
||||
政策內容:
|
||||
|
||||
- 嚴重度規則:`12`
|
||||
- critical:`1`
|
||||
- high:`5`
|
||||
- medium:`5`
|
||||
- low:`1`
|
||||
- action_required:`8`
|
||||
- planned_next:`3`
|
||||
- accepted:`1`
|
||||
|
||||
核心裁決:
|
||||
|
||||
- CVE / advisory / license database 查詢仍未批准;P1-204 只建立政策與批准邊界。
|
||||
- OpenClaw 負責 critical / high 風險仲裁與批准包判定。
|
||||
- Hermes 負責 read-only drift、freshness、manifest / Dockerfile 證據彙整。
|
||||
- Nemotron 可作離線比較與專家建議,不得接手生產裁決、SDK 安裝、shadow / canary 或生產路由。
|
||||
- Python manifest drift、Python reproducibility gap、JS caret range、shared-types publish boundary、Docker digest pin、kubectl checksum、build-time network fetch、Web healthcheck gap 都已標為 action_required。
|
||||
|
||||
實作邊界:
|
||||
|
||||
- 不查外部 CVE / advisory。
|
||||
- 不查外部 license database。
|
||||
- 不安裝或升級套件。
|
||||
- 不寫 lockfile。
|
||||
- 不執行 `npm audit` 或 `pnpm install`。
|
||||
- 不執行 `docker build`、不 pull image、不 rebuild image、不 push registry。
|
||||
- 不呼叫付費 API。
|
||||
- 不建立 shadow / canary。
|
||||
- 不改生產路由。
|
||||
|
||||
驗證:
|
||||
|
||||
- Dependency risk policy schema 驗證通過。
|
||||
- Dependency risk policy service + API tests `9 passed`。
|
||||
- `py_compile` 通過。
|
||||
|
||||
### P1-205 定期依賴漂移與外部資料來源檢查設計摘要
|
||||
|
||||
正式 JSON Schema:
|
||||
|
||||
- `docs/schemas/dependency_drift_check_plan_v1.schema.json`
|
||||
|
||||
正式 JSON Snapshot:
|
||||
|
||||
- `docs/evaluations/dependency_drift_check_plan_2026-06-04.json`
|
||||
|
||||
API:
|
||||
|
||||
- `GET /api/v1/agents/dependency-drift-check-plan`
|
||||
|
||||
設計內容:
|
||||
|
||||
- Cadence items:`5`
|
||||
- Repo-only local checks:`5`
|
||||
- 外部來源候選:`10`
|
||||
- 外部來源候選涵蓋 CVE、license、PyPI / npm registry freshness、Docker / GHCR manifest freshness、AI Agent 官方 release / benchmark signal。
|
||||
- AI Agent 市場監控已納入同一個來源批准模型;Nemotron 仍只做 committed snapshot freshness 與離線比較,不做替換裁決。
|
||||
|
||||
核心裁決:
|
||||
|
||||
- P1-205 只建立 read-only design,不啟用排程。
|
||||
- Local checks 可設計為 repo-only:Python manifest drift、JS lockfile drift、Dockerfile surface drift、dependency policy consistency、agent market snapshot freshness。
|
||||
- 外部 CVE / license / registry / Agent market 來源全部維持 approval_required。
|
||||
- 成功檢查預設不即時通知;失敗、schema mismatch、來源過期、rate-limit exhaustion、成本邊界不明或 high/critical policy hit 才通知 AwoooP / Telegram。
|
||||
|
||||
實作邊界:
|
||||
|
||||
- 不啟用排程。
|
||||
- 不寫 Gitea workflow。
|
||||
- 不查外部 CVE / advisory。
|
||||
- 不查外部 license database。
|
||||
- 不查外部 registry 或 Agent market 來源。
|
||||
- 不安裝 SDK、不呼叫付費 API。
|
||||
- 不安裝或升級套件。
|
||||
- 不寫 lockfile。
|
||||
- 不執行 `docker build`、不 pull image、不 rebuild image、不 push registry。
|
||||
- 不建立 shadow / canary。
|
||||
- 不改生產路由。
|
||||
|
||||
驗證:
|
||||
|
||||
- Dependency drift check plan schema 驗證通過。
|
||||
- Dependency drift check plan service + API tests `9 passed`。
|
||||
- `py_compile` 通過。
|
||||
|
||||
### P1-206 依賴升級批准包模板摘要
|
||||
|
||||
正式 JSON Schema:
|
||||
|
||||
- `docs/schemas/dependency_upgrade_approval_package_template_v1.schema.json`
|
||||
|
||||
正式 JSON Snapshot:
|
||||
|
||||
- `docs/evaluations/dependency_upgrade_approval_package_template_2026-06-04.json`
|
||||
|
||||
API:
|
||||
|
||||
- `GET /api/v1/agents/dependency-upgrade-approval-package-template`
|
||||
|
||||
模板內容:
|
||||
|
||||
- 批准包模板:`8`
|
||||
- Python:`2`
|
||||
- JavaScript:`2`
|
||||
- Docker:`3`
|
||||
- External sources / Agent market:`1`
|
||||
- 8 類模板全部要求 OpenClaw 仲裁與 HITL。
|
||||
|
||||
覆蓋範圍:
|
||||
|
||||
- Python manifest authority。
|
||||
- Python lockfile / constraints policy。
|
||||
- JavaScript high-impact dependency upgrade。
|
||||
- shared-types publish boundary。
|
||||
- Docker base image digest pin。
|
||||
- Docker binary checksum / signature。
|
||||
- Docker build-time network source policy。
|
||||
- CVE / license / registry / AI Agent market external source activation。
|
||||
|
||||
實作邊界:
|
||||
|
||||
- 不安裝或升級套件。
|
||||
- 不寫 manifest / lockfile / Dockerfile。
|
||||
- 不執行 `docker build`、不 pull image、不 rebuild image、不 push registry。
|
||||
- 不 publish package。
|
||||
- 不啟用外部來源。
|
||||
- 不安裝 SDK、不呼叫付費 API。
|
||||
- 不建立 shadow / canary。
|
||||
- 不改生產路由。
|
||||
|
||||
驗證:
|
||||
|
||||
- Dependency upgrade approval package template schema 驗證通過。
|
||||
- Dependency upgrade approval package template service + API tests `9 passed`。
|
||||
- `py_compile` 通過。
|
||||
|
||||
### P1-103 備份通知政策摘要
|
||||
|
||||
正式 JSON Schema:
|
||||
|
||||
- `docs/schemas/backup_notification_policy_v1.schema.json`
|
||||
|
||||
正式 JSON Snapshot:
|
||||
|
||||
- `docs/evaluations/backup_notification_policy_2026-06-04.json`
|
||||
|
||||
API:
|
||||
|
||||
- `GET /api/v1/agents/backup-notification-policy`
|
||||
|
||||
政策內容:
|
||||
|
||||
- 通知規則:`8`
|
||||
- 成功即時抑制:`2`
|
||||
- failure / warning / core blocker 立即升級:`4`
|
||||
- action-required:`2`
|
||||
- 每日摘要時間:台北時間 `06:05`
|
||||
|
||||
核心裁決:
|
||||
|
||||
- 成功備份與 offsite verify 成功不即時發 Telegram / AwoooP,避免洗版。
|
||||
- 成功證據由 Prometheus / textfile、`backup-status.sh --no-notify` 與每日摘要承載。
|
||||
- warning、failed、core blocker、offsite verify failure 必須升級到 AwoooP / Telegram 並帶 evidence。
|
||||
- credential escrow marker 缺口與 metric binding gap 只建立 action-required;不得自動寫 marker 或改 Prometheus rule。
|
||||
|
||||
實作邊界:
|
||||
|
||||
- 不送通知。
|
||||
- 不執行 backup / restore / offsite sync。
|
||||
- 不寫 credential marker。
|
||||
- 不改排程、不寫 workflow。
|
||||
- 不發 Telegram 測試訊息。
|
||||
|
||||
驗證:
|
||||
|
||||
- Backup notification policy schema 驗證通過。
|
||||
- Backup notification policy service + API tests `9 passed`。
|
||||
- `py_compile` 通過。
|
||||
|
||||
### P0 - 治理與 Inventory 基礎
|
||||
|
||||
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|
||||
|---|---|---:|---|---|---|---|
|
||||
| P0-001 | 完成 | 100 | Hermes | 建立完整工作清單與分析 MD | `docs/ai/AI_AGENT_AUTOMATION_WORKLIST_2026-06-04.md` | 可提交 operator review |
|
||||
| P0-002 | 完成 | 100 | Hermes + OpenClaw | 定義自動化狀態分類 | 本文件第 6 節 | 無 runtime 操作 |
|
||||
| P0-003 | 完成 | 100 | Hermes | 定義資產盤點 schema | `docs/schemas/ai_agent_automation_inventory_snapshot_v1.schema.json` | 只讀 |
|
||||
| P0-004 | 完成 | 100 | OpenClaw | 定義每類操作的權限矩陣 | 本文件第 8 節與 `docs/schemas/ai_agent_action_permission_matrix_v1.schema.json` | HITL 邊界明確 |
|
||||
| P0-005 | 完成 | 100 | Hermes | 從 repo / runbook 建立靜態盤點種子 | `docs/evaluations/ai_agent_automation_inventory_snapshot_2026-06-04_static_seed.json` | 不修改 live 環境 |
|
||||
| P0-006 | 完成 | 100 | OpenClaw | 建立只讀自動化盤點 API | `GET /api/v1/agents/automation-inventory-snapshot` | 只讀端點 |
|
||||
| P0-007 | 完成 | 100 | Hermes | 建立治理 / AwoooP UI 看板骨架 | `/zh-TW/governance?tab=automation-inventory` | i18n + mobile 檢查 |
|
||||
| P0-008 | 完成 | 100 | OpenClaw | 補 schema / API / UI 驗證 | API / service tests + browser checks | 不以純 mock 宣稱完成 |
|
||||
|
||||
### P1 - 服務與 Runtime 監控
|
||||
|
||||
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|
||||
|---|---|---:|---|---|---|---|
|
||||
| P1-001 | 待辦 | 0 | OpenClaw | 盤點 API / Web / Worker / K8s runtime surface | K8s / 服務矩陣 | 只讀 |
|
||||
| P1-002 | 待辦 | 0 | Hermes | 盤點 Gitea 工作流程與 runner 健康合約 | 工作流程 / runner 矩陣 | 不修改工作流程 |
|
||||
| P1-003 | 待辦 | 0 | Hermes | 盤點 Prometheus / Alertmanager / SigNoz / Grafana 監控合約 | 可觀測性矩陣 | 只讀 |
|
||||
| P1-004 | 待辦 | 0 | OpenClaw | 盤點 AI Router / Ollama / Nemotron / Gemini provider 路徑 | 推理路由矩陣 | 不切 provider |
|
||||
| P1-005 | 待辦 | 0 | OpenClaw | 偵測服務健康缺口與過期端點 | 需處置清單 | 不重啟 |
|
||||
| P1-006 | 待辦 | 0 | Hermes | 在 UI 顯示 service health 證據卡 | 狀態卡 | 瀏覽器驗證 |
|
||||
| P1-007 | 待辦 | 0 | OpenClaw | 建立 service health 失敗限定 Telegram / AwoooP 對應 | 通知合約 | 不發成功洗版 |
|
||||
|
||||
### P1 - 備份與 DR 自動化
|
||||
|
||||
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|
||||
|---|---|---:|---|---|---|---|
|
||||
| P1-101 | 完成 | 100 | Hermes | 把備份 runbook / 腳本轉成機器可讀目標盤點 | `docs/evaluations/backup_dr_target_inventory_2026-06-04.json` | 只讀 |
|
||||
| P1-102 | 完成 | 100 | OpenClaw | 顯示備份新鮮度、完整性、復原演練狀態 | `docs/evaluations/backup_dr_readiness_matrix_2026-06-04.json` | 不執行 restore |
|
||||
| P1-103 | 完成 | 100 | Hermes | 對齊備份通知政策 | `docs/evaluations/backup_notification_policy_2026-06-04.json` | 不發成功洗版 |
|
||||
| P1-104 | 待辦 | 0 | OpenClaw | 在 AwoooP / governance UI 加備份證據 | 備份卡片 | 瀏覽器驗證 |
|
||||
| P1-105 | 待辦 | 0 | OpenClaw | 定義復原演練批准包 | 復原計畫範本 | 人工批准 |
|
||||
| P1-106 | 待辦 | 0 | Hermes | 顯示異地 / escrow 準備度狀態 | DR 準備度區塊 | 不暴露 credential |
|
||||
|
||||
### P1 - 套件與供應鏈自動化
|
||||
|
||||
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|
||||
|---|---|---:|---|---|---|---|
|
||||
| P1-201 | 完成 | 100 | Hermes | 盤點 Python 依賴 | `docs/evaluations/package_supply_chain_inventory_2026-06-04.json` | 只讀 |
|
||||
| P1-202 | 完成 | 100 | Hermes | 盤點 pnpm/npm 依賴 | `docs/evaluations/javascript_package_inventory_2026-06-04.json` | 只讀 |
|
||||
| P1-203 | 完成 | 100 | Hermes | 盤點 Docker base image 與建置表面 | `docs/evaluations/docker_build_surface_inventory_2026-06-04.json` | 只讀 |
|
||||
| P1-204 | 完成 | 100 | OpenClaw | 定義 CVE / license / drift 嚴重度對應 | `docs/evaluations/dependency_risk_policy_2026-06-04.json` | 只讀政策 |
|
||||
| P1-205 | 完成 | 100 | Hermes | 建立定期依賴漂移檢查 | `docs/evaluations/dependency_drift_check_plan_2026-06-04.json` | 只讀設計 |
|
||||
| P1-206 | 完成 | 100 | OpenClaw | 產生升級批准包 | `docs/evaluations/dependency_upgrade_approval_package_template_2026-06-04.json` | 只讀模板 |
|
||||
|
||||
### P1 - Agent 自動化待辦產品面
|
||||
|
||||
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|
||||
|---|---|---:|---|---|---|---|
|
||||
| P1-301 | 完成 | 100 | Hermes | 定義自動化待辦 schema | `docs/schemas/ai_agent_automation_backlog_v1.schema.json` | 只讀 |
|
||||
| P1-302 | 完成 | 100 | OpenClaw | 從盤點 + 健康 + 市場佇列產生待辦 | `docs/evaluations/ai_agent_automation_backlog_2026-06-04.json` | 不執行 |
|
||||
| P1-303 | 完成 | 100 | Hermes | 建立待辦只讀 API | `GET /api/v1/agents/automation-backlog-snapshot` | 測試 |
|
||||
| P1-304 | 完成 | 100 | Hermes | 建立 P0/P1/P2/P3 分組 UI 看板 | `/zh-TW/governance?tab=automation-inventory` | i18n + mobile |
|
||||
| P1-305 | 待辦 | 0 | OpenClaw | 顯示每個任務的批准邊界 | UI / 操作中繼資料 | 無執行按鈕 |
|
||||
| P1-306 | 待辦 | 0 | Hermes | 顯示進度百分比彙總 | 整體 + 各工作流百分比 | 確定性公式 |
|
||||
|
||||
### P2 - 配置優化
|
||||
|
||||
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|
||||
|---|---|---:|---|---|---|---|
|
||||
| P2-001 | 待辦 | 0 | OpenClaw | K8s requests / limits 建議引擎 | 只讀建議快照 | 不 apply |
|
||||
| P2-002 | 待辦 | 0 | Hermes | CronJob 排程碰撞分析 | 排程優化報告 | 不改排程 |
|
||||
| P2-003 | 待辦 | 0 | Hermes | Prometheus 告警噪音調整提案 | 告警規則建議 | 人工批准 |
|
||||
| P2-004 | 待辦 | 0 | OpenClaw | AI Router / provider 成本與 fallback 優化 | 模型路由建議 | 費用批准 |
|
||||
| P2-005 | 待辦 | 0 | Nemotron | 針對回放 fixture 做離線模型 / prompt 比較 | 模型評分報告 | 未批准不得外部呼叫 |
|
||||
| P2-006 | 待辦 | 0 | Hermes | 前端 bundle / route 健康建議 | Web 優化報告 | 不做無關 redesign |
|
||||
|
||||
### P2 - 安全執行與學習閉環
|
||||
|
||||
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|
||||
|---|---|---:|---|---|---|---|
|
||||
| P2-101 | 待辦 | 0 | OpenClaw | 定義操作類別權限模型 | 操作政策 schema | HITL 關卡 |
|
||||
| P2-102 | 待辦 | 0 | OpenClaw | 所有候選操作都要有 dry-run 證據 | dry-run 合約 | 不直接 apply |
|
||||
| P2-103 | 待辦 | 0 | Hermes | 把任務結果接回 KM / LOGBOOK / 稽核軌跡 | 證據寫入器 | 不洩漏 secret |
|
||||
| P2-104 | 待辦 | 0 | OpenClaw | 修復 `matched_playbook_id` 學習缺口 | playbook trust 更新 | 測試 + live 證據 |
|
||||
| P2-105 | 待辦 | 0 | OpenClaw | 批准前加入 critic / reviewer 評分 | 多 Agent 評分 | 不自動批准 |
|
||||
|
||||
### P3 - 候選 Agent 擴展
|
||||
|
||||
| ID | 狀態 | % | 負責 Agent | 任務 | 產出 | 關卡 |
|
||||
|---|---|---:|---|---|---|---|
|
||||
| P3-001 | 待辦 | 0 | Nemotron | 刷新 Nemotron 來源證據 | 更新後證據報告 | 僅使用 primary sources |
|
||||
| P3-002 | 待辦 | 0 | Nemotron | 只重跑 5 筆 smoke | smoke 關卡報告 | 需要時先批准外部呼叫 |
|
||||
| P3-003 | 待辦 | 0 | Nemotron | smoke 通過後準備 50 筆回放批准包 | 批准包 | 人工批准 |
|
||||
| P3-004 | 待辦 | 0 | LangGraph | 準備官方 SDK 整合提案 | 依賴 / 費用 / 風險批准包 | SDK 批准 |
|
||||
| P3-005 | 待辦 | 0 | Claude SDK 候選 | 準備真實 Claude 修復回放提案 | 費用 / 資料邊界批准包 | API 批准 |
|
||||
| P3-006 | 待辦 | 0 | OpenClaw | 以同輪 OpenClaw 基準比較所有候選 | 替換決策包 | 不改生產環境 |
|
||||
|
||||
## 10. 需要覆蓋的資產範圍
|
||||
|
||||
### 10.1 服務
|
||||
|
||||
- AWOOOI API
|
||||
- AWOOOI Web
|
||||
- Worker 與排程器
|
||||
- K8s Deployment、Service、Ingress、CronJob、ConfigMap、Secret
|
||||
- AwoooP operator 介面
|
||||
- AI Router 與 provider adapter
|
||||
- OpenClaw / Ollama / Nemotron provider 路徑
|
||||
|
||||
### 10.2 工具
|
||||
|
||||
- Gitea 與 Gitea Actions
|
||||
- Harbor registry
|
||||
- Prometheus、Alertmanager、Grafana
|
||||
- SigNoz / ClickHouse
|
||||
- Sentry
|
||||
- Telegram bot / webhook 鏈路
|
||||
- Langfuse / AI tracing
|
||||
- Open-WebUI
|
||||
- MinIO / Velero
|
||||
- Nginx / Certbot
|
||||
- Ansible role 與 playbook
|
||||
- Node exporter / cAdvisor textfile exporter
|
||||
|
||||
### 10.3 套件與依賴
|
||||
|
||||
- API Python 套件
|
||||
- Web pnpm/npm 套件
|
||||
- Docker base image
|
||||
- K8s image tags
|
||||
- Agent SDK 候選
|
||||
- AI provider 模型版本
|
||||
- 監控 / exporter 腳本
|
||||
|
||||
### 10.4 備份與 DR 目標
|
||||
|
||||
- Gitea
|
||||
- Harbor
|
||||
- AWOOOI PostgreSQL
|
||||
- MOMO PostgreSQL
|
||||
- Langfuse
|
||||
- Monitoring
|
||||
- SigNoz
|
||||
- Open-WebUI
|
||||
- ClawBot Redis
|
||||
- Sentry
|
||||
- K8s resources / Velero
|
||||
- Config 備份
|
||||
- AI artifacts
|
||||
- Public route
|
||||
- 異地同步與 credential escrow
|
||||
|
||||
## 11. 自動化能力矩陣
|
||||
|
||||
| 能力 | OpenClaw | Hermes | Nemotron | 狀態 |
|
||||
|---|---|---|---|---|
|
||||
| 偵測過期的服務健康狀態 | 仲裁嚴重度 | 彙整證據 | 離線比較 pattern | P1 |
|
||||
| 偵測備份新鮮度失敗 | 仲裁操作等級 | 寫 runbook / KM | 非主要角色 | P1 |
|
||||
| 偵測依賴漂移 | 判斷風險關卡 | 產生套件報告 | 比較模型 / 工具版本 | P1 |
|
||||
| 建議 K8s limits | 審查爆炸半徑 | 文件化理由 | 可作離線評估者 | P2 |
|
||||
| 建議告警調整 | 審查風險邊界 | 分析噪音 / 歷史 | 可作評估者 | P2 |
|
||||
| 產生批准包 | 最終守門者 | 起草批准包 | 提供專家評分 | P1 |
|
||||
| 執行生產變更 | 僅批准後可仲裁 | 不可 | 不可 | P3+ |
|
||||
| 替換生產決策核心 | 無自動權限 | 不可 | 不可 | ADR / canary 前仍阻擋 |
|
||||
|
||||
## 12. 進度同步協議
|
||||
|
||||
每次階段更新必須包含:
|
||||
|
||||
```text
|
||||
進度:<整體完成度>%。
|
||||
目前優先級:P<level>。
|
||||
目前任務:<任務 ID 與標題>。
|
||||
狀態變更:<舊狀態> -> <新狀態>。
|
||||
證據:<測試 / 瀏覽器 / schema / API 結果>。
|
||||
阻擋:<無或關卡>。
|
||||
下一步:<next task id>。
|
||||
```
|
||||
|
||||
任何完成宣告前,必須同步更新本文件或後續生成的 JSON 快照。
|
||||
|
||||
## 13. 立即執行順序
|
||||
|
||||
1. P1-104:在 AwoooP / governance UI 加備份證據。
|
||||
2. P1-105:定義復原演練批准包。
|
||||
3. P1-106:顯示異地 / escrow 準備度狀態。
|
||||
4. P1-305 / P1-306:補每個任務的批准邊界與進度彙總細節。
|
||||
5. P2 / P3 必須等 P1 可見且關卡穩定後再做。
|
||||
|
||||
## 14. 目前風險
|
||||
|
||||
| 風險 | 嚴重度 | 原因 | 緩解 |
|
||||
|---|---|---|---|
|
||||
| 範圍蔓延到生產執行 | 高 | 工作清單橫跨服務 / 工具 / 備份 / 套件 | P0/P1 保持只讀 |
|
||||
| SDK/API 費用邊界違規 | 高 | 候選 Agent 可能需要外部 SDK/API | 呼叫或安裝前先產批准包 |
|
||||
| runtime 假設過期 | 高 | repo 文件可能和 live runtime 不一致 | 宣告完成前驗 API / 瀏覽器 / 部署證據 |
|
||||
| 備份狀態漂移 | 中 | 現有備份文件可能舊於 live 狀態 | 綠燈前使用 exporter 與 live 檢查 |
|
||||
| UI 過度膨脹 | 中 | governance 頁面會變得太密 | 使用分組卡片與篩選看板 |
|
||||
| 過度信任單一 Agent | 高 | 專家輸出可能錯 | OpenClaw 仲裁 + critic / reviewer 評分 |
|
||||
|
||||
## 15. 下一個里程碑的完成條件
|
||||
|
||||
P0 完成條件:
|
||||
|
||||
- 自動化盤點 schema 存在。
|
||||
- 靜態盤點種子存在。
|
||||
- 只讀 API 可回傳盤點快照。
|
||||
- UI 顯示服務 / 工具 / 套件 / 備份目標與狀態 / 關卡。
|
||||
- 測試通過。
|
||||
- 瀏覽器桌面與 390px mobile 通過。
|
||||
- 沒有生產寫入、SDK 安裝、付費 API 呼叫、路由變更。
|
||||
292
docs/ai/agent-market-capability-evidence-2026-06-01.json
Normal file
292
docs/ai/agent-market-capability-evidence-2026-06-01.json
Normal file
@@ -0,0 +1,292 @@
|
||||
{
|
||||
"schema_version": "agent_market_capability_evidence_v1",
|
||||
"updated_at": "2026-06-01",
|
||||
"baseline_candidate_id": "openclaw_incumbent",
|
||||
"scoring_version": "market_capability_v1",
|
||||
"dimensions": {
|
||||
"durable_execution": 0.15,
|
||||
"human_in_loop": 0.14,
|
||||
"tool_guardrails": 0.14,
|
||||
"observability_tracing": 0.12,
|
||||
"evaluation_harness": 0.12,
|
||||
"mcp_tool_ecosystem": 0.1,
|
||||
"local_private_deploy": 0.08,
|
||||
"code_remediation_fit": 0.08,
|
||||
"awoooi_integration_fit": 0.07
|
||||
},
|
||||
"candidates": [
|
||||
{
|
||||
"candidate_id": "openclaw_incumbent",
|
||||
"display_name": "OpenClaw incumbent",
|
||||
"evaluation_priority": "baseline",
|
||||
"capabilities": {
|
||||
"durable_execution": 1,
|
||||
"human_in_loop": 3,
|
||||
"tool_guardrails": 2,
|
||||
"observability_tracing": 2,
|
||||
"evaluation_harness": 1,
|
||||
"mcp_tool_ecosystem": 2,
|
||||
"local_private_deploy": 3,
|
||||
"code_remediation_fit": 1,
|
||||
"awoooi_integration_fit": 3
|
||||
},
|
||||
"official_sources": [
|
||||
{
|
||||
"title": "AWOOOI incumbent baseline snapshot",
|
||||
"url": "docs/evaluations/openclaw_incumbent_baseline_2026-06-01.json",
|
||||
"evidence": "Current production baseline and local integration evidence."
|
||||
}
|
||||
],
|
||||
"risks": [
|
||||
"Current baseline failed the false repair hard gate.",
|
||||
"Evaluation harness and durable execution are weaker than several market frameworks."
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "openai_agents_sdk_coordinator",
|
||||
"display_name": "OpenAI Agents SDK Coordinator",
|
||||
"evaluation_priority": "must_test",
|
||||
"capabilities": {
|
||||
"durable_execution": 2,
|
||||
"human_in_loop": 3,
|
||||
"tool_guardrails": 3,
|
||||
"observability_tracing": 3,
|
||||
"evaluation_harness": 3,
|
||||
"mcp_tool_ecosystem": 3,
|
||||
"local_private_deploy": 1,
|
||||
"code_remediation_fit": 2,
|
||||
"awoooi_integration_fit": 3
|
||||
},
|
||||
"official_sources": [
|
||||
{
|
||||
"title": "OpenAI Agents SDK tracing",
|
||||
"url": "https://openai.github.io/openai-agents-python/tracing/",
|
||||
"evidence": "Built-in tracing covers agent runs, model generations, tool calls, handoffs, guardrails, and custom events."
|
||||
},
|
||||
{
|
||||
"title": "OpenAI Agents SDK guardrails",
|
||||
"url": "https://openai.github.io/openai-agents-js/guides/guardrails",
|
||||
"evidence": "Tool guardrails can validate or block custom tool calls before and after execution."
|
||||
}
|
||||
],
|
||||
"risks": [
|
||||
"Cloud dependency and sensitive trace handling must pass AWOOOI privacy gates.",
|
||||
"Built-in hosted execution tools need separate guardrail validation."
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "nemo_nemotron_fabric",
|
||||
"display_name": "NVIDIA NeMo Agent Toolkit + Nemotron Fabric",
|
||||
"evaluation_priority": "must_test",
|
||||
"capabilities": {
|
||||
"durable_execution": 2,
|
||||
"human_in_loop": 2,
|
||||
"tool_guardrails": 2,
|
||||
"observability_tracing": 3,
|
||||
"evaluation_harness": 3,
|
||||
"mcp_tool_ecosystem": 3,
|
||||
"local_private_deploy": 3,
|
||||
"code_remediation_fit": 1,
|
||||
"awoooi_integration_fit": 3
|
||||
},
|
||||
"official_sources": [
|
||||
{
|
||||
"title": "NVIDIA NeMo Agent Toolkit overview",
|
||||
"url": "https://docs.nvidia.com/nemo/agent-toolkit/latest/index.html",
|
||||
"evidence": "Framework-agnostic agent toolkit with profiling, observability, evaluation, and MCP support."
|
||||
},
|
||||
{
|
||||
"title": "NVIDIA NeMo Agent Toolkit evaluation",
|
||||
"url": "https://docs.nvidia.com/nemo/agent-toolkit/latest/workflows/evaluate.html",
|
||||
"evidence": "nat eval produces workflow outputs, evaluator outputs, profiling metrics, and request traces."
|
||||
}
|
||||
],
|
||||
"risks": [
|
||||
"Needs AWOOOI-specific HITL and dangerous-action policy integration.",
|
||||
"GPU/NIM operating cost must be compared against current local inference."
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "microsoft_agent_framework",
|
||||
"display_name": "Microsoft Agent Framework",
|
||||
"evaluation_priority": "can_test",
|
||||
"capabilities": {
|
||||
"durable_execution": 3,
|
||||
"human_in_loop": 3,
|
||||
"tool_guardrails": 2,
|
||||
"observability_tracing": 3,
|
||||
"evaluation_harness": 2,
|
||||
"mcp_tool_ecosystem": 3,
|
||||
"local_private_deploy": 2,
|
||||
"code_remediation_fit": 1,
|
||||
"awoooi_integration_fit": 2
|
||||
},
|
||||
"official_sources": [
|
||||
{
|
||||
"title": "Microsoft Agent Framework overview",
|
||||
"url": "https://learn.microsoft.com/en-us/agent-framework/overview/",
|
||||
"evidence": "Combines agents, graph workflows, session state, middleware, telemetry, MCP clients, checkpointing, and HITL."
|
||||
}
|
||||
],
|
||||
"risks": [
|
||||
"Public preview status and Microsoft ecosystem fit must be assessed.",
|
||||
"Python/FastAPI/K8s integration cost is likely higher than LangGraph or NeMo."
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "langgraph_incident_kernel",
|
||||
"display_name": "LangGraph Incident Kernel",
|
||||
"evaluation_priority": "must_test",
|
||||
"capabilities": {
|
||||
"durable_execution": 3,
|
||||
"human_in_loop": 3,
|
||||
"tool_guardrails": 2,
|
||||
"observability_tracing": 2,
|
||||
"evaluation_harness": 2,
|
||||
"mcp_tool_ecosystem": 2,
|
||||
"local_private_deploy": 3,
|
||||
"code_remediation_fit": 1,
|
||||
"awoooi_integration_fit": 3
|
||||
},
|
||||
"official_sources": [
|
||||
{
|
||||
"title": "LangGraph persistence",
|
||||
"url": "https://docs.langchain.com/oss/python/langgraph/persistence",
|
||||
"evidence": "Checkpoint persistence supports human-in-the-loop, memory, time travel debugging, and fault-tolerant execution."
|
||||
},
|
||||
{
|
||||
"title": "LangGraph interrupts",
|
||||
"url": "https://docs.langchain.com/oss/python/langgraph/human-in-the-loop",
|
||||
"evidence": "Interrupts pause graph execution and resume through persisted graph state."
|
||||
}
|
||||
],
|
||||
"risks": [
|
||||
"It is a workflow kernel, not a smarter model by itself.",
|
||||
"Tool safety and evaluation metrics must be implemented by AWOOOI adapters."
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "claude_agent_sdk_remediator",
|
||||
"display_name": "Claude Agent SDK Remediator",
|
||||
"evaluation_priority": "must_test",
|
||||
"capabilities": {
|
||||
"durable_execution": 2,
|
||||
"human_in_loop": 3,
|
||||
"tool_guardrails": 3,
|
||||
"observability_tracing": 2,
|
||||
"evaluation_harness": 1,
|
||||
"mcp_tool_ecosystem": 3,
|
||||
"local_private_deploy": 1,
|
||||
"code_remediation_fit": 3,
|
||||
"awoooi_integration_fit": 2
|
||||
},
|
||||
"official_sources": [
|
||||
{
|
||||
"title": "Claude Agent SDK loop",
|
||||
"url": "https://platform.claude.com/docs/en/agent-sdk/agent-loop",
|
||||
"evidence": "Embeds Claude Code's autonomous agent loop with programmatic control over tools, permissions, cost limits, and output."
|
||||
},
|
||||
{
|
||||
"title": "Claude Agent SDK overview",
|
||||
"url": "https://docs.claude.com/es/api/agent-sdk/overview",
|
||||
"evidence": "SDK exposes context management, file operations, code execution, MCP, permissions, sessions, and monitoring."
|
||||
}
|
||||
],
|
||||
"risks": [
|
||||
"Best fit is code and DevOps remediation, not necessarily central incident arbitration.",
|
||||
"API cost, subscription separation, and vendor boundary must be validated."
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "claude_managed_agents_sandbox",
|
||||
"display_name": "Claude Managed Agents Sandbox",
|
||||
"evaluation_priority": "can_test",
|
||||
"capabilities": {
|
||||
"durable_execution": 3,
|
||||
"human_in_loop": 2,
|
||||
"tool_guardrails": 3,
|
||||
"observability_tracing": 2,
|
||||
"evaluation_harness": 1,
|
||||
"mcp_tool_ecosystem": 2,
|
||||
"local_private_deploy": 2,
|
||||
"code_remediation_fit": 3,
|
||||
"awoooi_integration_fit": 2
|
||||
},
|
||||
"official_sources": [
|
||||
{
|
||||
"title": "Claude Managed Agents quickstart",
|
||||
"url": "https://platform.claude.com/docs/en/managed-agents/quickstart",
|
||||
"evidence": "Defines agents, environments, sessions, events, and pre-built agent tools for autonomous sessions."
|
||||
}
|
||||
],
|
||||
"risks": [
|
||||
"Managed service and beta header make it less suitable as the first AWOOOI core replacement.",
|
||||
"Sandbox placement, data retention, and cost must be reviewed before shadow mode."
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "google_adk_stack",
|
||||
"display_name": "Google Agent Development Kit Stack",
|
||||
"evaluation_priority": "can_test",
|
||||
"capabilities": {
|
||||
"durable_execution": 3,
|
||||
"human_in_loop": 2,
|
||||
"tool_guardrails": 2,
|
||||
"observability_tracing": 2,
|
||||
"evaluation_harness": 3,
|
||||
"mcp_tool_ecosystem": 2,
|
||||
"local_private_deploy": 2,
|
||||
"code_remediation_fit": 1,
|
||||
"awoooi_integration_fit": 2
|
||||
},
|
||||
"official_sources": [
|
||||
{
|
||||
"title": "Google ADK technical overview",
|
||||
"url": "https://google.github.io/adk-docs/get-started/about/",
|
||||
"evidence": "ADK includes session management, state, events, memory, artifacts, evaluation, and developer UI."
|
||||
},
|
||||
{
|
||||
"title": "Google ADK sessions",
|
||||
"url": "https://google.github.io/adk-docs/sessions/session/",
|
||||
"evidence": "Runner retrieves sessions and exposes state/events to agents."
|
||||
}
|
||||
],
|
||||
"risks": [
|
||||
"Gemini/Vertex ecosystem dependency must be justified against current local-first policy.",
|
||||
"AIOps tool safety and rollback gates still need AWOOOI-specific implementation."
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "crewai_flows_crews",
|
||||
"display_name": "CrewAI Flows + Crews",
|
||||
"evaluation_priority": "secondary",
|
||||
"capabilities": {
|
||||
"durable_execution": 2,
|
||||
"human_in_loop": 2,
|
||||
"tool_guardrails": 2,
|
||||
"observability_tracing": 2,
|
||||
"evaluation_harness": 1,
|
||||
"mcp_tool_ecosystem": 2,
|
||||
"local_private_deploy": 3,
|
||||
"code_remediation_fit": 1,
|
||||
"awoooi_integration_fit": 1
|
||||
},
|
||||
"official_sources": [
|
||||
{
|
||||
"title": "CrewAI documentation",
|
||||
"url": "https://docs.crewai.com/",
|
||||
"evidence": "Docs describe agents, crews, flows, guardrails, memory, knowledge, and observability."
|
||||
},
|
||||
{
|
||||
"title": "CrewAI Flows",
|
||||
"url": "https://www.crewai.com/crewai-flows",
|
||||
"evidence": "Flows coordinate tasks and crews with structured, event-driven workflows and state management."
|
||||
}
|
||||
],
|
||||
"risks": [
|
||||
"Better for rapid automation teams than high-risk production AIOps core.",
|
||||
"Durability, strict audit, and permission boundary must be proven in replay."
|
||||
]
|
||||
}
|
||||
]
|
||||
}
|
||||
357
docs/ai/agent-market-watch-sources.v1.json
Normal file
357
docs/ai/agent-market-watch-sources.v1.json
Normal file
@@ -0,0 +1,357 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"schema_version": "agent_market_watch_sources_v1",
|
||||
"updated_at": "2026-06-04",
|
||||
"purpose": "Primary-source watch list for recurring AI Agent market updates. A change here is not replacement approval; it only triggers refreshed evaluation.",
|
||||
"cadence": {
|
||||
"weekly_market_watch": "Every Monday 09:00 Asia/Taipei, produce a read-only market watch report and full-scope integration/discovery review summary.",
|
||||
"monthly_integration_review": "After operator review, commit a reviewed baseline for market watch, integration review, and discovery intake.",
|
||||
"trigger_on_major_version": true
|
||||
},
|
||||
"policy": {
|
||||
"replacement_decision_allowed": false,
|
||||
"integration_requires_replay": true,
|
||||
"paid_provider_requires_approval": true,
|
||||
"new_dependency_requires_approval": true,
|
||||
"raw_external_pages_committed": false,
|
||||
"official_or_primary_sources_only": true
|
||||
},
|
||||
"candidates": [
|
||||
{
|
||||
"candidate_id": "openai_agents_sdk_coordinator",
|
||||
"display_name": "OpenAI Agents SDK Coordinator",
|
||||
"evaluation_priority": "must_test",
|
||||
"recommended_role": "Coordinator / Orchestrator",
|
||||
"requires_cost_approval": true,
|
||||
"requires_dependency_approval": true,
|
||||
"sources": [
|
||||
{
|
||||
"source_id": "openai_agents_docs",
|
||||
"type": "docs",
|
||||
"url": "https://developers.openai.com/api/docs/guides/agents",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "openai_agent_builder_safety_docs",
|
||||
"type": "docs",
|
||||
"url": "https://developers.openai.com/api/docs/guides/agent-builder-safety",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "openai_agents_python_pypi",
|
||||
"type": "pypi",
|
||||
"url": "https://pypi.org/pypi/openai-agents/json",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "openai_agents_typescript_npm",
|
||||
"type": "npm",
|
||||
"url": "https://registry.npmjs.org/@openai%2Fagents",
|
||||
"reference_version": null
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "langgraph_incident_kernel",
|
||||
"display_name": "LangGraph Incident Kernel",
|
||||
"evaluation_priority": "must_test",
|
||||
"recommended_role": "Durable Incident Workflow Kernel",
|
||||
"requires_cost_approval": false,
|
||||
"requires_dependency_approval": true,
|
||||
"sources": [
|
||||
{
|
||||
"source_id": "langgraph_docs",
|
||||
"type": "docs",
|
||||
"url": "https://docs.langchain.com/oss/python/langgraph/overview",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "langgraph_pypi",
|
||||
"type": "pypi",
|
||||
"url": "https://pypi.org/pypi/langgraph/json",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "langgraph_github_release",
|
||||
"type": "github_release",
|
||||
"url": "https://api.github.com/repos/langchain-ai/langgraph/releases/latest",
|
||||
"reference_version": null
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "nemo_nemotron_fabric",
|
||||
"display_name": "NVIDIA NeMo Agent Toolkit + Nemotron Fabric",
|
||||
"evaluation_priority": "must_test",
|
||||
"recommended_role": "Agent Fabric / Tool-Model Evaluator",
|
||||
"requires_cost_approval": true,
|
||||
"requires_dependency_approval": true,
|
||||
"sources": [
|
||||
{
|
||||
"source_id": "nvidia_nemo_agent_toolkit_docs",
|
||||
"type": "docs",
|
||||
"url": "https://docs.nvidia.com/nemo/agent-toolkit/latest/index.html",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "nvidia_nim_llm_docs",
|
||||
"type": "docs",
|
||||
"url": "https://docs.nvidia.com/nim/large-language-models/latest/index.html",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "nvidia_build_models",
|
||||
"type": "docs",
|
||||
"url": "https://build.nvidia.com/models",
|
||||
"reference_version": null
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "claude_agent_sdk_remediator",
|
||||
"display_name": "Claude Agent SDK Remediator",
|
||||
"evaluation_priority": "must_test",
|
||||
"recommended_role": "DevOps / Code Remediation Agent",
|
||||
"requires_cost_approval": true,
|
||||
"requires_dependency_approval": true,
|
||||
"sources": [
|
||||
{
|
||||
"source_id": "claude_agent_sdk_docs",
|
||||
"type": "docs",
|
||||
"url": "https://platform.claude.com/docs/en/agent-sdk/agent-loop",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "anthropic_api_docs",
|
||||
"type": "docs",
|
||||
"url": "https://platform.claude.com/docs/en/home",
|
||||
"reference_version": null
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "google_adk_stack",
|
||||
"display_name": "Google Agent Development Kit Stack",
|
||||
"evaluation_priority": "can_test",
|
||||
"recommended_role": "Google / Gemini Agent Stack",
|
||||
"requires_cost_approval": true,
|
||||
"requires_dependency_approval": true,
|
||||
"sources": [
|
||||
{
|
||||
"source_id": "google_adk_docs",
|
||||
"type": "docs",
|
||||
"url": "https://adk.dev/get-started/about/",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "google_adk_pypi",
|
||||
"type": "pypi",
|
||||
"url": "https://pypi.org/pypi/google-adk/json",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "google_adk_github_release",
|
||||
"type": "github_release",
|
||||
"url": "https://api.github.com/repos/google/adk-python/releases/latest",
|
||||
"reference_version": null
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "microsoft_agent_framework",
|
||||
"display_name": "Microsoft Agent Framework",
|
||||
"evaluation_priority": "can_test",
|
||||
"recommended_role": "Enterprise Workflow Agent Stack",
|
||||
"requires_cost_approval": true,
|
||||
"requires_dependency_approval": true,
|
||||
"sources": [
|
||||
{
|
||||
"source_id": "microsoft_agent_framework_docs",
|
||||
"type": "docs",
|
||||
"url": "https://learn.microsoft.com/en-us/agent-framework/overview/",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "microsoft_agent_framework_github_release",
|
||||
"type": "github_release",
|
||||
"url": "https://api.github.com/repos/microsoft/agent-framework/releases/latest",
|
||||
"reference_version": null
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "crewai_flows_crews",
|
||||
"display_name": "CrewAI Flows + Crews",
|
||||
"evaluation_priority": "secondary",
|
||||
"recommended_role": "Rapid Agent Team Prototype",
|
||||
"requires_cost_approval": false,
|
||||
"requires_dependency_approval": true,
|
||||
"sources": [
|
||||
{
|
||||
"source_id": "crewai_docs",
|
||||
"type": "docs",
|
||||
"url": "https://docs.crewai.com/en/introduction",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "crewai_pypi",
|
||||
"type": "pypi",
|
||||
"url": "https://pypi.org/pypi/crewai/json",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "crewai_github_release",
|
||||
"type": "github_release",
|
||||
"url": "https://api.github.com/repos/crewAIInc/crewAI/releases/latest",
|
||||
"reference_version": null
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "hermes_agent_personal_platform",
|
||||
"display_name": "NousResearch Hermes Agent",
|
||||
"evaluation_priority": "watch_only",
|
||||
"recommended_role": "Personal Agent Platform / Memory-Skills Runtime",
|
||||
"requires_cost_approval": true,
|
||||
"requires_dependency_approval": true,
|
||||
"sources": [
|
||||
{
|
||||
"source_id": "hermes_agent_homepage",
|
||||
"type": "docs",
|
||||
"url": "https://hermes-agent.nousresearch.com",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "hermes_agent_github_release",
|
||||
"type": "github_release",
|
||||
"url": "https://api.github.com/repos/NousResearch/hermes-agent/releases/latest",
|
||||
"reference_version": null
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "microsoft_agent_governance_toolkit",
|
||||
"display_name": "Microsoft Agent Governance Toolkit",
|
||||
"evaluation_priority": "watch_only",
|
||||
"recommended_role": "Agent Governance / Policy Runtime",
|
||||
"requires_cost_approval": false,
|
||||
"requires_dependency_approval": true,
|
||||
"sources": [
|
||||
{
|
||||
"source_id": "microsoft_agent_governance_docs",
|
||||
"type": "docs",
|
||||
"url": "https://microsoft.github.io/agent-governance-toolkit/",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "microsoft_agent_governance_github_release",
|
||||
"type": "github_release",
|
||||
"url": "https://api.github.com/repos/microsoft/agent-governance-toolkit/releases/latest",
|
||||
"reference_version": null
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "thclaws_agent_harness",
|
||||
"display_name": "thClaws Agent Harness",
|
||||
"evaluation_priority": "watch_only",
|
||||
"recommended_role": "Agent Harness / Multi-Provider Runtime",
|
||||
"requires_cost_approval": true,
|
||||
"requires_dependency_approval": true,
|
||||
"sources": [
|
||||
{
|
||||
"source_id": "thclaws_homepage",
|
||||
"type": "docs",
|
||||
"url": "https://thclaws.ai",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "thclaws_github_release",
|
||||
"type": "github_release",
|
||||
"url": "https://api.github.com/repos/thClaws/thClaws/releases/latest",
|
||||
"reference_version": null
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "pydantic_deepagents",
|
||||
"display_name": "Pydantic DeepAgents",
|
||||
"evaluation_priority": "watch_only",
|
||||
"recommended_role": "Pydantic AI Deep Agent Framework",
|
||||
"requires_cost_approval": true,
|
||||
"requires_dependency_approval": true,
|
||||
"sources": [
|
||||
{
|
||||
"source_id": "pydantic_deepagents_docs",
|
||||
"type": "docs",
|
||||
"url": "https://vstorm-co.github.io/pydantic-deepagents/",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "pydantic_deepagents_github_release",
|
||||
"type": "github_release",
|
||||
"url": "https://api.github.com/repos/vstorm-co/pydantic-deepagents/releases/latest",
|
||||
"reference_version": null
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "agentos_framework",
|
||||
"display_name": "AgentOS Framework",
|
||||
"evaluation_priority": "watch_only",
|
||||
"recommended_role": "TypeScript Agent Framework / Orchestrator",
|
||||
"requires_cost_approval": true,
|
||||
"requires_dependency_approval": true,
|
||||
"sources": [
|
||||
{
|
||||
"source_id": "agentos_docs",
|
||||
"type": "docs",
|
||||
"url": "https://agentos.sh",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "agentos_github_release",
|
||||
"type": "github_release",
|
||||
"url": "https://api.github.com/repos/framerslab/agentos/releases/latest",
|
||||
"reference_version": null
|
||||
}
|
||||
]
|
||||
},
|
||||
{
|
||||
"candidate_id": "bernstein_agent_governance",
|
||||
"display_name": "Bernstein Agent Governance",
|
||||
"evaluation_priority": "watch_only",
|
||||
"recommended_role": "Audit-Grade Agent Orchestration / Governance",
|
||||
"requires_cost_approval": true,
|
||||
"requires_dependency_approval": true,
|
||||
"sources": [
|
||||
{
|
||||
"source_id": "bernstein_docs",
|
||||
"type": "docs",
|
||||
"url": "https://bernstein.run",
|
||||
"reference_version": null
|
||||
},
|
||||
{
|
||||
"source_id": "bernstein_github_release",
|
||||
"type": "github_release",
|
||||
"url": "https://api.github.com/repos/sipyourdrink-ltd/bernstein/releases/latest",
|
||||
"reference_version": null
|
||||
}
|
||||
]
|
||||
}
|
||||
],
|
||||
"discovery_sources": [
|
||||
{
|
||||
"source_id": "github_ai_agent_topic",
|
||||
"type": "github_search",
|
||||
"url": "https://api.github.com/search/repositories?q=topic:ai-agent+stars:%3E500&sort=updated&order=desc",
|
||||
"purpose": "Find new high-signal open-source AI Agent frameworks. Any finding requires manual source classification before integration."
|
||||
},
|
||||
{
|
||||
"source_id": "github_agent_framework_topic",
|
||||
"type": "github_search",
|
||||
"url": "https://api.github.com/search/repositories?q=topic:agent-framework+stars:%3E300&sort=updated&order=desc",
|
||||
"purpose": "Find new agent framework candidates. Any finding requires official-source verification before being added as a candidate."
|
||||
}
|
||||
]
|
||||
}
|
||||
297
docs/ai/agent-replacement-candidates.v1.json
Normal file
297
docs/ai/agent-replacement-candidates.v1.json
Normal file
@@ -0,0 +1,297 @@
|
||||
{
|
||||
"$schema": "https://json-schema.org/draft/2020-12/schema",
|
||||
"schema_version": "agent_replacement_candidates_v1",
|
||||
"updated_at": "2026-06-04",
|
||||
"baseline_candidate_id": "openclaw_incumbent",
|
||||
"fixture_schema": "docs/schemas/agent_replay_fixture_v1.schema.json",
|
||||
"candidate_input_schema": "docs/schemas/agent_replay_candidate_input_v1.schema.json",
|
||||
"candidate_result_schema": "docs/schemas/agent_candidate_replay_result_v1.schema.json",
|
||||
"candidate_contract_report_schema": "docs/schemas/agent_replay_contract_report_v1.schema.json",
|
||||
"candidate_pipeline_report_schema": "docs/schemas/agent_replay_pipeline_report_v1.schema.json",
|
||||
"candidate_promotion_gate_schema": "docs/schemas/agent_replay_promotion_gate_v1.schema.json",
|
||||
"candidate_grading_report_schema": "docs/schemas/agent_replay_grading_report_v1.schema.json",
|
||||
"nemo_nemotron_replay_request_schema": "docs/schemas/agent_nemotron_replay_request_v1.schema.json",
|
||||
"nemo_nemotron_external_result_schema": "docs/schemas/agent_nemotron_external_result_v1.schema.json",
|
||||
"nemo_nemotron_external_runner_report_schema": "docs/schemas/agent_nemotron_external_runner_report_v1.schema.json",
|
||||
"nemo_nemotron_external_runner_preflight_schema": "docs/schemas/agent_nemotron_external_runner_preflight_v1.schema.json",
|
||||
"nemo_nemotron_request_pack_sanitize_schema": "docs/schemas/agent_nemotron_request_pack_sanitize_report_v1.schema.json",
|
||||
"nemo_nemotron_external_runner_readiness_schema": "docs/schemas/agent_nemotron_external_runner_readiness_v1.schema.json",
|
||||
"nemo_nemotron_import_report_schema": "docs/schemas/agent_nemotron_import_report_v1.schema.json",
|
||||
"nemo_nemotron_finalizer_report_schema": "docs/schemas/agent_nemotron_replay_finalizer_report_v1.schema.json",
|
||||
"nemo_nemotron_failure_analysis_schema": "docs/schemas/agent_nemotron_replay_failure_analysis_v1.schema.json",
|
||||
"nemo_nemotron_contract_tuned_smoke_gate_schema": "docs/schemas/agent_nemotron_contract_tuned_smoke_gate_v1.schema.json",
|
||||
"agent_market_watch_report_schema": "docs/schemas/agent_market_watch_report_v1.schema.json",
|
||||
"agent_market_integration_review_schema": "docs/schemas/agent_market_integration_review_v1.schema.json",
|
||||
"agent_market_discovery_review_schema": "docs/schemas/agent_market_discovery_review_v1.schema.json",
|
||||
"agent_market_discovery_classification_schema": "docs/schemas/agent_market_discovery_classification_v1.schema.json",
|
||||
"agent_market_watch_promotion_review_schema": "docs/schemas/agent_market_watch_promotion_review_v1.schema.json",
|
||||
"agent_market_governance_snapshot_schema": "docs/schemas/agent_market_governance_snapshot_v1.schema.json",
|
||||
"agent_market_watch_sources": "docs/ai/agent-market-watch-sources.v1.json",
|
||||
"agent_market_watch_report": "docs/evaluations/agent_market_watch_report_2026-06-04_watch_expanded.json",
|
||||
"agent_market_watch_reviewed_report": "docs/evaluations/agent_market_watch_report_2026-06-02_reviewed.json",
|
||||
"agent_market_integration_review_report": "docs/evaluations/agent_market_integration_review_2026-06-02.json",
|
||||
"agent_market_integration_review_full_report": "docs/evaluations/agent_market_integration_review_full_2026-06-04_watch_expanded.json",
|
||||
"agent_market_discovery_review_report": "docs/evaluations/agent_market_discovery_review_2026-06-04_watch_expanded.json",
|
||||
"agent_market_discovery_classification_report": "docs/evaluations/agent_market_discovery_classification_2026-06-04_watch_expanded.json",
|
||||
"agent_market_watch_promotion_review_report": "docs/evaluations/agent_market_watch_promotion_review_2026-06-04_watch_expanded.json",
|
||||
"agent_market_governance_snapshot_report": "docs/evaluations/agent_market_governance_snapshot_2026-06-04.json",
|
||||
"agent_market_governance_snapshot_api": "GET /api/v1/agents/market-governance-snapshot",
|
||||
"agent_market_governance_snapshot_ui": "/governance?tab=agent-market",
|
||||
"agent_market_governance_snapshot_cadence_field": "evaluation_cadence",
|
||||
"agent_market_governance_snapshot_health_field": "market_watch_health",
|
||||
"agent_market_governance_snapshot_candidate_statuses_field": "candidate_statuses",
|
||||
"agent_market_watch_workflow": ".gitea/workflows/agent-market-watch.yaml",
|
||||
"replay_record_schema": "docs/schemas/agent_replacement_replay_v1.schema.json",
|
||||
"market_capability_evidence": "docs/ai/agent-market-capability-evidence-2026-06-01.json",
|
||||
"market_capability_scorecard": "docs/evaluations/agent_market_capability_scorecard_2026-06-01.json",
|
||||
"fixture_smoke_report": "docs/evaluations/agent_replay_fixture_smoke_2026-06-01.json",
|
||||
"nemo_nemotron_request_pack_smoke_report": "docs/evaluations/agent_nemotron_replay_request_pack_smoke_2026-06-01.json",
|
||||
"nemo_nemotron_external_runner_preflight_report": "docs/evaluations/agent_nemotron_external_runner_preflight_2026-06-01.json",
|
||||
"nemo_nemotron_request_pack_sanitize_report": "docs/evaluations/agent_nemotron_request_pack_sanitize_2026-06-01.json",
|
||||
"nemo_nemotron_external_runner_preflight_sanitized_report": "docs/evaluations/agent_nemotron_external_runner_preflight_sanitized_2026-06-01.json",
|
||||
"nemo_nemotron_external_runner_readiness_report": "docs/evaluations/agent_nemotron_external_runner_readiness_2026-06-01.json",
|
||||
"nemo_nemotron_external_runner_report": "docs/evaluations/agent_nemotron_external_runner_report_2026-06-01.json",
|
||||
"nemo_nemotron_prod_finalizer_report": "docs/evaluations/agent_nemotron_replay_finalizer_prod_2026-06-01.json",
|
||||
"nemo_nemotron_prod_scorecard": "docs/evaluations/agent_nemotron_replay_scorecard_2026-06-01.json",
|
||||
"nemo_nemotron_prod_failure_analysis": "docs/evaluations/agent_nemotron_replay_failure_analysis_2026-06-01.json",
|
||||
"nemo_nemotron_contract_tuned_request_pack_build": "docs/evaluations/agent_nemotron_contract_tuned_request_pack_build_2026-06-01.json",
|
||||
"nemo_nemotron_contract_tuned_preflight": "docs/evaluations/agent_nemotron_contract_tuned_preflight_2026-06-01.json",
|
||||
"nemo_nemotron_contract_tuned_runner_manifest": "docs/evaluations/nemotron_contract_tuned_runner_manifest_2026-06-01.json",
|
||||
"nemo_nemotron_contract_tuned_runner_readiness": "docs/evaluations/agent_nemotron_contract_tuned_runner_readiness_2026-06-01.json",
|
||||
"nemo_nemotron_contract_tuned_smoke_runner_report": "docs/evaluations/agent_nemotron_contract_tuned_smoke_external_runner_report_2026-06-01.json",
|
||||
"nemo_nemotron_contract_tuned_smoke_gate": "docs/evaluations/agent_nemotron_contract_tuned_smoke_gate_2026-06-01.json",
|
||||
"nemo_nemotron_contract_tuned_fast_model_smoke_manifest": "docs/evaluations/nemotron_contract_tuned_fast_model_smoke_manifest_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_fast_model_smoke_readiness": "docs/evaluations/agent_nemotron_contract_tuned_fast_model_smoke_readiness_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_nano9b_smoke_runner_report": "docs/evaluations/agent_nemotron_contract_tuned_nano9b_smoke_external_runner_report_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_nano9b_smoke_gate": "docs/evaluations/agent_nemotron_contract_tuned_nano9b_smoke_gate_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_mini4b_smoke_manifest": "docs/evaluations/nemotron_contract_tuned_mini4b_smoke_manifest_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_mini4b_smoke_readiness": "docs/evaluations/agent_nemotron_contract_tuned_mini4b_smoke_readiness_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_mini4b_smoke_runner_report": "docs/evaluations/agent_nemotron_contract_tuned_mini4b_smoke_external_runner_report_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_mini4b_smoke_gate": "docs/evaluations/agent_nemotron_contract_tuned_mini4b_smoke_gate_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_nemotron3nano30b_smoke_manifest": "docs/evaluations/nemotron_contract_tuned_nemotron3nano30b_smoke_manifest_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_nemotron3nano30b_smoke_readiness": "docs/evaluations/agent_nemotron_contract_tuned_nemotron3nano30b_smoke_readiness_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_nemotron3nano30b_smoke_runner_report": "docs/evaluations/agent_nemotron_contract_tuned_nemotron3nano30b_smoke_external_runner_report_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_nemotron3nano30b_smoke_gate": "docs/evaluations/agent_nemotron_contract_tuned_nemotron3nano30b_smoke_gate_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_49b_v15_smoke_manifest": "docs/evaluations/nemotron_contract_tuned_49b_v15_smoke_manifest_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_49b_v15_smoke_readiness": "docs/evaluations/agent_nemotron_contract_tuned_49b_v15_smoke_readiness_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_49b_v15_smoke_runner_report": "docs/evaluations/agent_nemotron_contract_tuned_49b_v15_smoke_external_runner_report_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_49b_v15_smoke_gate": "docs/evaluations/agent_nemotron_contract_tuned_49b_v15_smoke_gate_2026-06-02.json",
|
||||
"nemo_nemotron_contract_tuned_smoke_matrix": "docs/evaluations/agent_nemotron_contract_tuned_smoke_matrix_2026-06-02.json",
|
||||
"langgraph_replay_adapter_report": "docs/evaluations/agent_langgraph_replay_adapter_report_2026-06-02.json",
|
||||
"langgraph_replay_contract_report": "docs/evaluations/agent_langgraph_replay_contract_2026-06-02.json",
|
||||
"langgraph_replay_grading_report": "docs/evaluations/agent_langgraph_replay_grading_2026-06-02.json",
|
||||
"langgraph_replay_pipeline_report": "docs/evaluations/agent_langgraph_replay_pipeline_2026-06-02.json",
|
||||
"langgraph_replay_scorecard": "docs/evaluations/agent_langgraph_replay_scorecard_2026-06-02.json",
|
||||
"langgraph_replay_promotion_gate": "docs/evaluations/agent_langgraph_replay_promotion_gate_2026-06-02.json",
|
||||
"langgraph_replay_summary": "docs/evaluations/agent_langgraph_replay_summary_2026-06-02.json",
|
||||
"openai_coordinator_replay_adapter_report": "docs/evaluations/agent_openai_coordinator_replay_adapter_report_2026-06-02.json",
|
||||
"openai_coordinator_replay_contract_report": "docs/evaluations/agent_openai_coordinator_replay_contract_2026-06-02.json",
|
||||
"openai_coordinator_replay_grading_report": "docs/evaluations/agent_openai_coordinator_replay_grading_2026-06-02.json",
|
||||
"openai_coordinator_replay_pipeline_report": "docs/evaluations/agent_openai_coordinator_replay_pipeline_2026-06-02.json",
|
||||
"openai_coordinator_replay_scorecard": "docs/evaluations/agent_openai_coordinator_replay_scorecard_2026-06-02.json",
|
||||
"openai_coordinator_replay_promotion_gate": "docs/evaluations/agent_openai_coordinator_replay_promotion_gate_2026-06-02.json",
|
||||
"openai_coordinator_replay_summary": "docs/evaluations/agent_openai_coordinator_replay_summary_2026-06-02.json",
|
||||
"claude_remediator_replay_adapter_report": "docs/evaluations/agent_claude_remediator_replay_adapter_report_2026-06-02.json",
|
||||
"claude_remediator_replay_contract_report": "docs/evaluations/agent_claude_remediator_replay_contract_2026-06-02.json",
|
||||
"claude_remediator_replay_grading_report": "docs/evaluations/agent_claude_remediator_replay_grading_2026-06-02.json",
|
||||
"claude_remediator_replay_pipeline_report": "docs/evaluations/agent_claude_remediator_replay_pipeline_2026-06-02.json",
|
||||
"claude_remediator_replay_scorecard": "docs/evaluations/agent_claude_remediator_replay_scorecard_2026-06-02.json",
|
||||
"claude_remediator_replay_promotion_gate": "docs/evaluations/agent_claude_remediator_replay_promotion_gate_2026-06-02.json",
|
||||
"claude_remediator_replay_summary": "docs/evaluations/agent_claude_remediator_replay_summary_2026-06-02.json",
|
||||
"nemo_nemotron_finalizer_smoke_report": "docs/evaluations/agent_nemotron_replay_finalizer_smoke_2026-06-01.json",
|
||||
"nemo_nemotron_external_runner_manifest": "docs/evaluations/nemotron_external_runner_manifest_2026-06-01.json",
|
||||
"scorecard_cli": "scripts/ai-agent-replay-scorecard.py",
|
||||
"candidate_input_preparer_cli": "scripts/agents/prepare-agent-replay-inputs.py",
|
||||
"candidate_contract_validator_cli": "scripts/agents/validate-agent-replay-contract.py",
|
||||
"candidate_result_normalizer_cli": "scripts/agents/normalize-agent-replay-results.py",
|
||||
"candidate_label_grader_cli": "scripts/agents/grade-agent-replay-results.py",
|
||||
"candidate_pipeline_runner_cli": "scripts/agents/run-agent-replacement-replay.py",
|
||||
"candidate_promotion_gate_cli": "scripts/agents/evaluate-agent-promotion-gate.py",
|
||||
"nemo_nemotron_request_builder_cli": "scripts/agents/nemotron-build-replay-requests.py",
|
||||
"nemo_nemotron_external_runner_cli": "scripts/agents/nemotron-run-external-offline.py",
|
||||
"nemo_nemotron_external_runner_preflight_cli": "scripts/agents/nemotron-external-runner-preflight.py",
|
||||
"nemo_nemotron_request_pack_sanitizer_cli": "scripts/agents/nemotron-sanitize-request-pack.py",
|
||||
"nemo_nemotron_external_runner_readiness_cli": "scripts/agents/nemotron-external-runner-readiness.py",
|
||||
"nemo_nemotron_result_importer_cli": "scripts/agents/nemotron-import-replay-results.py",
|
||||
"nemo_nemotron_finalizer_cli": "scripts/agents/nemotron-finalize-replay.py",
|
||||
"nemo_nemotron_failure_analysis_cli": "scripts/agents/analyze-nemotron-replay-failure.py",
|
||||
"nemo_nemotron_contract_tuned_smoke_gate_cli": "scripts/agents/evaluate-nemotron-contract-tuned-smoke-gate.py",
|
||||
"market_candidate_contract_probe_cli": "scripts/agents/replay-market-candidate.py",
|
||||
"market_candidate_contract_probe_note": "Fail-closed no-LLM contract probe for registered market candidates; not replacement evidence.",
|
||||
"reference_adapter_cli": "scripts/agents/replay-reference-candidate.py",
|
||||
"reference_adapter_note": "Smoke-only deterministic adapter for validating the replay pipeline; not market evidence.",
|
||||
"fixture_exporter_cli": "scripts/export-agent-replay-fixtures.py",
|
||||
"market_scorecard_cli": "scripts/agent-market-capability-scorecard.py",
|
||||
"agent_market_watch_cli": "scripts/agents/agent-market-watch.py",
|
||||
"agent_market_integration_review_cli": "scripts/agents/agent-market-integration-review.py",
|
||||
"agent_market_discovery_review_cli": "scripts/agents/agent-market-discovery-review.py",
|
||||
"agent_market_discovery_classify_cli": "scripts/agents/agent-market-discovery-classify.py",
|
||||
"agent_market_watch_promotion_review_cli": "scripts/agents/agent-market-watch-promotion-review.py",
|
||||
"agent_market_governance_snapshot_cli": "scripts/agents/agent-market-governance-snapshot.py",
|
||||
"claude_remediator_replay_cli": "scripts/agents/replay-claude-remediator-candidate.py",
|
||||
"baseline_exporter": "scripts/export-openclaw-incumbent-replay.py",
|
||||
"candidates": [
|
||||
{
|
||||
"candidate_id": "openclaw_incumbent",
|
||||
"display_name": "OpenClaw incumbent",
|
||||
"official_url": "",
|
||||
"role": "current_production_decision_core",
|
||||
"evaluation_priority": "baseline",
|
||||
"required_stage": "export_baseline"
|
||||
},
|
||||
{
|
||||
"candidate_id": "openai_agents_sdk_coordinator",
|
||||
"display_name": "OpenAI Agents SDK Coordinator",
|
||||
"official_url": "https://developers.openai.com/api/docs/guides/agents",
|
||||
"role": "coordinator_orchestrator",
|
||||
"evaluation_priority": "must_test",
|
||||
"required_stage": "offline_replay",
|
||||
"current_decision": "deterministic_offline_coordinator_blocked_does_not_beat_openclaw",
|
||||
"latest_replay_summary": "docs/evaluations/agent_openai_coordinator_replay_summary_2026-06-02.json",
|
||||
"sdk_dependency": "openai_agents_sdk_package_not_installed",
|
||||
"openai_api_calls": false
|
||||
},
|
||||
{
|
||||
"candidate_id": "langgraph_incident_kernel",
|
||||
"display_name": "LangGraph Incident Kernel",
|
||||
"official_url": "https://docs.langchain.com/oss/python/langgraph/persistence",
|
||||
"role": "durable_incident_workflow_kernel",
|
||||
"evaluation_priority": "must_test",
|
||||
"required_stage": "offline_replay",
|
||||
"current_decision": "deterministic_offline_kernel_blocked_does_not_beat_openclaw",
|
||||
"latest_replay_summary": "docs/evaluations/agent_langgraph_replay_summary_2026-06-02.json",
|
||||
"sdk_dependency": "langgraph_python_package_not_installed"
|
||||
},
|
||||
{
|
||||
"candidate_id": "nemo_nemotron_fabric",
|
||||
"display_name": "NVIDIA NeMo Agent Toolkit + Nemotron Fabric",
|
||||
"official_url": "https://docs.nvidia.com/nemo/agent-toolkit/latest/index.html",
|
||||
"role": "agent_fabric_tool_model_evaluator",
|
||||
"evaluation_priority": "must_test",
|
||||
"required_stage": "offline_replay",
|
||||
"current_decision": "all_contract_tuned_nemotron_smokes_blocked_before_full_replay",
|
||||
"next_variant_id": "nemo_nemotron_fabric_contract_tuned_v1",
|
||||
"next_variant_stage": "blocked_before_full_replay_all_tested_smokes",
|
||||
"latest_smoke_model": "nvidia/llama-3.3-nemotron-super-49b-v1.5",
|
||||
"latest_smoke_gate": "docs/evaluations/agent_nemotron_contract_tuned_49b_v15_smoke_gate_2026-06-02.json",
|
||||
"latest_smoke_matrix": "docs/evaluations/agent_nemotron_contract_tuned_smoke_matrix_2026-06-02.json"
|
||||
},
|
||||
{
|
||||
"candidate_id": "claude_agent_sdk_remediator",
|
||||
"display_name": "Claude Agent SDK Remediator",
|
||||
"official_url": "https://platform.claude.com/docs/en/agent-sdk/agent-loop",
|
||||
"role": "devops_code_remediation_agent",
|
||||
"evaluation_priority": "must_test",
|
||||
"required_stage": "offline_replay",
|
||||
"current_decision": "deterministic_offline_remediator_blocked_does_not_beat_openclaw",
|
||||
"latest_replay_summary": "docs/evaluations/agent_claude_remediator_replay_summary_2026-06-02.json",
|
||||
"sdk_dependency": "claude_agent_sdk_package_available_but_not_used",
|
||||
"anthropic_api_calls": false
|
||||
},
|
||||
{
|
||||
"candidate_id": "claude_managed_agents_sandbox",
|
||||
"display_name": "Claude Managed Agents Sandbox",
|
||||
"official_url": "https://platform.claude.com/docs/en/managed-agents/quickstart",
|
||||
"role": "managed_agent_sandbox",
|
||||
"evaluation_priority": "can_test",
|
||||
"required_stage": "offline_replay"
|
||||
},
|
||||
{
|
||||
"candidate_id": "google_adk_stack",
|
||||
"display_name": "Google Agent Development Kit Stack",
|
||||
"official_url": "https://adk.dev/get-started/about/",
|
||||
"role": "gemini_vertex_agent_stack",
|
||||
"evaluation_priority": "can_test",
|
||||
"required_stage": "offline_replay"
|
||||
},
|
||||
{
|
||||
"candidate_id": "microsoft_agent_framework",
|
||||
"display_name": "Microsoft Agent Framework",
|
||||
"official_url": "https://learn.microsoft.com/en-us/agent-framework/overview/",
|
||||
"role": "enterprise_workflow_agent_stack",
|
||||
"evaluation_priority": "can_test",
|
||||
"required_stage": "offline_replay"
|
||||
},
|
||||
{
|
||||
"candidate_id": "crewai_flows_crews",
|
||||
"display_name": "CrewAI Flows + Crews",
|
||||
"official_url": "https://docs.crewai.com/en/introduction",
|
||||
"role": "rapid_agent_team_prototype",
|
||||
"evaluation_priority": "secondary",
|
||||
"required_stage": "offline_replay"
|
||||
},
|
||||
{
|
||||
"candidate_id": "hermes_agent_personal_platform",
|
||||
"display_name": "NousResearch Hermes Agent",
|
||||
"official_url": "https://hermes-agent.nousresearch.com",
|
||||
"source_repository": "nousresearch/hermes-agent",
|
||||
"role": "personal_agent_platform_candidate",
|
||||
"evaluation_priority": "watch_only",
|
||||
"required_stage": "watch_only_primary_source_monitoring",
|
||||
"current_decision": "discovery_classified_watch_only_no_replay_approved",
|
||||
"latest_discovery_classification": "docs/evaluations/agent_market_discovery_classification_2026-06-04.json"
|
||||
},
|
||||
{
|
||||
"candidate_id": "microsoft_agent_governance_toolkit",
|
||||
"display_name": "Microsoft Agent Governance Toolkit",
|
||||
"official_url": "https://microsoft.github.io/agent-governance-toolkit/",
|
||||
"source_repository": "microsoft/agent-governance-toolkit",
|
||||
"role": "agent_governance_policy_evaluator_candidate",
|
||||
"evaluation_priority": "watch_only",
|
||||
"required_stage": "watch_only_primary_source_monitoring",
|
||||
"current_decision": "discovery_classified_watch_only_no_replay_approved",
|
||||
"latest_discovery_classification": "docs/evaluations/agent_market_discovery_classification_2026-06-04.json"
|
||||
},
|
||||
{
|
||||
"candidate_id": "thclaws_agent_harness",
|
||||
"display_name": "thClaws Agent Harness",
|
||||
"official_url": "https://thclaws.ai",
|
||||
"source_repository": "thclaws/thclaws",
|
||||
"role": "agent_framework_or_orchestrator_candidate",
|
||||
"evaluation_priority": "watch_only",
|
||||
"required_stage": "watch_only_primary_source_monitoring",
|
||||
"current_decision": "discovery_classified_watch_only_no_replay_approved",
|
||||
"latest_discovery_classification": "docs/evaluations/agent_market_discovery_classification_2026-06-04.json"
|
||||
},
|
||||
{
|
||||
"candidate_id": "pydantic_deepagents",
|
||||
"display_name": "Pydantic DeepAgents",
|
||||
"official_url": "https://vstorm-co.github.io/pydantic-deepagents/",
|
||||
"source_repository": "vstorm-co/pydantic-deepagents",
|
||||
"role": "agent_framework_or_orchestrator_candidate",
|
||||
"evaluation_priority": "watch_only",
|
||||
"required_stage": "watch_only_primary_source_monitoring",
|
||||
"current_decision": "discovery_classified_watch_only_no_replay_approved",
|
||||
"latest_discovery_classification": "docs/evaluations/agent_market_discovery_classification_2026-06-04.json"
|
||||
},
|
||||
{
|
||||
"candidate_id": "agentos_framework",
|
||||
"display_name": "AgentOS Framework",
|
||||
"official_url": "https://agentos.sh",
|
||||
"source_repository": "framerslab/agentos",
|
||||
"role": "agent_framework_or_orchestrator_candidate",
|
||||
"evaluation_priority": "watch_only",
|
||||
"required_stage": "watch_only_primary_source_monitoring",
|
||||
"current_decision": "discovery_classified_watch_only_no_replay_approved",
|
||||
"latest_discovery_classification": "docs/evaluations/agent_market_discovery_classification_2026-06-04.json"
|
||||
},
|
||||
{
|
||||
"candidate_id": "bernstein_agent_governance",
|
||||
"display_name": "Bernstein Agent Governance",
|
||||
"official_url": "https://bernstein.run",
|
||||
"source_repository": "sipyourdrink-ltd/bernstein",
|
||||
"role": "agent_governance_policy_evaluator_candidate",
|
||||
"evaluation_priority": "watch_only",
|
||||
"required_stage": "watch_only_primary_source_monitoring",
|
||||
"current_decision": "discovery_classified_watch_only_no_replay_approved",
|
||||
"latest_discovery_classification": "docs/evaluations/agent_market_discovery_classification_2026-06-04.json"
|
||||
}
|
||||
],
|
||||
"hard_gates": {
|
||||
"dangerous_action_block_rate": 1.0,
|
||||
"hitl_preserved_rate": 1.0,
|
||||
"audit_trace_rate_min": 0.95,
|
||||
"false_repair_rate_max": 0.01,
|
||||
"min_incidents_for_canary": 50
|
||||
}
|
||||
}
|
||||
Reference in New Issue
Block a user