docs(ops): record 110 fwupd cleanup [skip ci]

This commit is contained in:
Your Name
2026-06-14 08:32:03 +08:00
parent e0a6d33966
commit 2b22c9d606
3 changed files with 64 additions and 17 deletions

View File

@@ -11,13 +11,13 @@
| Area | Status | Completion | Evidence |
|------|--------|------------|----------|
| Overall recovery readiness | SERVICE_AVAILABLE_KM_VECTORIZE_FAILED_DR_ESCROW_BLOCKED | 96% | 2026-06-14 03:11 cold-start scorecard is `PASS=81 WARN=2 BLOCKED=0`; 120/121 K3s are both `Ready control-plane`, backup core blockers remain `0`, public route/API smoke remains green, deploy marker `ec03f0b7` put API/Web/Worker/CronJob image `8ddb80d6` live, and API/Web remain live-verified split across 120 / 121. The 03:00 official `km-vectorize-29689620` Job failed with `BackoffLimitExceeded`, and 110 has `fwupd` failed units, so full cold-start cannot be declared green. Evidence retention and tenant-context fix are live; DR remains blocked by five missing credential escrow evidence markers. |
| Overall recovery readiness | SERVICE_AVAILABLE_KM_VECTORIZE_FAILED_DR_ESCROW_BLOCKED | 97% | 2026-06-14 08:24 cold-start scorecard is `PASS=82 WARN=1 BLOCKED=0`; 120/121 K3s are both `Ready control-plane`, backup core blockers remain `0`, public route/API smoke remains green, deploy marker `ec03f0b7` put API/Web/Worker/CronJob image `8ddb80d6` live, and API/Web remain live-verified split across 120 / 121. The 110 `fwupd` failed-unit warning is cleared by intentionally disabling inactive `fwupd-refresh.timer` and `systemctl --failed` is clean. Full cold-start still cannot be declared green because the official `km-vectorize-29689620` Job failed with `BackoffLimitExceeded`; DR remains blocked by five missing credential escrow evidence markers. |
| P0 host / K3s recovery | DONE | 100% | 120 booted after console fsck at `2026-06-12 15:13`; host is reachable, root is mounted `rw`, failed units `0`, `mon` and `mon1` are both `Ready control-plane`, and cold-start P0/P1 checks are green. |
| P1 backup / alert / escrow | BLOCKED_DR_ESCROW | 92% | 2026-06-14 03:11 `backup-status` shows 110 `13/13 fresh failed=0`, 188 `2/2 fresh failed=0`, `core_blockers=0`, `escrow_missing=5`, last aggregate `2026-06-14 02:40:22`. Owner request package is ready; actual marker write remains blocked on real non-secret evidence IDs. |
| P2 service / data truth | VERIFIED_WORKLOAD_BALANCED_WITH_CRON_WARN | 98% | 2026-06-14 03:11 cold-start is degraded by warnings only; public route/API smoke is green, VIP API/Web are reachable, momo current-month parity remains covered by the scorecard, schedules/services are mostly green. API/Web both keep 120 / 121 split placement after latest ArgoCD revision `ec03f0b7`, with live API/Web/Worker image `8ddb80d6`; the exception is failed `km-vectorize-29689620`. |
| P3 docs / automation contracts | DONE_WITH_VALIDATION_GAP | 100% | Workplan, SOP v1.10, BACKUP-STATUS, LOGBOOK, 120 console/fsck recovery, Gitea backup stale-dump hardening, reboot ledger/version-comparison SOP, escrow evidence audit, 188 nginx Ansible baseline, 110 cold-start detector script, startup judgment layers, GO/NO-GO tree, host recovery cards, T+0/T+60 timeline checks, host role / load-balancing assessment, CD `known_hosts` guardrail, and `km-vectorize` remediation tracking are updated; Ansible syntax check is unavailable on this workstation. |
| P2 service / data truth | VERIFIED_WORKLOAD_BALANCED_WITH_KM_WARN | 99% | 2026-06-14 08:24 cold-start is degraded by one warning only; public route/API smoke is green, VIP API/Web are reachable, momo current-month parity remains covered by the scorecard, schedules/services are mostly green, and 110 failed units are now `0`. API/Web both keep 120 / 121 split placement after latest ArgoCD revision `ec03f0b7`, with live API/Web/Worker image `8ddb80d6`; the remaining exception is failed `km-vectorize-29689620`. |
| P3 docs / automation contracts | DONE_WITH_VALIDATION_GAP | 100% | Workplan, SOP v1.11, BACKUP-STATUS, LOGBOOK, 120 console/fsck recovery, Gitea backup stale-dump hardening, reboot ledger/version-comparison SOP, escrow evidence audit, 188 nginx Ansible baseline, 110 cold-start detector script, startup judgment layers, GO/NO-GO tree, host recovery cards, T+0/T+60 timeline checks, host role / load-balancing assessment, CD `known_hosts` guardrail, `fwupd-refresh.timer` rollback note, and `km-vectorize` remediation tracking are updated; Ansible syntax check is unavailable on this workstation. |
Full cold-start may be declared green only for the latest verified evidence set. As of 2026-06-14 03:11, the latest evidence set is degraded, not green. Do not declare DR scorecard complete while credential escrow evidence remains blocked.
Full cold-start may be declared green only for the latest verified evidence set. As of 2026-06-14 08:24, the latest evidence set is degraded by `km-vectorize` only, not green. Do not declare DR scorecard complete while credential escrow evidence remains blocked.
2026-06-13 01:26 refresh: full cold-start is again green for the current evidence set. AWOOOI API/Web workload balancing survived the next normal CD deploy: Gitea main `e4a349bc`, ArgoCD revision `e4a349bc`, images from `414413a5`, API/Web split across `mon` / `mon1`, and global `known_hosts` retained 120 / 188 after CD fix `80e6ec1a`. Do not declare DR complete while credential escrow is missing. `km-vectorize` remediation is `90%`: schedule/label fix is live, and the remaining gate is the next official 03:00 CronJob success readback.
@@ -64,6 +64,7 @@ Full cold-start may be declared green only for the latest verified evidence set.
| 2026-06-13 final goal audit refresh | SERVICE_GREEN_REMAINING_GATES_EXPLICIT | Clean worktree rebased onto `a520c32d` and reran source guards successfully; live ArgoCD tracks revision `a520c32d` with API/Web/Worker image `e897c8bf`, health `Degraded` only by `km-vectorize`; `km-vectorize` schedule remains `0 3 * * *`, `timeZone=Asia/Taipei`, `failedJobsHistoryLimit=3`, and no failed Job is currently retained. Public `/zh-TW/governance`, `/en/governance`, and `/api/v1/health` are green; backup core blockers remain `0`, `escrow_missing=5`; 14:16 cold-start is `PASS=83 WARN=0 BLOCKED=0`. Remaining gates: five credential escrow markers and next official 03:00 `km-vectorize` success readback. |
| 2026-06-14 `km-vectorize` official run follow-up | DEGRADED_EVIDENCE_RETENTION_LIVE | 03:00 official `km-vectorize-29689620` ran from CronJob and failed with `BackoffLimitExceeded`; ArgoCD later auto-synced revision `8868c025` and remains `Synced / Degraded`. Job is retained, but failed Pod `km-vectorize-29689620-nwpqz` was deleted before logs could be read, so root cause remains unproven for this run. Live CronJob is now `restartPolicy: Never` plus `terminationMessagePolicy: FallbackToLogsOnError`, so the next official failure should retain Pod/log evidence. Backup core remains green, `escrow_missing=5`, and 03:11 cold-start is `PASS=81 WARN=2 BLOCKED=0`. |
| 2026-06-14 `km-vectorize` tenant context follow-up | ROOT_CAUSE_CANDIDATE_LIVE | Source audit shows `cron_km_vectorize.py` calls `/api/v1/knowledge/embed-all` without project context, while API middleware and `get_db_context()` require `X-Project-ID` / tenant context for fail-closed RLS. API logs show matching `db_context_missing` / `Missing tenant context` patterns. Deploy marker `ec03f0b7` put image `8ddb80d6` live; CronJob now has `KM_PROJECT_ID=awoooi`, script sends `X-Project-ID`, targeted pytest `7 passed`, and no manual Job was created. Completion still waits for the next official 03:00 success or retained failed Pod/log. |
| 2026-06-14 110 failed-unit cleanup | SERVICE_AVAILABLE_KM_VECTORIZE_FAILED_DR_ESCROW_BLOCKED | `fwupd-refresh.timer` is intentionally `disabled / inactive` after non-runtime firmware metadata refresh failed units were classified; rollback is `sudo systemctl enable --now fwupd-refresh.timer`. `systemctl --failed` now returns `0 loaded units listed`; 08:24 cold-start improved to `PASS=82 WARN=1 BLOCKED=0`. Remaining warning is only K8s failed Job `km-vectorize-29689620`; backup core remains green and `escrow_missing=5`. |
---
@@ -165,7 +166,7 @@ Next: <single next action>
| P3-005 | DONE | 100 | Update cold-start SOP | SOP now includes start, shutdown, reboot, record, comparison, and 120 blocker handling. | Increment SOP version after each process change. | SOP has controlled power-operation sections and ledger template. |
| P3-006 | DONE | 100 | Update backup status | Backup status now reflects current cron, rclone latest-only, failure-only alert posture, and escrow blocker. | Refresh after 120 backup rerun. | Backup status no longer claims noisy success Telegram notifications. |
| P3-007 | DONE | 100 | Harden Gitea backup stale dump handling | 2026-06-05 manual Gitea backup failed because the container retained `/tmp/gitea-dump.zip` from the 02:00 failure. `scripts/backup/backup-gitea.sh` now renames stale container dump files to timestamped evidence before running a new dump, and the live 110 script is updated. | Watch the next 02:00 Gitea backup. | `bash -n` passes locally and on 110; manual Gitea backup completed after stale evidence rename. |
| P3-008 | DONE | 100 | Continuously optimize host reboot SOP | SOP v1.9 adds startup judgment layers, GO/NO-GO decision tree, freeze execution checklist, host boot detection, 110/188/120/121 recovery cards, 2026-06-12 post-reboot anchor, 2026-06-13 post-CD trust/workload anchor, T+0/T+60 verification timeline, AA/AS 判定, workload 分散判定, CD SSH trust guardrail, CronJob failure evidence retention rule, and allowed declaration wording. | Use v1.9 for the next reboot record, then compare actual timing and blockers against §14.8 / §14.9 / §14.10. | SOP distinguishes `HOST_BOOTED`, `HOST_READY`, `SERVICE_READY`, `FULL_STACK_GREEN`, `K3S_CONTROL_PLANE_AA`, and `WORKLOAD_BALANCED`, and blocks false green while escrow or governed CronJob debt remain red. |
| P3-008 | DONE | 100 | Continuously optimize host reboot SOP | SOP v1.11 adds startup judgment layers, GO/NO-GO decision tree, freeze execution checklist, host boot detection, 110/188/120/121 recovery cards, 2026-06-12 post-reboot anchor, 2026-06-13 post-CD trust/workload anchor, 2026-06-14 110 failed-unit cleanup anchor, T+0/T+60 verification timeline, AA/AS 判定, workload 分散判定, CD SSH trust guardrail, CronJob failure evidence retention rule, `fwupd-refresh.timer` rollback note, and allowed declaration wording. | Use v1.11 for the next reboot record, then compare actual timing and blockers against §14.8 / §14.9 / §14.10 / §14.11. | SOP distinguishes `HOST_BOOTED`, `HOST_READY`, `SERVICE_READY`, `FULL_STACK_GREEN`, `K3S_CONTROL_PLANE_AA`, and `WORKLOAD_BALANCED`, and blocks false green while escrow or governed CronJob debt remain red. |
| P3-009 | DONE | 100 | Assess 120/121 AA/AS role and host load balancing | 2026-06-12 15:19 live check confirms 120 and 121 are both `Ready control-plane`, `k3s active`, `k3s-agent inactive`, with no taints; however most AWOOOI / ArgoCD / Velero workload remains on 121 after 120 fsck recovery. New assessment defines control-plane AA vs workload AA, migration candidates from 110/188, and stateful migration blockers. | After P0 backup/offsite/cold-start green, implement topology spread for AWOOOI API/Web before moving additional services. | `docs/runbooks/HOST-ROLE-LOAD-BALANCING-ASSESSMENT.md` exists; SOP v1.6 links AA/AS and load-balancing checks; migration implementation remains explicitly `0%`. |
| P3-010 | DONE | 100 | Update workload balancing docs with 2026-06-13 live truth | Host role assessment, workplan, SOP, backup status, and LOGBOOK are refreshed with current cold-start, backup, 188 certbot degraded, ArgoCD `km-vectorize` degraded, Gitea main `acaae999`, ArgoCD sync, and final pod placement evidence. | Keep updating this file after the next reboot or deploy. | Docs separate service-green status from DR escrow, workload rollout, and non-service governance debt. |
| P3-011 | DONE | 100 | Record `km-vectorize` remediation status | LOGBOOK, this workplan, and SOP now state the schedule/label fix, ArgoCD sync evidence, the invalid manual Job boundary, and the 90% waiting-for-next-schedule gate. | After next 03:00 run, update this row and the top verdict with `lastSuccessfulTime` / ArgoCD health evidence. | No document claims ArgoCD green before official CronJob success evidence exists. |
@@ -203,6 +204,16 @@ Do not run `truncate`, whole DB restore, force-push, DROP, or online root filesy
## 9. Progress Updates
```text
2026-06-14 08:24 Asia/Taipei
Phase: P0/P1/P2/P3
Before: Overall 96%, P1 92%, P2 98%, P3 100%
After: Overall 97%, P1 92%, P2 99%, P3 100%
Evidence: 110 fwupd-refresh.timer disabled/inactive with rollback command recorded; systemctl --failed returned 0 loaded units listed; backup-status 110 13/13 fresh failed=0 and 188 2/2 fresh failed=0 with core_blockers=0 and escrow_missing=5; cold-start PASS=82 WARN=1 BLOCKED=0; ArgoCD/CronJob still waiting for official km-vectorize lastSuccessfulTime after deploy marker ec03f0b7 / image 8ddb80d6.
Blocked: yes for full cold-start green, because km-vectorize-29689620 remains failed until the next official 03:00 success or retained failed Pod/log evidence; yes for DR complete, because credential escrow evidence markers still missing 5.
Next: after the next 03:00 Asia/Taipei official km-vectorize schedule, read-only verify lastSuccessfulTime, latest Job/Pod/log, and ArgoCD health; do not manual-run, delete, patch, or fake evidence.
```
```text
2026-06-13 01:29 Asia/Taipei
Phase: P0/P1/P2/P3