Files
CASAN/packages/casan-harness/security/incident-runbook.md
T
thanhnvandClaude Opus 4.8 36a4812ef3 refactor(structure): promote app to repo root + remove redundant workspace cruft
Standard production layout: the OKR app (was nested under AINative_OKR_CASAN5/) is now
the repository root. No more wrapper directory.

- Promote AINative_OKR_CASAN5/* -> repo root (backend/ frontend/ packages/ apps/
  .specify/ docs/ infra/ nginx/ scripts/ + configs). Merge tool dirs: .gitea (kept the
  active deploy ci.yml, added harness-ci.yml + runbooks), .claude (agents/commands +
  launch.json), .github moved up.
- Remove redundant: 00_SUBMISSION_PACKAGE, scattered root notes (FPT_CASAN_Full.md,
  tu-tuong-casan.md, casan-tu-sinh..., casan_harness_assessment.md, source-review...,
  README_CASAN5_REFINED.md), casan-next-plans/ and optimize-docs/ (competition/planning
  artifacts — roadmap + design history preserved in git log / commit messages).
- Update all references to the old layout:
  - .gitea/workflows/{ci,harness-ci}.yml, .github/workflows/{ci,deploy}.yml:
    working-directory .; drop AINative_OKR_CASAN5/ prefix; .specify/{tests,scripts}
    -> packages/casan-harness/... (.specify/logs state kept)
  - .claude/launch.json, .gitea/*-runbook.md: path prefixes
  - CLAUDE.md, README.md: docs/input -> apps/okr/domain/input
  - policy-bundle.yaml: 8 policy paths -> packages/casan-harness/...; manifest re-signed
- secrets-scan.sh: fixture excludes -> new package/domain paths.

Full gate from the new root: PASS=64 FAIL=0 SKIP=3.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-08 13:26:36 +09:00

1.8 KiB

CASAN Incident Runbook (C7 / V23)

When a gate raises an incident (incident.sh raise <event>), it is classified, recorded to logs/level5/incidents.jsonl, and for HIGH/CRIT the scoped kill-switch is engaged automatically + an alert is dispatched.

Severity → owner → response

Severity Owner (on-call) Auto-action Human step
CRIT security-oncall kill-switch engaged + alert Contain now; verify blast radius; do NOT clear until root cause known
HIGH ops-oncall kill-switch engaged + alert Assess; clear switch only after fix + reviewer sign-off
MED tech-lead recorded + alert Triage within SLA; batch-fix
LOW triage recorded Review in retro

Kill-switch operations

kill-switch.sh status                     # what is engaged
kill-switch.sh check <scope> <id>         # gates honor this (exit 2 = stop)
kill-switch.sh clear <scope> <id> <reason># turn off (production: reviewer-approved)

Scopes: project · model · provider · global (global stops everything).

Event → severity

See incident-severity.map. Examples: secret-to-cloud=CRIT, tool-write-sensitive=CRIT, dependency-postinstall=HIGH, audit-chain-broken=HIGH, cost-budget-exceeded=MED.

Postmortem template (fill after resolution)

  • Incident: <id / timestamp / event / severity>
  • Detection: which gate fired, what signal
  • Blast radius: scope, what was stopped by the kill-switch
  • Root cause:
  • Fix:
  • Prevent recurrence: new test/gate added (link the fail-able check)
  • Kill-switch cleared by: at

Production TODO

Managed alert channel (Slack/PagerDuty) + on-call rota + auto issue creation; kill-switch clear gated by reviewer approval (tie to approval-identity C4).