Files
CASAN/casan-next-plans/CASAN_HARDENING_STATUS.md
T
thanhnvandClaude Opus 4.8 42115e3361 docs+demo: deep-gap closers — 211/0 re-score, HD1-HD4 scenes, status/README/scoring
Full authoritative run 2026-07-06: all 12 suites 211 PASS / 0 FAIL (KMS + container
isolation live via Vault dev + Docker; security-gate 11/0).
- run-hardening.sh: new "Vá đường lọt sâu" section (HD1 incident/kill-switch,
  HD2 multilingual VI/JA, HD3 true container isolation, HD4 split+classifier),
  closer updated to 211 checks.
- CASAN_HARDENING_STATUS.md: Phase 6 deep-gap closers table; test inventory
  175→211 (12 suites); C7/multilingual moved out of planned; C6 planned→partial
  (real isolation done); honest claim → H4 83, H2 82, H5/H6 stay 80 (infra-bound).
- scoring-report-02-after-competition.md: current state — 211/0, H4 80→83,
  H2 80→82, avg 80.9→81.6, lowest harness still 80 (H5/H6), 3-milestone table.
- README claim boundary: deep-gap closers listed; totals 175→211; H4/H2 bumps.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 09:51:39 +09:00

117 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# CASAN Production Hardening — Implementation Status
> Live record of what has been **implemented and tested** from Plan-07
> (Production Hardening) and Plan-09 (Evidence Pack), versus what is still
> **planned**. Honest by design: a control is only listed as done when it has an
> executable test that would fail if the control were removed.
>
> Labels: **[implemented+tested]** · **[scaffold+tested]** (works, but a stronger
> production form is documented as TODO) · **[planned]** (not built).
## 1. What is implemented now
### Phase 1 — Plan-07 Track A (low-risk hardening) — [implemented+tested]
| ID | Control | Where | Test |
|---|---|---|---|
| A1 | H4 semantic **strict fail-closed** (`CASAN_SECURITY_STRICT=1`): model unavailable ⇒ BLOCK, never silent SKIP; non-strict logs `SEMANTIC_SKIPPED` loudly | `security-check.sh` | phase1 A1 |
| A2 | H4 **unicode/encoding normalization**: NFKC + zero-width strip + homoglyph fold + base64/hex decode & rescan | `security-check.sh`, `unicode-normalize.py`, `decode-suspicious.py` | phase1 A2 |
| A3 | H4 **tool-output injection scan** before output re-enters model context (mode off/warn/block; strict⇒block) | `tool-output-scan.sh`, `casan-harness.sh` | phase1 A3 |
| A4 | H5 **telemetry integrity**: provider-usage + cost metrics bound to a signed manifest head (tamper ⇒ mismatch; re-forge ⇒ signature invalid) | `telemetry-integrity.sh`, `sign-audit-head.sh` | phase1 A4 |
| A5 | H6 **cost controls**: absolute per-call cap + cumulative budget + cold-start protection, keeping median×mult spike | `cost-spike-detect.sh` | phase1 A5 |
| A6 | **Benign corpus (EN/VI/JA, 95 samples) + FP budget gate**: FP ≤ 3%, adversarial block ≥ 95%, CRITICAL = 100% | `benign-fp-report.sh`, `benign-corpus/`, `redteam-vectors.jsonl` | phase1 A6 |
### Phase 2 — Plan-07 Track C-MVP (production minimum bar) — mixed
| ID | Control | Status | Where | Test |
|---|---|---|---|---|
| C1 | **Tool authorization / action gating** (ALLOW/WARN/REQUIRE_APPROVAL/BLOCK): sensitive-file writes, destructive/remote-exec commands, dep-install & network egress | [implemented+tested] | `action-gate.sh` | phase2 C1 |
| C2 | **Supply-chain gate**: manifest diff, denylist + typosquat + dangerous lifecycle-script block, new-dep approval, dep-diff report | [implemented+tested] | `supply-chain-gate.sh`, `supply-chain-scan.py` | phase2 C2 |
| C3 | **Data-exfil guard**: secret-to-cloud & artifact-leaks-env ⇒ BLOCK; PII-in-audit ⇒ mask | [implemented+tested] | `data-exfil-guard.sh` | phase2 C3 |
| C6 | **Runtime sandbox**: static policy block (read ~/.ssh, net egress, fork bomb, write-outside-workspace, huge-file) + ulimit/timeout backstops | [scaffold+tested] | `sandbox-run.sh` | phase2 C6 |
### Phase 3 — Plan-09 Evidence Pack (MVP) — [implemented+tested]
| Control | Where | Test |
|---|---|---|
| `casan pack` / `casan verify-pack` (mapped to `evidence-pack.sh`): standard 12-file pack, hash manifest, signed head, tamper-evident verify, certified-run gate (no false certification, no silent skip) | `evidence-pack.sh`, `evidence-pack-build.py`, `evidence-pack-verify.py` | phase3-evidence-pack |
### Phase 4 — H5 governance hardening (raises the lowest harness) — mixed
| ID | Control | Status | Where | Test |
|---|---|---|---|---|
| C4 | **Approval-identity**: high-risk approval trusted only when a REGISTERED reviewer cryptographically SIGNS the request and their role is authorized — env-var approver no longer enough (SoD still enforced) | [implemented+tested] | `approval-sign.sh`, `approval-verify.sh`, `reviewers.registry`, `governance-check.sh` (`CASAN_APPROVAL_STRICT=1`) | phase-h5-approval (8) |
| B3 | **KMS key management**: sign audit/telemetry head via Vault Transit (key `exportable:false` → never leaves KMS) + key **rotation**; validated live | [implemented+tested] (live when Vault present; skip-aware otherwise) | `vault-kms.sh` (`rotate`, `assert-nonexportable`), `sign-audit-head.sh` | phase-h5-infra (KMS) |
| C5 | **External WORM audit**: ship audit head to a hash-linked append-only ledger; detect local rollback (`AUDIT_GAP_DETECTED`) and ledger tamper (`AUDIT_LEDGER_TAMPERED`) | [implemented+tested] (local ledger MVP) | `worm-ledger.py`, `audit-ship.sh`, `verify-audit-gap.sh` | phase-h5-infra (WORM) |
### Phase 5 — H6 AgentOps hardening (raises the lowest harness) — mixed
| ID | Control | Status | Where | Test |
|---|---|---|---|---|
| D1 | **Live alert dispatch**: alerts POST to a real HTTP webhook (severity routing, dedup window, retry); undelivered alerts fail-loud (`ALERT_DELIVERY_FAILED`, strict) and queue to a **dead-letter** file with redelivery (`--flush-deadletter`); wired into `agent-metrics.sh` (failing step ⇒ live page) | [implemented+tested] (live local webhook; production points at Slack/PagerDuty + on-call) | `alert-dispatch.sh`, `agent-metrics.sh` | phase-h6-agentops (①) |
| D2 | **Provider-telemetry API + reconciliation**: pull usage from a provider usage HTTP API (all-or-nothing schema gate, `PROVIDER_API_UNREACHABLE`/`PROVIDER_USAGE_INVALID` fail-loud) + reconcile local metrics vs provider ground truth — under-reporting/hidden runs ⇒ `TELEMETRY_DISCREPANCY` | [implemented+tested] (live local API endpoint; production points at OpenAI/Anthropic usage APIs) | `provider-usage-fetch.sh`, `telemetry-reconcile.sh` | phase-h6-agentops (②) |
| D3 | **Hosted dashboard**: dashboard served over HTTP with a stale-aware `/healthz` probe (fresh ⇒ 200 ok; telemetry silent-death ⇒ 503 stale — page-able by any uptime monitor) | [implemented+tested] (local HTTP daemon; production host = nginx/container, same routes) | `dashboard-serve.sh`, `dashboard-server.py` | phase-h6-agentops (③) |
| D4 | **Sliding-window circuit breaker (V15)**: failure **rate** over the last N calls trips `CIRCUIT_OPEN_WINDOW` — interleaving successes no longer evades the consecutive-failure breaker | [implemented+tested] | `circuit-breaker-check.sh` | phase-h6-agentops (④) |
### Phase 6 — Deep-gap closers (Track B + C6/C7, post-competition) — mixed
| ID | Control | Status | Where | Test |
|---|---|---|---|---|
| C7 | **Incident response + kill-switch (V23)**: `incident.sh raise` grades severity (LOW/MED/HIGH/CRIT via `incident-severity.map`), records a routed entry (owner), and for HIGH/CRIT auto-engages the scoped `kill-switch.sh` (project/model/provider/global) + fires an alert; `casan-harness.sh` refuses to run under an engaged switch (opt-in) | [implemented+tested] | `incident.sh`, `kill-switch.sh`, `incident-runbook.md`, `incident-severity.map` | phase-c7-incident (15) |
| B1 | **Multilingual VI/JA injection (V2)**: VI/JA block-patterns (matched on raw UTF-8, anchored on the injection object) catch injections English regex missed, with 0 false positives on the benign VI/JA corpus | [implemented+tested] | `prompt-filter.yaml` (PI-VI-*, PI-JA-*) | phase-h4-multilingual (7) |
| C6 | **TRUE runtime isolation (V22)**: container sandbox (`--network=none --read-only --pids-limit --cap-drop=ALL`, workspace-only mount) — the kernel neutralises host-file reads / egress / out-of-workspace writes; upgrades the static scaffold | [implemented+tested] (live via Docker; skip-aware) | `sandbox-container.sh`, `sandbox-run.sh` (`CASAN_SANDBOX_MODE=container`) | phase-c6-sandbox (6) |
| B2 | **Split + classifier injection (V5,V6)**: `context-assemble-scan.sh` scans the concatenated context so a payload split across benign pieces is caught on assembly; verdict-steering patterns (PI-CLS-*) block content that tries to hijack the evaluator | [implemented+tested] | `context-assemble-scan.sh`, `prompt-filter.yaml` (PI-CLS-*) | phase-h4-split-inject (8) |
## 2. Test inventory (all suites)
| Suite | Checks | Purpose |
|---|---|---|
| `run-casan4-harness-tests.sh` | 35 | Baseline happy-path + Level-5 evidence (unchanged) |
| `adversarial-harness-tests.sh` | 44 | Original adversarial battery (unchanged) |
| `phase1-track-a-tests.sh` | 25 | Track A hardening |
| `phase2-track-c-tests.sh` | 29 | Track C-MVP |
| `phase3-evidence-pack-tests.sh` | 7 | Evidence Pack MVP |
| `phase-h5-approval-tests.sh` | 8 | Approval-identity (C4) |
| `phase-h5-infra-tests.sh` | 7 | KMS (B3, live/skip-aware) + WORM (C5) |
| `phase-h6-agentops-tests.sh` | 20 | live alerting (D1) + provider-API/reconcile (D2) + hosted dashboard (D3) + window breaker (D4); against live local HTTP endpoints |
| `phase-c7-incident-tests.sh` | 15 | **New** — incident severity + scoped kill-switch (C7) + wrapper enforcement |
| `phase-h4-multilingual-tests.sh` | 7 | **New** — VI/JA injection block + benign VI/JA 0-FP (B1) |
| `phase-c6-sandbox-tests.sh` | 6 | **New** — TRUE container isolation (C6, live via Docker / skip-aware) |
| `phase-h4-split-inject-tests.sh` | 8 | **New** — split-injection assembly scan + classifier-inject (B2) |
| **Total** | **211** | Baseline 79 preserved; +132 new hardening checks. Last full run 2026-07-05 @ head of `feat/plan07-track-a-hardening`, 0 fail (KMS + container isolation validated live via Vault dev + Docker). |
Run order note: `run-casan4-harness-tests.sh` does `rm -rf .specify/logs`, so run it
**first** and never concurrently with the other suites.
## 3. What is NOT done (still planned — do not claim as production-ready)
| Area | Status | Plan ref |
|---|---|---|
| Model-digest pinning | [planned] — sliding-window circuit breaker (V15) is now done (Phase 5 D4) | Plan-07 B4 (V16) |
| Live alerting to a managed channel (Slack/PagerDuty + on-call rota) | [partial] — webhook dispatch + dedup + dead-letter done; managed channel & escalation are config away, incident workflow is C7 | Plan-07 C7 / Phase 5 D1 |
| Hosted telemetry dashboard | [partial] — HTTP-served dashboard + stale-aware `/healthz` done locally; deployed host (nginx/container, auth) planned | Phase 5 D3 |
| Provider billing-API telemetry | [partial] — API fetch + schema gate + local-vs-provider reconciliation done against a live local endpoint; real OpenAI/Anthropic usage-API calls (needs keys) planned | Phase 5 D2 |
| True runtime isolation | [partial] — real container isolation done + validated live via Docker (C6 phase-6); nsjail/rootless + a hardened base image for CI still planned | Plan-07 C6 (V22) |
| KMS key management (rotation, non-exportable) | [partial] — Vault Transit path implemented + validated live; not yet the default (local-key fallback), no HSM/short-lived IdP tokens | Plan-07 B3 |
| Reviewer approval workflow | [partial] — cryptographic **approval-identity** done (signed reviewer + role); live **IdP (OIDC/JWT)** + policy versioning/diff still planned | Plan-07 C4 (V20) |
| External append-only (WORM) audit | [partial] — hash-linked local ledger + rollback/tamper detection done; true WORM store (S3 Object Lock/QLDB) + trusted timestamp planned | Plan-07 C5 (V21) |
| Live CVE/OSV scanning wired in | [partial] — availability detected; local denylist authoritative offline | Plan-07 C2 follow-up |
## 4. Honest claim
Track A + Track C-MVP + Evidence Pack + H5/H6 hardening + the deep-gap closers
(C7 incident/kill-switch, VI/JA multilingual, true container isolation, split &
classifier injection) raise H4/H5/H6 from "PoC/demo (~3.0/5)" to **early
internal-production hardening**, with executable adversarial tests for every
control (**211 checks, 0 fail** — last full run 2026-07-05; KMS + container
isolation validated live via Vault dev + Docker). Fair maturity score
(`00_SUBMISSION_PACKAGE/evidence/scoring-run-report.md`): **H4 80→83** (multilingual
+ split/classifier closed), **H2 80→82** (real sandbox isolation), C7 incident
dimension closed; **H5 and H6 remain at 80** (their remaining gaps are infra), so the
**lowest harness stays 80** — CASAN **Level 4**, proven by attack. This is **not** full
production readiness: crossing the whole pipeline into "Strong (81+)" still needs the
H5/H6 infra items — live IdP (OIDC/JWT), a true WORM store (S3 Object Lock),
KMS-by-default + HSM, a deployed dashboard host + managed alert channel/on-call, and
real billing-API telemetry — the [partial]/[planned] rows above and in
`CASAN_PLAN_07_PRODUCTION_HARDENING.md`.