thanhnv and Claude Opus 4.8
2af67ef6a5
docs: re-score from a real full run (155/0) after H5 hardening
...
Ran all 7 suites sequentially on 2026-07-04 @ 00aabfa with Vault dev live so KMS
runs (not skips): run-casan4 35 · adversarial 44 · phase1-track-a 25 ·
phase2-track-c 29 · phase3-evidence 7 · phase-h5-approval 8 · phase-h5-infra 7
= 155 PASS / 0 FAIL; security-gate PASS=11 FAIL=0.
- scoring-run-report.md: fair re-score — H5 76→80 (approval-identity + KMS live
rotate/non-exportable + WORM), lowest harness now H6=79, avg ~80.7/100, Level 4.
Evidence lists the live KMS + WORM results.
- CASAN_HARDENING_STATUS.md: new Phase 4 (C4 approval-identity, B3 KMS, C5 WORM =
implemented+tested); test inventory 140→155 (7 suites); planned→partial for
KMS/approval/WORM with honest remaining gaps (live IdP, S3 WORM store, KMS default).
- Plan-07 §2: key-mgmt 2.5→4, policy-approval 2.5→4, external-audit 1.5→3.5;
header now H5 76→80, lowest harness H6.
- INDEX row, README claim boundary, video-guide Q&A: 155 checks, H5=80, lowest H6=79.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-04 23:51:50 +09:00
thanhnv and Claude Opus 4.8
d8583fdb2e
docs(plan-07): re-score §2 readiness table after Track A + C-MVP + Evidence Pack
...
Update the 0–5 production-readiness table with Baseline → Nay per dimension and
point to the fair 0–100 report. Trio H4/H5/H6 ~3.0 → ~3.7/5 (per-harness fair
score 76–80/100); Track C dims lifted (tool-authz 2→4, supply-chain 1→3.5,
data-exfil 2→4, sandbox 1→2.5 scaffold, +Evidence Pack 4). Governance/Ops
(C4/C5/C7) still 1–2.5 [planned]. Honest headline: ~80/100 avg, lowest H5=76,
CASAN Level 4 at threshold. Source: evidence/scoring-run-report.md.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-04 22:56:03 +09:00
thanhnv and Claude Opus 4.8
2ffbda3fad
feat(video+deck): Part 2 hardening demo script + slide deck update
...
- run-hardening.sh: narrated "Part 2" battery in the same visual language as
run-all.sh (card/attack/guard/cmd + on-screen exit code with ⛔ /✋ /✅ verdict
chips). Scenes: Track A (HA1-7: homoglyph/zero-width/base64, strict fail-closed,
telemetry tamper, cost slow-boil/spray, benign FP=0%), Track C-MVP (HC1-4:
action-gate, supply-chain, data-exfil, sandbox), and the Evidence Pack
money-shot (HE1-4: pack → verify VALID → tamper 1 byte → TAMPERED exit=1 →
certified-only-when-earned). Verified end-to-end (exit 0).
- CASAN_SLIDE_DECK.html: reflect implemented vs planned honestly. Updated H4/H5/H6
AFTER columns, evidence terminal (140 checks), threat-model reframed to
"identified AND Track A closed", readiness meters bumped to ~3.8-4.0, Track C-MVP
reframed to DELIVERED (29/29), roadmap marks 07+09 done. Added 2 slides:
"Track A delivered" and "Evidence Pack money-shot" (18 slides total).
- video guide: Part 1 (run-all) + Part 2 (run-hardening) with pre-flight note.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-03 23:51:11 +09:00
thanhnv and Claude Opus 4.8
4cc78f74eb
docs: honest hardening status + claim boundary (Plan-07 A/C-MVP, Plan-09)
...
- CASAN_HARDENING_STATUS.md: canonical implemented/scaffold/planned record with
per-control test mapping and test inventory (baseline 79 + 61 new = 140 checks).
- README claim boundary: separate implemented+tested controls from planned;
explicitly does NOT claim full production-readiness (Track B, C-Gov/Ops, true
sandbox isolation, IdP/WORM still planned).
- INDEX status table: Plan-07 Track A + C-MVP done, Plan-09 MVP done.
- video guide: core demo battery counts unchanged (hardening lives in separate
suites); added deep-dive commands + 3 Q&A rows + sandbox honesty note.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-03 23:07:31 +09:00
thanhnv and Claude Opus 4.8
fbcef967e5
chore(freeze): snapshot demo state before Plan-07 hardening work
...
Freeze current submission/demo baseline:
- casan-next-plans/: full task-level plan set (Plan 00 index + 02/04/06/07/08/09/12, QA, slide deck)
- optimize-docs/video-steps/: per-vector scene breakdown (commands/screen-text/script) + start-tmux
- run-all.sh / scorecard.sh / map-live.sh: REAL=1 live-battery wiring
- regenerated evidence + audit/telemetry logs from live REAL=1 run
- submission README + video recording guide updates
- dry-run pipeline logs for 001-okr-web-app
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com >
2026-07-03 22:03:44 +09:00