Commit Graph
21 Commits
Author SHA1 Message Date
thanhnv 4419cd9eae update doc and optimize 2026-07-06 17:47:12 +09:00
thanhnv a832c40bdb update doc/log 2026-07-06 10:13:43 +09:00
thanhnvandClaude Opus 4.8 42115e3361 docs+demo: deep-gap closers — 211/0 re-score, HD1-HD4 scenes, status/README/scoring
Full authoritative run 2026-07-06: all 12 suites 211 PASS / 0 FAIL (KMS + container
isolation live via Vault dev + Docker; security-gate 11/0).
- run-hardening.sh: new "Vá đường lọt sâu" section (HD1 incident/kill-switch,
  HD2 multilingual VI/JA, HD3 true container isolation, HD4 split+classifier),
  closer updated to 211 checks.
- CASAN_HARDENING_STATUS.md: Phase 6 deep-gap closers table; test inventory
  175→211 (12 suites); C7/multilingual moved out of planned; C6 planned→partial
  (real isolation done); honest claim → H4 83, H2 82, H5/H6 stay 80 (infra-bound).
- scoring-report-02-after-competition.md: current state — 211/0, H4 80→83,
  H2 80→82, avg 80.9→81.6, lowest harness still 80 (H5/H6), 3-milestone table.
- README claim boundary: deep-gap closers listed; totals 175→211; H4/H2 bumps.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-06 09:51:39 +09:00
thanhnvandClaude Fable 5 571f5d8cc3 docs(arch): add projectable visual architecture page (casan-architecture.html)
Self-contained, full-screen-projectable HTML visualizing the base → competition
→ now journey: OLED-glass design, 3-milestone timeline, 7 harness bento cards
(color-coded H1–H7 with each file's role), 5 upgrade-track tables (Track A /
C-MVP / Evidence / H5+ / H6+) with gap-closed columns, honest-maturity panel +
one-line flow diagram. UTF-8 standalone doc (opens directly in a browser),
IntersectionObserver scroll reveals, custom cubic-bezier motion, mobile fallback.
Verified rendering in preview. Cross-linked from both architecture md files.
Adds .claude/launch.json (static server for local preview of optimize-docs).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 12:48:32 +09:00
thanhnvandClaude Fable 5 bb1dc8ad9f docs(arch): add before/after architecture files vs the vanilla spec-kit base
Two companion docs tracing the .specify/scripts evolution from the 5-file
spec-kit scaffold shown in the file tree:

- CASAN_ARCHITECTURE_BEFORE.md — state brought to the competition (freeze
  fbcef96): base 5 → 37 scripts implementing all 7 harnesses, each file's
  purpose grouped by H1–H7, + the 8 competition test suites. Honest maturity:
  demo/PoC (~3.0/5).
- CASAN_ARCHITECTURE_AFTER.md — the feat/plan07-track-a-hardening upgrades:
  37 → 60 scripts (+23) grouped by Track A / Track C-MVP / Evidence Pack /
  H5+ / H6+, each new file's purpose + the gap it closes, notes on in-place
  modifications (strict fail-closed, cost caps, KMS rotate, window breaker),
  + the 6 new test suites (+96 checks). Honest maturity: Level 4 proven by
  attack (~4.0/5), not full production.

File counts verified against git (ls-tree fbcef96 vs HEAD).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 12:21:28 +09:00
thanhnvandClaude Fable 5 fa9f91eb37 docs(video): add per-attack rebuttal companion (CASAN_VIDEO_DEFENSE.md)
Deep-dive defense ammo tied to the video scenes so a presenter can field
follow-up questions live. For each of the 26 attacks across the 3 videos:
plain-language what-it-is, why-dangerous, how-CASAN-blocks (flagging which
layer is deterministic vs the 3-4 AI scenes), likely challenge question +
answer, and the honest limit. Opens with 3 "mantra" lines that cover most
hard questions and closes with a 10-toughest-questions cheat sheet.

Link it from CASAN_SCRIPT_2VIDEO.md. Numbers kept consistent: V1 = 23 vectors,
V2 = 140 checks, V3 = 175 checks; ~90% of controls are deterministic.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 01:51:12 +09:00
thanhnvandClaude Fable 5 459fe7f4a3 docs(video): fix Video 1 close — "4 layers" conflated checkpoints with harnesses
The cross-layer showpiece stops ONE attack at 4 sequential checkpoints
(injection block → tool-input schema reject → runaway timeout → signed audit),
which span H4 + the Tool layer (H2) + H5 — not "four harness layers". That
contradicted the intro's "this video focuses on 3 layers (H4/H5/H6)". Reframe
the close as one attack caught at each successive stage, listing the 4 concrete
checkpoints, so it's accurate to the demo without mislabeling them as harnesses.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 01:38:08 +09:00
thanhnvandClaude Fable 5 71f85f3e80 docs(video): frame the 3 videos as basic → advanced → deepest-gap
Add an explicit difficulty/depth progression the presenter can say out loud:
- Video 1 = basic & classic attacks (everyone must block these)
- Video 2 = advanced evasion + proof-you-can-trust (not just harder attacks —
  half of it is production maturity: fail-closed, evidence pack, FP budget)
- Video 3 = deepest gaps — real governance & ops hardening

Update each video's subtitle and the quick-reference table with a Level column.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 01:34:53 +09:00
thanhnvandClaude Fable 5 e399c1f844 docs(video): rewrite 2-video script for non-technical audience + add Video 3 (H5+/H6+ hardening)
- Rewrite Video 1 (attack battery) and Video 2 (Track A + C-MVP + Evidence Pack)
  in plain, presentation-friendly language: each scene now has a 【what's on screen】
  cue + 🎙️ spoken line, with everyday analogies (sealed ledger, fresh signature,
  key-in-a-vault) so a non-technical viewer follows while the terminal video plays.
- Add Video 3 covering the feat/plan07-track-a-hardening work that was NOT in the
  recorded videos: H5+ governance (approval-identity, KMS-in-vault, external WORM)
  and H6+ AgentOps (live alerting + dead-letter, provider-API cost reconciliation,
  stale-aware hosted dashboard, sliding-window breaker). Framed as the
  "lift the weakest link" arc, ending on the honest 175/0 · Level 4 · not-yet-full-
  production close.
- Add a quick-reference table + 60-second highlight cut for presenters.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 01:30:24 +09:00
thanhnvandClaude Fable 5 da66a36f97 feat(h6): AgentOps hardening — live alerting, provider-API reconcile, hosted dashboard, window breaker (79→80)
Close the three gaps the scoring report itself flagged for H6 plus V15,
each as a real MVP + fail-able adversarial test (same pattern that lifted H5):

- D1 alert-dispatch.sh: alerts POST to a real HTTP webhook (severity routing,
  dedup window, retry) + dead-letter queue with redelivery; fail-loud in strict.
  Wired into agent-metrics.sh so a failing step pages live end-to-end.
- D2 provider-usage-fetch.sh + telemetry-reconcile.sh: pull usage from a provider
  usage HTTP API (all-or-nothing schema gate, fail-loud) + reconcile local vs
  provider ground truth — token under-reporting/hidden runs => TELEMETRY_DISCREPANCY.
- D3 dashboard-serve.sh + dashboard-server.py: serve the dashboard over HTTP with
  a stale-aware /healthz probe (fresh=200 ok, telemetry silent-death=503 stale).
- D4 circuit-breaker-check.sh: sliding-window failure-rate breaker (V15) — interleaved
  successes no longer evade the consecutive-failure breaker (CIRCUIT_OPEN_WINDOW).

New suite phase-h6-agentops-tests.sh: 20/20, all live against local HTTP endpoints
(webhook sink, mock provider API, dashboard server) — deterministic, no model needed.

Also fix sign-policy-bundle.sh key-sync invariant: the local-fallback branch only
exported policy-public.pem when generating a NEW key, so a Vault-DOWN run after a
Vault-signed run verified a local-key signature against the Vault pubkey (RSA padding
error, run-casan4 died mid-suite). Now always re-exports the pubkey before signing —
same fix class as tool-audit-lib.sh / governance-check.sh.

Full battery re-run sequentially: 175/175 PASS, 0 FAIL across 8 suites
(KMS SKIP this run — Vault down; validated live 2026-07-04). Docs synced:
scoring-run-report (H6 79→80, no harness below 80, 155→175), CASAN_HARDENING_STATUS
(Phase 5 D1–D4), Plan-07, submission README, and run-hardening.sh (H6+ scenes HO1–HO4).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
2026-07-05 01:14:46 +09:00
thanhnvandClaude Opus 4.8 00aabfac4d feat(video): add H5+ governance-hardening scenes to run-hardening.sh
New "H5+ — Governance Hardening" section (before Evidence Pack) demoing the three
lowest-harness fixes with on-screen verdicts:
- HG1 approval-identity: env-var approver → BLOCKED; reviewer-signed → APPROVED.
- HG2 KMS key: sign/verify via Vault Transit + rotate + non-exportable (skip-aware
  when no Vault; shows the docker one-liner to enable it live).
- HG3 WORM audit: ship anchors → in-sync; roll back local audit → AUDIT_GAP_DETECTED.
Closer updated: H5 76→~85 line, total 155 checks. Verified end-to-end (exit 0,
live Vault: rotate v3 + non-exportable confirmed).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 23:19:52 +09:00
thanhnvandClaude Opus 4.8 9d63f7338c docs: short closing-only narration for each video (CASAN_CHOT_2VIDEO.md)
One plain-language closing paragraph per video for time-constrained delivery:
Video 1 (attacks blocked, proven by exit codes) and Video 2 (production
hardening + signed evidence pack). Jargon-light, keeps both closers.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 22:19:39 +09:00
thanhnvandClaude Opus 4.8 b1cb7b59de docs: condensed 2-video script (Video 1 + Video 2 only)
CASAN_SCRIPT_2VIDEO.md — tight summary voiceover for just the two demo videos,
one screen each: Video 1 (battery: H4/H5/H6 + chain) and Video 2 (hardening
Track A + C-MVP + Evidence Pack). Strips the per-scene cues/notes/soundbites of
the full narration; keeps the spoken beats and both closers.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 22:17:03 +09:00
thanhnvandClaude Opus 4.8 7870eb51e4 docs: professional voiceover/teleprompter script for both attack demo videos
CASAN_VIDEO_NARRATION.md — spoken lines timed to each on-screen scene of
run-all.sh (Video 1) and run-hardening.sh (Video 2). Golden rule: terminal says
WHAT, narrator says WHY — never read the screen. Per-scene ▶cue / 🎙️line / ⏸pause
markers, hero-scene emphasis (A5, B1, B6, D1, chain, HA1, HA4, HE3 money-shot),
clustered narration for fast scenes, delivery notes (tone/pace/verdict timing),
a soundbite bank for judge Q&A, and a duration table (~13-16 min). Aligned with
existing claims: Level 4 proven, AI as optional escalation, sandbox scaffold
honesty, "casan-old wins a demo; casan5 survives production".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 22:01:08 +09:00
thanhnvandClaude Opus 4.8 7f81120999 docs: attack catalog (visual) + playbook (internal) covering both demo scripts
- CASAN_ATTACK_CATALOG.md: one-page visual map of ALL ~38 attack scenes from
  run-all.sh (A1-A9, B1-B8, D1-D6, chain) + run-hardening.sh (HA1-7, HC1-4,
  HE1-4). Legend for 12 attack directions (input/output/artifact/tool-out/audit/
  telemetry/cost/action/supply/exfil/runtime/evidence), a defense-in-depth
  diagram, per-scene matrix with control + verdict + AI marker.
- CASAN_ATTACK_PLAYBOOK.md: internal deep-dive — per attack: scenario, why
  dangerous, exact blocking mechanism, AI-or-deterministic, verify command +
  expected result, threat-model ref. Prominent AI-usage answer: only A3/A8/D4
  invoke the model; everything else is deterministic. Semantic AI is an optional
  escalation that only ADDS a block, never removes one.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 00:21:21 +09:00
thanhnvandClaude Opus 4.8 2ffbda3fad feat(video+deck): Part 2 hardening demo script + slide deck update
- run-hardening.sh: narrated "Part 2" battery in the same visual language as
  run-all.sh (card/attack/guard/cmd + on-screen exit code with ⛔/✋/✅ verdict
  chips). Scenes: Track A (HA1-7: homoglyph/zero-width/base64, strict fail-closed,
  telemetry tamper, cost slow-boil/spray, benign FP=0%), Track C-MVP (HC1-4:
  action-gate, supply-chain, data-exfil, sandbox), and the Evidence Pack
  money-shot (HE1-4: pack → verify VALID → tamper 1 byte → TAMPERED exit=1 →
  certified-only-when-earned). Verified end-to-end (exit 0).
- CASAN_SLIDE_DECK.html: reflect implemented vs planned honestly. Updated H4/H5/H6
  AFTER columns, evidence terminal (140 checks), threat-model reframed to
  "identified AND Track A closed", readiness meters bumped to ~3.8-4.0, Track C-MVP
  reframed to DELIVERED (29/29), roadmap marks 07+09 done. Added 2 slides:
  "Track A delivered" and "Evidence Pack money-shot" (18 slides total).
- video guide: Part 1 (run-all) + Part 2 (run-hardening) with pre-flight note.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-03 23:51:11 +09:00
thanhnvandClaude Opus 4.8 fbcef967e5 chore(freeze): snapshot demo state before Plan-07 hardening work
Freeze current submission/demo baseline:
- casan-next-plans/: full task-level plan set (Plan 00 index + 02/04/06/07/08/09/12, QA, slide deck)
- optimize-docs/video-steps/: per-vector scene breakdown (commands/screen-text/script) + start-tmux
- run-all.sh / scorecard.sh / map-live.sh: REAL=1 live-battery wiring
- regenerated evidence + audit/telemetry logs from live REAL=1 run
- submission README + video recording guide updates
- dry-run pipeline logs for 001-okr-web-app

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-03 22:03:44 +09:00
thanhnvandClaude Opus 4.8 392f190b7e feat(model): implement OpenAI/Anthropic cloud backends + cloud-aware judge gate
The cloud branch of model-call.py was a stub (cloud_backend_not_implemented,
failed even with a key set); casan-step.mjs gated the judge on a hard-coded
Ollama ping. Wire up the real cloud path so a model can run without Ollama.

- model-call.py: add call_openai() and call_anthropic() (raw urllib, no new
  dependency — matches the existing call_ollama). Endpoints hard-pinned to the
  SSRF allowlist; keys read from env, never logged. Anthropic sends no
  temperature/thinking (rejected as 400 on Opus 4.8/4.7; omitting thinking
  keeps the terse one-word classify/judge answer). main() routes by
  ollama:/openai:/anthropic: prefix; key-unset still fails closed honestly.
  provider-usage.jsonl cost_source is per-backend, keeping ollama's exact
  "ollama_local_real_tokens" tag that evidence/tests key on.
- casan-step.mjs: ollamaAvailable() -> modelAvailable() — when
  CASAN_MODEL_PRIMARY is a cloud spec with its key set, the judge runs through
  the cloud path; otherwise it pings local Ollama as before. Default
  (unset CASAN_MODEL_PRIMARY) is unchanged.
- CASAN_MASTER_RUNBOOK.md: update sections 0/1/4/7/8 — cloud is now
  implemented (not a stub); keep the honest "untested with a real key" +
  CA-cert caveats.

Not verified against a live API key (none available); confirmed key-set makes
a real HTTPS call and key-unset fails closed. Gates unchanged:
security-gate PASS=11 FAIL=0, adversarial PASS=44 FAIL=0.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 11:38:59 +09:00
thanhnvandClaude Opus 4.8 c08d119381 feat(demo): REAL=1 runs the live attack battery through the production wrapper
Add REAL=1 to the video-steps demo so attack vectors flow through the real
production entry-point instead of calling sub-scripts directly.

- run-all.sh: REAL=1 feeds each H4 vector (A1/A2/A4/A6/A7 + cross-layer step
  1) as the INPUT of an agent step run through casan-harness.sh, so the
  BLOCK/PASS verdict is produced by the wrapper itself (H4-in -> H5 -> H6 ->
  exec -> H4-out) exactly as when the real pipeline meets malicious input.
  After the battery it runs a real pipeline slice (STEP1 okr.srs via
  casan-harness.sh -- node casan-step.mjs) and shows audit.jsonl growing by a
  real record. An inline inventory documents which vectors intentionally keep
  calling a single control directly (artifact-scan, audit tamper/re-forge,
  detectors on synthetic telemetry) and why. Default mode (no REAL) unchanged.
- map-live.sh: show the PIPELINE (STEP1) row only under REAL=1, driven by a
  mode sidecar file written by run-all.sh.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-07-03 10:05:06 +09:00
thanhnvandClaude Sonnet 4.6 2f06662f5d scorecard: chấm live 2 mục hardcode; pipeline: fallback real + honest scoring doc
Hướng A — scorecard.sh (video demo):
- h5_1 approval workflow: hardcode 0 → governance-check deploy live (approval_required)
- h6_2 hallucination rate: hardcode 0 → hallucination-scan phân biệt dirty>clean live
- "N/5 mục" chuyển từ text cứng sang đếm động
- H4/H5/H6 → 100/100 (5/5 gate live), Average 57.9 → 90.0

Hướng B — run-casan-pipeline.mjs:
- fallback: stub 'exit 9' → 'cat /nonexistent' (real failure, nhất quán adversarial T3)
- drift: giữ so fallback-output vs golden (clean run=1.0); năng lực phát hiện
  drift thật chứng minh ở adversarial suite
- Full 12-step run verify: H1 CONTEXT_VALID=24, H2 tool-audit records=25 signed,
  H5 audit-chain records=22 signed, H6 provider_telemetry per-step thật, H7 rollback real

phase3-real-run-scoring.md: giải thích vì sao scorecard cũ cho H5=60/H6=80
(hardcode), phân biệt scorecard-90 vs re-score-84 (2 mục đích khác nhau).

Verify: adversarial 44/0, security-gate 11/0/0, pipeline 12 steps OK.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
2026-07-03 00:18:43 +09:00
thanhnv 928f074b5f feat: add optimize dsoc 2026-07-02 22:17:03 +09:00