Files
CASAN/00_SUBMISSION_PACKAGE/README.md
T
thanhnvandClaude Opus 4.8 2af67ef6a5 docs: re-score from a real full run (155/0) after H5 hardening
Ran all 7 suites sequentially on 2026-07-04 @ 00aabfa with Vault dev live so KMS
runs (not skips): run-casan4 35 · adversarial 44 · phase1-track-a 25 ·
phase2-track-c 29 · phase3-evidence 7 · phase-h5-approval 8 · phase-h5-infra 7
= 155 PASS / 0 FAIL; security-gate PASS=11 FAIL=0.

- scoring-run-report.md: fair re-score — H5 76→80 (approval-identity + KMS live
  rotate/non-exportable + WORM), lowest harness now H6=79, avg ~80.7/100, Level 4.
  Evidence lists the live KMS + WORM results.
- CASAN_HARDENING_STATUS.md: new Phase 4 (C4 approval-identity, B3 KMS, C5 WORM =
  implemented+tested); test inventory 140→155 (7 suites); planned→partial for
  KMS/approval/WORM with honest remaining gaps (live IdP, S3 WORM store, KMS default).
- Plan-07 §2: key-mgmt 2.5→4, policy-approval 2.5→4, external-audit 1.5→3.5;
  header now H5 76→80, lowest harness H6.
- INDEX row, README claim boundary, video-guide Q&A: 155 checks, H5=80, lowest H6=79.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-04 23:51:50 +09:00

87 lines
4.5 KiB
Markdown

# HarnessAthon Submission Package
## One-Line Positioning
We upgraded the SDD Speckit OKR pipeline from the CASAN Level 3→4 transition to **CASAN Level 4 — proven by attack**: H4/H5/H6 (former GAPs 20/25/30) now defeat ~23 live adversarial vectors with on-screen exit codes. Level 5 is stated as **roadmap** (IdP, WORM log storage, provider-telemetry API, hosted dashboard) — not claimed as achieved.
## Open These First
1. `presentation/HarnessAthon_CASAN_Level5_Demo.pptx`
2. `docs/00_task_breakdown.md`
3. `docs/01_submission_checklist.md`
4. `docs/02_pitch_script.md`
5. `docs/03_judge_qna.md`
6. `docs/04_evidence_map.md`
7. `video/01_video_recording_guide.md`
8. `ai_context/AI_README.md`
## Main Evidence
| Evidence | Source |
|---|---|
| Harness test report | `../AINative_OKR_CASAN5/docs/output/casan/evidence/harness-test-report.md` |
| CASAN assessment | `../AINative_OKR_CASAN5/docs/output/casan/casan-level4-assessment.md` |
| Before/after scorecard | `../AINative_OKR_CASAN5/docs/output/casan/before-after-scorecard.md` |
| Demo pipeline context | `../AINative_OKR_CASAN5/docs/output/output_logs/casan-demo/pipeline-context.yaml` |
| Central dashboard | `../AINative_OKR_CASAN5/docs/output/casan/central-agentops-dashboard.html` |
| Final zip | `../AINative_OKR_CASAN5.zip` |
## Folder Purpose
| Folder | Purpose |
|---|---|
| `docs/` | Checklist, pitch script, judge Q&A, evidence map, AI-optimized structure notes |
| `presentation/` | PPT deck and deterministic generation source |
| `video/` | Screen-recording guide and narration outline |
| `evidence/` | Evidence index pointing to canonical generated evidence |
| `ai_context/` | One-page context for AI/teammate review |
## Verification Command
```bash
cd ../AINative_OKR_CASAN5
bash .specify/tests/run-casan4-harness-tests.sh
```
Expected: all PASS, including H4/H5/H6 and Level 5 evidence checks.
## Claim Boundary
- CASAN Level 4: achieved.
- CASAN Level 5: demonstrated locally with signed policy, provider telemetry import, shared harness registry, fallback, drift detection, rollback, KPI feedback, and central dashboard.
- Enterprise production Level 5 still needs live IdP, WORM storage, live provider API integration, and hosted dashboard.
### Production hardening — implemented vs planned
Following Plan-07/Plan-09, the following hardening is **implemented and tested**
(each control has an executable adversarial test that fails if the control is
removed). Full status: `casan-next-plans/CASAN_HARDENING_STATUS.md`.
- **Implemented + tested (Plan-07 Track A):** H4 strict semantic fail-closed
(`CASAN_SECURITY_STRICT`), unicode/encoding normalization (homoglyph, zero-width,
fullwidth, base64/hex), tool-output injection scan, H5 telemetry-integrity signing,
H6 absolute + cumulative + cold-start cost controls, benign/false-positive budget
gate (FP ≤ 3%, adversarial block ≥ 95%, CRITICAL = 100%).
- **Implemented + tested (Plan-07 Track C-MVP):** tool-authorization / action gating,
supply-chain gate, data-exfiltration guard.
- **Scaffold + tested (Track C-MVP):** runtime sandbox — static policy + `ulimit`
backstops. **Not** kernel isolation; production needs container `--network=none
--read-only --pids-limit` / nsjail.
- **Implemented + tested (Plan-09 MVP):** Evidence Pack — `casan pack` / `casan
verify-pack` (tamper-evident manifest + signed head + certified-run gate).
- **Implemented + tested (H5 governance hardening):** approval-identity (reviewer
cryptographically signs the request + role authorization — env-var approver no
longer enough); KMS key management (Vault Transit sign + rotation + non-exportable,
**validated live**); external WORM audit ledger (rollback + tamper detection).
- **Planned (NOT done — do not claim as production-ready):** multilingual H4,
classifier/split-injection resistance, HSM + KMS-by-default, live IdP (OIDC/JWT),
true WORM store (S3 Object Lock), incident kill-switch, true sandbox isolation.
Test totals: baseline 79 (run-casan4 35 + adversarial 44) preserved, **+76 new**
hardening checks (Track A 25, Track C-MVP 29, Evidence Pack 7, H5-approval 8,
H5-infra KMS+WORM 7) = **155**, 0 fail — last full run 2026-07-04 with KMS live via
Vault (`evidence/scoring-run-report.md`). Fair maturity ~80/100 per harness; H5 rose
76→80 so the **lowest harness is now H6=79** (CASAN Level 4, proven by attack). See
`CASAN_HARDENING_STATUS.md`. Because these live in **separate** suites, the demo
attack battery counts in `video/01_video_recording_guide.md` are unchanged.