HarnessAthon Submission Package
One-Line Positioning
We upgraded the SDD Speckit OKR pipeline from the CASAN Level 3→4 transition to CASAN Level 4 — proven by attack: H4/H5/H6 (former GAPs 20/25/30) now defeat ~23 live adversarial vectors with on-screen exit codes. Level 5 is stated as roadmap (IdP, WORM log storage, provider-telemetry API, hosted dashboard) — not claimed as achieved.
Open These First
presentation/HarnessAthon_CASAN_Level5_Demo.pptxdocs/00_task_breakdown.mddocs/01_submission_checklist.mddocs/02_pitch_script.mddocs/03_judge_qna.mddocs/04_evidence_map.mdvideo/01_video_recording_guide.mdai_context/AI_README.md
Main Evidence
| Evidence | Source |
|---|---|
| Harness test report | ../AINative_OKR_CASAN5/docs/output/casan/evidence/harness-test-report.md |
| CASAN assessment | ../AINative_OKR_CASAN5/docs/output/casan/casan-level4-assessment.md |
| Before/after scorecard | ../AINative_OKR_CASAN5/docs/output/casan/before-after-scorecard.md |
| Demo pipeline context | ../AINative_OKR_CASAN5/docs/output/output_logs/casan-demo/pipeline-context.yaml |
| Central dashboard | ../AINative_OKR_CASAN5/docs/output/casan/central-agentops-dashboard.html |
| Final zip | ../AINative_OKR_CASAN5.zip |
Folder Purpose
| Folder | Purpose |
|---|---|
docs/ |
Checklist, pitch script, judge Q&A, evidence map, AI-optimized structure notes |
presentation/ |
PPT deck and deterministic generation source |
video/ |
Screen-recording guide and narration outline |
evidence/ |
Evidence index pointing to canonical generated evidence |
ai_context/ |
One-page context for AI/teammate review |
Verification Command
cd ../AINative_OKR_CASAN5
bash .specify/tests/run-casan4-harness-tests.sh
Expected: all PASS, including H4/H5/H6 and Level 5 evidence checks.
Claim Boundary
- CASAN Level 4: achieved.
- CASAN Level 5: demonstrated locally with signed policy, provider telemetry import, shared harness registry, fallback, drift detection, rollback, KPI feedback, and central dashboard.
- Enterprise production Level 5 still needs live IdP, WORM storage, live provider API integration, and hosted dashboard.
Production hardening — implemented vs planned
Following Plan-07/Plan-09, the following hardening is implemented and tested
(each control has an executable adversarial test that fails if the control is
removed). Full status: casan-next-plans/CASAN_HARDENING_STATUS.md.
- Implemented + tested (Plan-07 Track A): H4 strict semantic fail-closed
(
CASAN_SECURITY_STRICT), unicode/encoding normalization (homoglyph, zero-width, fullwidth, base64/hex), tool-output injection scan, H5 telemetry-integrity signing, H6 absolute + cumulative + cold-start cost controls, benign/false-positive budget gate (FP ≤ 3%, adversarial block ≥ 95%, CRITICAL = 100%). - Implemented + tested (Plan-07 Track C-MVP): tool-authorization / action gating, supply-chain gate, data-exfiltration guard.
- Scaffold + tested (Track C-MVP): runtime sandbox — static policy +
ulimitbackstops. Not kernel isolation; production needs container--network=none --read-only --pids-limit/ nsjail. - Implemented + tested (Plan-09 MVP): Evidence Pack —
casan pack/casan verify-pack(tamper-evident manifest + signed head + certified-run gate). - Implemented + tested (H5 governance hardening): approval-identity (reviewer cryptographically signs the request + role authorization — env-var approver no longer enough); KMS key management (Vault Transit sign + rotation + non-exportable, validated live); external WORM audit ledger (rollback + tamper detection).
- Implemented + tested (H6 AgentOps hardening): live alert dispatch (webhook +
dedup + dead-letter, fail-loud, end-to-end from a failing step); provider-telemetry
API fetch + local-vs-provider reconciliation (catches token under-reporting);
hosted dashboard with stale-aware
/healthz; sliding-window circuit breaker (V15). - Implemented + tested (deep-gap closers, post-competition): incident response
- scoped kill-switch (C7, severity→auto-stop); multilingual VI/JA injection detection (0 FP on benign VI/JA); TRUE container runtime isolation (C6, kernel neutralises egress/host-read, validated live via Docker); split-injection (assembled-context scan) + classifier-injection resistance.
- Planned (NOT done — do not claim as production-ready): HSM + KMS-by-default, live IdP (OIDC/JWT), true WORM store (S3 Object Lock), deployed dashboard host + managed alert channel, real billing-API telemetry, model-digest pinning.
Test totals: baseline 79 (run-casan4 35 + adversarial 44) preserved, +132 new
hardening checks (Track A 25, Track C-MVP 29, Evidence Pack 7, H5-approval 8,
H5-infra KMS+WORM 7, H6-agentops 20, C7-incident 15, H4-multilingual 7, C6-sandbox 6,
H4-split-inject 8) = 211, 0 fail — last full run 2026-07-05 (KMS + container
isolation validated live via Vault + Docker; evidence/scoring-run-report.md).
Fair maturity: H4 rose to 83 and H2 to 82 (deep-gap closers); H5/H6 stay at 80
(remaining gaps are infra) so the lowest harness is still 80 — CASAN Level 4, proven
by attack. See CASAN_HARDENING_STATUS.md. Because these live in separate suites,
the demo attack battery counts in video/01_video_recording_guide.md are unchanged.