Standard production layout: the OKR app (was nested under AINative_OKR_CASAN5/) is now
the repository root. No more wrapper directory.
- Promote AINative_OKR_CASAN5/* -> repo root (backend/ frontend/ packages/ apps/
.specify/ docs/ infra/ nginx/ scripts/ + configs). Merge tool dirs: .gitea (kept the
active deploy ci.yml, added harness-ci.yml + runbooks), .claude (agents/commands +
launch.json), .github moved up.
- Remove redundant: 00_SUBMISSION_PACKAGE, scattered root notes (FPT_CASAN_Full.md,
tu-tuong-casan.md, casan-tu-sinh..., casan_harness_assessment.md, source-review...,
README_CASAN5_REFINED.md), casan-next-plans/ and optimize-docs/ (competition/planning
artifacts — roadmap + design history preserved in git log / commit messages).
- Update all references to the old layout:
- .gitea/workflows/{ci,harness-ci}.yml, .github/workflows/{ci,deploy}.yml:
working-directory .; drop AINative_OKR_CASAN5/ prefix; .specify/{tests,scripts}
-> packages/casan-harness/... (.specify/logs state kept)
- .claude/launch.json, .gitea/*-runbook.md: path prefixes
- CLAUDE.md, README.md: docs/input -> apps/okr/domain/input
- policy-bundle.yaml: 8 policy paths -> packages/casan-harness/...; manifest re-signed
- secrets-scan.sh: fixture excludes -> new package/domain paths.
Full gate from the new root: PASS=64 FAIL=0 SKIP=3.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
67 lines
4.7 KiB
Markdown
67 lines
4.7 KiB
Markdown
# CASAN Phase 3 — Wave 2 Results
|
||
|
||
**Date:** 2026-06-30
|
||
**Executed by:** Claude (Opus 4.8), in CASAN5, against the live local model (`ollama:ornith:9b`).
|
||
**Integrity:** real numbers only; blocked items recorded as BLOCKED, not pass. No score >90 claimed.
|
||
**Evidence:** `docs/output/casan/phase3-evidence/wave2/`.
|
||
|
||
---
|
||
|
||
## Implemented (real, verified)
|
||
|
||
### WP-A + WP-S2 — Quantitative H4 (semantic injection, measured)
|
||
Built a labeled red-team corpus (`.specify/security/redteam-corpus.jsonl`: 10 novel injection
|
||
paraphrases that are NOT in the regex blocklist + 6 benign) and a runner
|
||
(`.specify/tests/phase3-redteam-metrics.sh`) that scores both layers against the **live model**.
|
||
|
||
Real result (`phase3-evidence/wave2/redteam-metrics.log`):
|
||
|
||
| Layer | precision | recall |
|
||
|---|:--:|:--:|
|
||
| regex-only | 1.00 | **0.00** (caught 0/10 paraphrases) |
|
||
| model layer | 1.00 | **0.90** (caught 9/10, 0 benign false-positives) |
|
||
|
||
`GATE model_recall>=0.8 AND model_recall>regex_recall → PASS`. This **quantifies** the model
|
||
layer's value: the regex blocklist is blind to novel paraphrasing (recall 0), the model recovers
|
||
90% recall at 100% precision. One paraphrase still slipped the local 9B (recall 0.9) — recorded
|
||
honestly; cloud or a larger model would raise coverage.
|
||
**Honest caveat:** corpus is 16 samples; the security DoD targets ≥30 — this is a real but
|
||
not-yet-DoD-sized corpus.
|
||
|
||
### WP-C — H6 real per-step billing + spike detection
|
||
- `.specify/scripts/bash/cost-spike-detect.sh` (median-based, flags steps > 3× median, exits non-zero).
|
||
- `provider-usage.jsonl` now has **18 real records with 9 distinct token values** (e.g. 218, 219, 220, 216) from real router calls — **not** the recycled 2778 sample.
|
||
- **Fail-able proof:** injecting a 4× step → `COST_SPIKE_DETECTED` (exit 2). The detector also flags the leftover 2778 sample records as anomalies vs. the real ~220-token calls — working as intended.
|
||
- `cost_usd=0.0`, `cost_source=ollama_local_real_tokens` — local inference is genuinely free; not faked pricing.
|
||
|
||
### WP-S8 — One-command security gate
|
||
`.specify/scripts/bash/security-gate.sh` aggregates the security-relevant checks into one verdict.
|
||
Real run (`phase3-evidence/wave2/security-gate.log`): **PASS=6 FAIL=0 SKIP=0** —
|
||
run-casan4 suite, adversarial suite, audit hash-chain (signed), tool-call audit (signed),
|
||
model-router tests, red-team H4 metrics. Live-model gates SKIP (not fail) if the tunnel is down.
|
||
|
||
## Status of the other work packages
|
||
| WP | Status |
|
||
|---|---|
|
||
| WP-D (H7 fallback/drift/undo) | Largely done across Wave 1 + earlier Phase 3: failure-driven fallback via real router (Wave 1 test 7), real drift between different artifacts, exception-style undo via `rollback-manager checkpoint`. Drift against the in-pipeline attempt1-vs-attempt2 artifacts is wired but the full pipeline run is blocked by baseline. |
|
||
| WP-E (H1 context) | Done — `context-validate.sh` red-able; on the real context it correctly flags 12 missing trace files (a real retention gap to fix, not a validator defect). |
|
||
| WP-F (H2 rate-limit + schema) | Done in Phase 3 push (`rate_limit_per_run`, `validate-tool-input.sh`), adversarially tested. |
|
||
| WP-G (H5 governance) | Offline done (signed re-forge-detection, key off-repo). KMS/WORM **blocked** (no AWS). |
|
||
| WP-B (H3 judge + frontend tests) | **Partial/blocked.** The router `--role judge` gives a real single-model judge usable now; wiring it into `casan-step.mjs` review gates and the frontend Vitest/RTL tests need the app baseline (`npm install` → network) which is blocked. True 2-of-3 consensus needs a 2nd independent provider (no cloud key). → Wave 3. |
|
||
|
||
## Blockers (recorded, NOT passed)
|
||
1. **App baseline** — `node_modules` absent, `npm install` needs network (sandbox-blocked). Blocks WP-B frontend tests, `casan-step.mjs` judge wiring, `npm audit` (WP-S3), full pipeline run.
|
||
2. **Cloud provider** — keys unset → no multi-provider judge diversity (WP-B), no cloud-grade classify coverage.
|
||
3. **AWS/KMS/Object Lock** — not configured → WP-G KMS/WORM not claimed.
|
||
4. **Corpus size** — red-team corpus is 16, DoD wants ≥30 (expand in Wave 3).
|
||
|
||
## Files added/changed (Wave 2)
|
||
- new: `.specify/scripts/bash/cost-spike-detect.sh`, `security-gate.sh`; `.specify/security/redteam-corpus.jsonl`; `.specify/tests/phase3-redteam-metrics.sh`
|
||
- (Wave 1 carried in: `model-router.sh`, `model-call.py`, `phase3-model-router-tests.sh`, semantic escalation in `security-check.sh`)
|
||
|
||
## Score statement
|
||
**No average >90 claimed.** Wave 2 added *measured* H4 (recall 0→0.90), real per-step H6 telemetry +
|
||
spike detection, and a one-command security gate (6/6 green). App-dependent harnesses (WP-B frontend,
|
||
judge wiring) and KMS/WORM remain blocked. Scores update only after a full independent audit on a
|
||
green baseline.
|