Standard production layout: the OKR app (was nested under AINative_OKR_CASAN5/) is now
the repository root. No more wrapper directory.
- Promote AINative_OKR_CASAN5/* -> repo root (backend/ frontend/ packages/ apps/
.specify/ docs/ infra/ nginx/ scripts/ + configs). Merge tool dirs: .gitea (kept the
active deploy ci.yml, added harness-ci.yml + runbooks), .claude (agents/commands +
launch.json), .github moved up.
- Remove redundant: 00_SUBMISSION_PACKAGE, scattered root notes (FPT_CASAN_Full.md,
tu-tuong-casan.md, casan-tu-sinh..., casan_harness_assessment.md, source-review...,
README_CASAN5_REFINED.md), casan-next-plans/ and optimize-docs/ (competition/planning
artifacts — roadmap + design history preserved in git log / commit messages).
- Update all references to the old layout:
- .gitea/workflows/{ci,harness-ci}.yml, .github/workflows/{ci,deploy}.yml:
working-directory .; drop AINative_OKR_CASAN5/ prefix; .specify/{tests,scripts}
-> packages/casan-harness/... (.specify/logs state kept)
- .claude/launch.json, .gitea/*-runbook.md: path prefixes
- CLAUDE.md, README.md: docs/input -> apps/okr/domain/input
- policy-bundle.yaml: 8 policy paths -> packages/casan-harness/...; manifest re-signed
- secrets-scan.sh: fixture excludes -> new package/domain paths.
Full gate from the new root: PASS=64 FAIL=0 SKIP=3.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
11 KiB
CASAN Phase 3 — Wave 4 Prompt
Mang file này sang máy mới và paste làm system prompt / đầu conversation cho Claude.
Repo cần clone: Output_CASAN5_REFINED/AINative_OKR_CASAN5 (branch main).
Context — trạng thái sau Wave 3
Đây là dự án AINative_OKR_CASAN5 đang được hardening theo framework FPT CASAN (7 harnesses H1–H7, mỗi harness 0–100 điểm). Mục tiêu: tất cả harness > 80, target ~90, chỉ real evidence — không fake.
Scores ước tính sau Wave 3 (cần independent audit để xác nhận)
| Harness | Score | Trạng thái |
|---|---|---|
| H1 Context | ~82 | context-validate.sh hoạt động nhưng có 12 trace files thiếu |
| H2 Tool | ~82 | rate-limit, schema validation, tool-exec timeout wired |
| H3 Evaluation | ~84 | Real app + backend tests; model judge gate wired; frontend test = tsc --noEmit only |
| H4 Security | ~85 | Semantic injection (recall=0.85 với 30 samples); artifact scan; secrets scan; circuit breaker |
| H5 Governance | ~82 | RSA-anchored audit chain; signing key off-repo; KMS/WORM BLOCKED (no AWS) |
| H6 AgentOps | ~82 | Real per-step tokens; cost-spike detect; hallucination signals |
| H7 Orchestration | ~82 | Real DAG; real rollback/drift/fallback |
| Average | ~83 | Cần ≥85 cho mọi harness để push thêm |
Verify baseline (chạy trước khi làm bất kỳ thứ gì)
cd Output_CASAN5_REFINED/AINative_OKR_CASAN5
bash .specify/tests/run-casan4-harness-tests.sh # expect 35 PASS / 0 FAIL
bash .specify/tests/adversarial-harness-tests.sh # expect 40 PASS / 0 FAIL
bash .specify/scripts/bash/verify-audit-chain.sh # expect AUDIT_CHAIN_VALID anchor=signed
bash .specify/scripts/bash/security-gate.sh # expect PASS=9 FAIL=0 (nếu Ollama up)
Tất cả phải PASS trước khi bắt đầu Wave 4. Nếu có FAIL, dừng lại và điều tra.
Môi trường Wave 4 cần
| Requirement | Lý do |
|---|---|
npm install được (network) |
Mở khóa frontend Vitest (H3 gap lớn nhất) |
| Node.js ≥ 18 | Chạy backend tests + pipeline |
Ollama với ornith:9b (hoặc model khác) |
Tiếp tục semantic layer (H4) |
(Optional) ANTHROPIC_API_KEY hoặc OPENAI_API_KEY |
Nâng recall từ 0.85 lên gần 1.0 |
Wave 4 Tasks
WV4-A: H3 Frontend Runtime Tests (CRITICAL — gap lớn nhất)
Hiện trạng: frontend/package.json test script = tsc --noEmit (chỉ type-check, không runtime). vitest không có trong devDependencies.
Yêu cầu:
- Thêm
vitestvà@testing-library/react+@testing-library/jest-dom+jsdomvàofrontend/devDependencies. - Cập nhật
frontend/package.jsonscript:"test": "vitest run". - Viết ít nhất 5 real Vitest tests trong
frontend/src/— phải cover:- Component render (ít nhất 1 test kiểm tra DOM output thực sự)
- Role-based access logic (ADMIN/MANAGER/EMPLOYEE)
- Form validation (Zod schema — invalid quarter format returns error)
- Progress calculation (0–100 range)
- API error handling (mock Axios, kiểm tra error state)
npm test -w frontendphải chạy được và có thể FAIL (không phải hardcoded pass).- Fail-before proof: viết 1 test sai trước, chạy → FAIL, sửa → PASS. Ghi lại cả hai trong evidence.
Acceptance test:
npm install -w frontend
npm test -w frontend # must show real test results with pass/fail counts
# Modify one assertion to be wrong → re-run → must FAIL
# Restore → re-run → PASS
Files: frontend/package.json, frontend/src/__tests__/ (tạo mới), frontend/vite.config.ts (thêm test config).
WV4-B: H1 Fix 12 Missing Trace Files
Hiện trạng:
bash .specify/scripts/bash/context-validate.sh \
docs/output/output_logs/001-okr-web-app/pipeline-context.yaml
# → CONTEXT_INVALID missing=12: .specify/logs/trace/agentops-*.json
12 trace files được referenced trong pipeline-context.yaml nhưng không tồn tại trên disk (chúng là traces từ run cũ, đã bị xóa).
Yêu cầu — chọn 1 trong 2 approach (không fake):
Option A (preferred): Tạo stub trace files hợp lệ cho 12 missing traces:
- Mỗi file là JSON hợp lệ với
trace_id,step,status,latency_ms,timestamp. - Dữ liệu phải consistent với pipeline run thực (không điền số tùy tiện).
- Sau khi tạo:
context-validate.shphải PASS.
Option B: Cập nhật pipeline-context.yaml:
- Xóa references đến 12 trace files không còn tồn tại.
- Thêm annotation
retention_gap: truevới giải thích. - Thêm
ttl_seconds: 86400cho các entries còn lại. - Sau khi sửa:
context-validate.shphải PASS.
Acceptance test:
bash .specify/scripts/bash/context-validate.sh \
docs/output/output_logs/001-okr-web-app/pipeline-context.yaml
# → CONTEXT_VALID (exit 0)
WV4-C: H6 Real Pipeline End-to-End Run
Hiện trạng: Pipeline đã chạy một lần (Phase 2), nhưng sau nhiều thay đổi (judge gate, tool-exec wiring) chưa chạy lại. Per-step telemetry dùng records từ Wave 1-3 model router calls, không phải từ pipeline step execution.
Yêu cầu:
- Chạy
node scripts/run-casan-pipeline.mjstừ đầu (hoặc từ step 01-srs nếu pipeline incremental). - Xác nhận từng step trong
docs/output/output_logs/001-okr-web-app/reports/được tạo ra. - Kiểm tra
.specify/logs/level5/provider-usage.jsonlcó records mới từ run này (distinguish bằngrun_idhoặc timestamp). - Kiểm tra
.specify/logs/cost/metrics.jsonlcó latency/cost records từ pipeline steps. - Chạy
cost-spike-detect.sh— phải detect được bất kỳ step nào có token > 3× median.
Nếu pipeline bị lỗi (backend chưa start):
- Ghi rõ:
PIPELINE_RUN_BLOCKED reason=<lý do cụ thể>trong evidence. - KHÔNG fake pipeline output.
Acceptance test:
node scripts/run-casan-pipeline.mjs
bash .specify/scripts/bash/cost-spike-detect.sh # should show per-step data
WV4-D: H4 Multi-Provider Recall (nếu có cloud key)
Hiện trạng: Local ornith:9b đạt recall=0.85 trên 30-sample corpus. 3/20 novel paraphrases vẫn trốn thoát.
Nếu ANTHROPIC_API_KEY hoặc OPENAI_API_KEY có sẵn:
- Implement cloud backend trong
.specify/scripts/bash/model-call.py:anthropic:prefix → gọiapi.anthropic.comvới key từ envopenai:prefix → gọiapi.openai.com- Vẫn giữ SSRF guard (chỉ allow api.anthropic.com / api.openai.com)
- Vẫn fail-closed (malformed → INJECTION/REJECTED)
- Chạy
phase3-redteam-metrics.shvớiCASAN_MODEL_PRIMARY=anthropic:claude-haiku-4-5-20251001(hoặcopenai:gpt-4o-mini). - So sánh recall: local=0.85 vs cloud=?
- Nếu cloud recall > 0.85 → ghi rõ improvement trong evidence.
- KHÔNG claim cloud backend nếu key unset. Script đã có logic
cloud_backend_unavailable.
Nếu không có cloud key: Skip WV4-D, ghi BLOCKED: no cloud API key.
WV4-E: Adversarial Suite Phải Vẫn PASS=40+ FAIL=0
Sau mỗi thay đổi ở WV4-A/B/C/D, bắt buộc chạy:
bash .specify/tests/adversarial-harness-tests.sh
# expect PASS ≥ 40, FAIL = 0
Nếu có thêm tests mới từ WV4-A (frontend runtime), add vào adversarial suite:
- 1 test: frontend test passes khi code đúng
- 1 test: frontend test fails khi assertion sai (fail-before)
WV4-F: Final Security Gate
Sau tất cả wave:
bash .specify/scripts/bash/security-gate.sh
# expect PASS ≥ 9 (thêm nếu add gates mới), FAIL = 0
Add gate mới nếu WV4-A thành công:
run "frontend runtime tests (WV4-A)" npm test -w frontend
Integrity Rules (bắt buộc — không ngoại lệ)
- Không fake PASS, APPROVED, token usage, model verdicts, cost, recall, rollback, audit results.
- Không hardcode expected PASS results.
- Mọi test phải có thể FAIL — nếu test không thể fail thì đó không phải test.
- Không dùng
.specify/tests/generate-casan-demo-context.py(script demo, không phải real evidence). - Không claim "semantic detection" nếu chỉ dùng regex. Gọi đúng tên:
model-layerhoặcregex-only. - Không claim cloud/hybrid judging nếu key unset. Ghi
BLOCKEDhoặccloud_backend_unavailable. - Không claim KMS/WORM — AWS không được cấu hình. Ghi là local key, document KMS là production requirement.
- Không claim scores > thực tế. Ghi
estimated+ note cần independent audit. - Nếu infra thiếu, ghi BLOCKED với lý do cụ thể. Không mark PASS thay.
- Không tăng score mà không có evidence thực. Evidence = real command output, real file content, real test results.
Files quan trọng cần đọc trước khi làm
.specify/scripts/bash/
security-check.sh # H4 input/output filter (regex + semantic)
model-call.py # Ollama/cloud router (SSRF-guarded, fail-closed)
model-router.sh # thin wrapper over model-call.py
casan-harness.sh # unified H4→H5→H6→H4 wrapper (tool-exec wired)
tool-exec.sh # hard timeout wrapper
governance-check.sh # H5 audit chain (11-field, RSA-anchored)
tool-registry-gate.sh # H2 per-agent permission + rate-limit
artifact-scan.sh # H4 indirect injection scanner (Wave 3)
secrets-scan.sh # H4 secrets lifecycle (Wave 3)
circuit-breaker-check.sh # H4 no-bypass + circuit breaker (Wave 3)
security-gate.sh # one-command aggregate (9 gates)
.specify/tests/
run-casan4-harness-tests.sh # 35 deterministic tests
adversarial-harness-tests.sh # 40 adversarial tests (includes Wave 3)
phase3-judge-gate-tests.sh # 5 WP-B judge gate tests
phase3-model-router-tests.sh # 7 model router tests
phase3-redteam-metrics.sh # red-team corpus metrics runner
.specify/security/
redteam-corpus.jsonl # 30 samples (20 injection, 10 benign)
prompt-filter.yaml # regex blocklist
scripts/
casan-step.mjs # pipeline step runner (judge gate wired into 04/06/10)
run-casan-pipeline.mjs # full pipeline orchestrator
docs/output/casan/
phase3-wave3-results.md # Wave 3 complete results
phase3-wave2-results.md # Wave 2 results (WP-C cost, WP-S2 corpus, WP-S8 gate)
phase3-wave1-results.md # Wave 1 results (model router, semantic injection)
TEAM-HANDOFF-PLAN.md # Full history + scores + verify commands
Output cần tạo
docs/output/casan/phase3-wave4-results.md— kết quả thực, honest, theo format của Wave 1-3 results.docs/output/casan/phase3-evidence/wave4/— evidence logs (command outputs, test results).- Cập nhật
TEAM-HANDOFF-PLAN.mdvới scores sau Wave 4. - Nếu có tests mới: update
adversarial-harness-tests.shvàsecurity-gate.sh.
Thứ tự ưu tiên
- WV4-A (H3 frontend Vitest) — impact lớn nhất, unblocked trên máy mới
- WV4-B (H1 trace files) — context-validate cần PASS
- WV4-E (adversarial suite vẫn green) — sau mỗi change
- WV4-C (H6 pipeline re-run) — nếu backend stack hoạt động
- WV4-D (H4 cloud recall) — nếu có cloud key
- WV4-F (final security gate) — cuối cùng
Bắt đầu bằng verify baseline. Nếu baseline không PASS → dừng và báo cáo.