diff --git a/00_SUBMISSION_PACKAGE/evidence/scoring-run-report.md b/00_SUBMISSION_PACKAGE/evidence/scoring-run-report.md index 9736895..323505e 100644 --- a/00_SUBMISSION_PACKAGE/evidence/scoring-run-report.md +++ b/00_SUBMISSION_PACKAGE/evidence/scoring-run-report.md @@ -6,20 +6,23 @@ ## 1. Môi trường chạy | | | |---|---| -| Thời điểm | **2026-07-05 (UTC)** — chạy tuần tự lại toàn bộ sau H6-hardening | -| Git | `feat/plan07-track-a-hardening` (sau `2af67ef` + H6-hardening) | +| Thời điểm | **2026-07-06 (JST)** — chạy tuần tự lại sau B4 digest pinning + C4 mock IdP/OIDC + Plan-10 traceability | +| Git | working tree handoff continuation (commit pending) | | Model | Ollama `ornith:9b` @127.0.0.1:11434 — live | | KMS | Vault @127.0.0.1:8200 — **down lần chạy này** (KMS SKIP; đã validate **live** 2026-07-04 @ `00aabfa`) | | Runtime | node v24.12.0 · python 3.9.0 · openssl 3.6.3 · macOS | ## 2. Bằng chứng thô (100% thật, exit code on-screen — chạy tuần tự lại lần này) -- **175 test PASS / 0 FAIL** trên **8 suite**: +- **218 test PASS / 0 FAIL** trên **13 core harness suite**: run-casan4 **35** · adversarial **44** · phase1-track-a **25** · phase2-track-c **29** · - phase3-evidence-pack **7** · phase-h5-approval **8** · phase-h5-infra **7** · **phase-h6-agentops 20 (mới)**. + phase3-evidence-pack **7** · phase-h5-approval **12** · phase-h5-infra **7** · phase-h6-agentops **20** · + phase-c7-incident **15** · phase-h4-multilingual **7** · phase-c6-sandbox **6** · phase-h4-split-inject **8** · phase10-traceability **3**. +- Direct model-router suite: **10 PASS / 0 FAIL** (includes model-digest pin OK, mismatch BLOCK, mismatch WARN rollout mode). +- Frontend Vitest: **16 PASS / 0 FAIL**. Backend `npm test` is blocked by pre-existing app-test infra mismatch (`schema.prisma` provider MySQL but `setup-sqlite.mjs` applies the MySQL migration to SQLite). - `security-gate` aggregate: **verdict PASS=11 FAIL=0 SKIP=0** (run 2026-07-04). - H4 recall: model 0.85 > regex 0.00 (GATE PASS). Benign-FP: **fp_rate 0.00% · block_rate 100.00%** (95 mẫu benign EN/VI/JA + 12 vector). - **H5 (run 2026-07-04, Vault live):** - - Approval-identity: env-var approver → BLOCK; reviewer **ký** request + đúng role → APPROVED; chữ ký giả / sai role / replay / tự-duyệt → BLOCK (8/8). + - Approval-identity: env-var approver → BLOCK; reviewer **ký** request + đúng role → APPROVED; mock IdP JWT hợp lệ → APPROVED; JWT hết hạn / sai role / chữ ký giả → BLOCK (12/12). - KMS **live** (Vault Transit): sign→verify (v1) → **rotate** → sign→verify (v2) → **khoá NON-exportable** (export bị từ chối). - WORM: ship anchor → in-sync; rollback log local → **AUDIT_GAP_DETECTED**; sửa ledger → **AUDIT_LEDGER_TAMPERED**. - **H6 mới (chạy thật lần này — mọi check chạy LIVE qua endpoint HTTP local, cùng chuẩn "live" như Vault dev):** @@ -35,9 +38,9 @@ |----|---------|:--:|:--:|:--:|---|---|---| | H1 | Context | 90 | 84 | **84** | Good+ | pipeline-context pointer store, log-levels, context-validate | chưa nén/RAG context lớn (Plan-08 planned) | | H2 | Tool | 75 | 80 | **80** | Good+ | registry, idempotency, rate-limit, action-gate, supply-chain, tool-audit | sandbox mới scaffold (chưa cô lập thật) | -| H3 | Evaluation | 85 | 82 | **82** | Strong- | LLM-judge multi-gate live, golden dataset, auto-retry, judge-gate 5/0 | judge chỉ 1 model local, chưa eval-set độc lập quy mô | +| H3 | Evaluation | 85 | 82 | **82** | Strong- | LLM-judge multi-gate live, golden dataset, auto-retry, judge-gate 5/0, Plan-10 FR→code→test matrix 3/0 | judge chỉ 1 model local, chưa eval-set độc lập quy mô/line-level traceability | | H4 | Security | 20 | 80 | **80** | Good(đỉnh) | injection (direct/paraphrase→model/obfus/unicode/base64), indirect-artifact, secret in+out, PII, tool-output scan, strict fail-closed, recall 0.85, FP 0% | multilingual VI/JA & split-injection [planned]; sandbox scaffold; 1 model | -| H5 | Governance | 25 | 76 | **80** ⬆ | Good(đỉnh) | hash-chain + RSA head, SoD, least-privilege, rate-limit, secrets-scan, no-bypass, tool-audit, telemetry-integrity **+ approval-identity ký-danh-tính + KMS live (rotate/non-exportable) + WORM audit ngoài (gap/tamper)** | **IdP live** (registry pubkey tĩnh, chưa OIDC/JWT); **WORM store thật** (ledger local, chưa S3-Object-Lock); KMS chưa mặc định (fallback khoá local) | +| H5 | Governance | 25 | 76 | **80** ⬆ | Good(đỉnh) | hash-chain + RSA head, SoD, least-privilege, rate-limit, secrets-scan, no-bypass, tool-audit, telemetry-integrity **+ approval-identity ký-danh-tính + mock IdP/OIDC JWT + KMS live (rotate/non-exportable) + WORM audit ngoài (gap/tamper)** | **IdP live/JWKS thật**; **WORM store thật** (ledger local, chưa S3-Object-Lock); KMS chưa mặc định (fallback khoá local) | | H6 | AgentOps | 30 | 79 | **80** ⬆ | Good(đỉnh) | cost-spike 4 chế độ (rel+abs+cumulative+cold-start), drift, hallucination-rate, telemetry token thật + ký toàn vẹn **+ alerting LIVE (webhook + dedup + dead-letter, end-to-end từ step fail) + provider-telemetry API (fetch + đối soát bắt under-reporting) + dashboard hosted (/healthz stale-aware) + window circuit-breaker (V15)** | dashboard host thật (deploy nginx/container + auth); kênh alert managed (Slack/PagerDuty + on-call, C7 incident); billing-API thật (OpenAI/Anthropic, cần key) | | H7 | Orchestration | 80 | 80 | **80** | Good(đỉnh) | Boss DAG, BACK-TO-PLAN/retry, rollback thật, model-fallback thật, drift | chưa transaction-rollback xuyên nhiều step | @@ -52,7 +55,7 @@ - (5/5 gate)×100 chỉ đo **độ phủ control**, không đo **độ trưởng thành/vận-hành-thật** — report này tách bạch: mục 2 = coverage/pass thật, mục 3 = trưởng thành công tâm. ## 5. Ranh giới trung thực -- CASAN **Level 4 chứng minh bằng tấn công (175 test)**. Level 5 các control hiện thực + test cục bộ; production Level 5 cần **IdP thật**, **WORM store (S3 Object Lock)**, KMS mặc định + HSM, **billing-API thật** (usage endpoint OpenAI/Anthropic), **dashboard deploy thật** (nginx/container + auth), **kênh alert managed + on-call (C7)**, sandbox isolation thật. +- CASAN **Level 4 chứng minh bằng tấn công (218 core harness tests)**. Level 5 các control hiện thực + test cục bộ; production Level 5 cần **IdP thật/JWKS**, **WORM store (S3 Object Lock)**, KMS mặc định + HSM, **billing-API thật** (usage endpoint OpenAI/Anthropic), **dashboard deploy thật** (nginx/container + auth), **kênh alert managed + on-call**, sandbox isolation rootless/nsjail/base image CI. - KMS đã chạy **live qua Vault dev** ở lần chấm 2026-07-04 (đường Transit thật, khoá non-exportable); lần chạy 2026-07-05 Vault down → suite KMS **SKIP đúng thiết kế** (không tính là fail). Production thay bằng Vault/AWS-KMS/CloudHSM. - H6 "live" nghĩa là: webhook, provider-usage API, dashboard `/healthz` đều là **endpoint HTTP thật chạy local** (cùng chuẩn Vault-dev) — chưa phải dịch vụ hosted/managed bên ngoài. - Model = Ollama ornith:9b **local**; đường cloud (OpenAI/Anthropic) đã hiện thực trong `model-call.py` nhưng **chưa test bằng key thật**. @@ -66,12 +69,18 @@ bash .specify/tests/adversarial-harness-tests.sh # 44/0 bash .specify/tests/phase1-track-a-tests.sh # 25/0 bash .specify/tests/phase2-track-c-tests.sh # 29/0 bash .specify/tests/phase3-evidence-pack-tests.sh # 7/0 -bash .specify/tests/phase-h5-approval-tests.sh # 8/0 +bash .specify/tests/phase-h5-approval-tests.sh # 12/0 # KMS live: bật Vault dev trước để phần KMS chạy thật (không SKIP) docker run -d -p 8200:8200 -e VAULT_DEV_ROOT_TOKEN_ID=root hashicorp/vault VAULT_ADDR=http://127.0.0.1:8200 VAULT_TOKEN=root \ bash .specify/tests/phase-h5-infra-tests.sh # 7/0 (KMS live + WORM) bash .specify/tests/phase-h6-agentops-tests.sh # 20/0 (alerting live + provider-API + hosted dashboard + V15) +bash .specify/tests/phase-c7-incident-tests.sh # 15/0 +bash .specify/tests/phase-h4-multilingual-tests.sh # 7/0 +bash .specify/tests/phase-c6-sandbox-tests.sh # 6/0 +bash .specify/tests/phase-h4-split-inject-tests.sh # 8/0 +bash .specify/tests/phase10-traceability-tests.sh # 3/0 +bash .specify/tests/phase3-model-router-tests.sh # 10/0 (direct model-router/digest suite) bash .specify/scripts/bash/security-gate.sh # verdict PASS=11 FAIL=0 ``` > Mục 3 là **đánh giá trưởng thành theo rubric** (người chấm, neo vào bằng chứng + gap thật), không phải output tự động của scorecard.sh (vốn chỉ đo coverage). Suite model cần Ollama live; suite KMS cần Vault live để chạy (không có thì SKIP, không tính là fail). Suite H6 **tự dựng** webhook sink / mock provider-API / dashboard server trên cổng ephemeral local — deterministic, không cần model. diff --git a/AINative_OKR_CASAN5/.specify/scripts/bash/approval-jwt-mint.py b/AINative_OKR_CASAN5/.specify/scripts/bash/approval-jwt-mint.py new file mode 100755 index 0000000..88e309e --- /dev/null +++ b/AINative_OKR_CASAN5/.specify/scripts/bash/approval-jwt-mint.py @@ -0,0 +1,69 @@ +#!/usr/bin/env python3 +"""Mint a mock IdP RS256 approval JWT for CASAN tests/dev. + +The token is bound to the same high-risk request that governance-check verifies: +sub=, role=, action, actor, and input_sha256. +""" +import argparse +import base64 +import hashlib +import json +import os +import subprocess +import sys +import tempfile +import time + + +def b64u(data: bytes) -> str: + return base64.urlsafe_b64encode(data).decode().rstrip("=") + + +def main() -> int: + ap = argparse.ArgumentParser() + ap.add_argument("--key", required=True) + ap.add_argument("--sub", required=True) + ap.add_argument("--role", required=True) + ap.add_argument("--action", required=True) + ap.add_argument("--actor", required=True) + ap.add_argument("--input", required=True) + ap.add_argument("--exp-offset", type=int, default=300) + args = ap.parse_args() + + with open(args.input, "rb") as f: + input_sha = hashlib.sha256(f.read()).hexdigest() + now = int(time.time()) + header = {"alg": "RS256", "typ": "JWT"} + claims = { + "iss": "casan-mock-idp", + "sub": args.sub, + "role": args.role, + "action": args.action, + "actor": args.actor, + "input_sha256": input_sha, + "iat": now, + "exp": now + args.exp_offset, + } + signing_input = ".".join([ + b64u(json.dumps(header, separators=(",", ":"), sort_keys=True).encode()), + b64u(json.dumps(claims, separators=(",", ":"), sort_keys=True).encode()), + ]) + with tempfile.TemporaryDirectory() as td: + msg = os.path.join(td, "msg.txt") + sig = os.path.join(td, "sig.bin") + open(msg, "wb").write(signing_input.encode()) + rc = subprocess.run( + ["openssl", "dgst", "-sha256", "-sign", args.key, "-out", sig, msg], + stdout=subprocess.DEVNULL, + stderr=subprocess.DEVNULL, + ).returncode + if rc != 0: + print("approval-jwt-mint: signing failed", file=sys.stderr) + return 1 + token = signing_input + "." + b64u(open(sig, "rb").read()) + print(token) + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/AINative_OKR_CASAN5/.specify/scripts/bash/approval-verify.sh b/AINative_OKR_CASAN5/.specify/scripts/bash/approval-verify.sh index b008555..9e4eee0 100755 --- a/AINative_OKR_CASAN5/.specify/scripts/bash/approval-verify.sh +++ b/AINative_OKR_CASAN5/.specify/scripts/bash/approval-verify.sh @@ -1,7 +1,7 @@ #!/usr/bin/env bash set -uo pipefail -# CASAN H5 — Signed-approval verifier (Approval-identity MVP · C4 / V20). +# CASAN H5 — Signed/JWT approval verifier (Approval-identity MVP · C4 / V20). # # Problem: high-risk approval used to trust a plain env var (CASAN_APPROVER=bob) — # anyone who can set the env can "approve". This binds an approval to a REGISTERED @@ -13,11 +13,14 @@ set -uo pipefail # # Usage: # approval-verify.sh +# CASAN_APPROVAL_JWT= approval-verify.sh - # Registry (line format, no yaml dep): # reviewer # action # Env: CASAN_REVIEWERS_FILE (default governance/reviewers.registry) # CASAN_REVIEWERS_DIR (default governance/reviewers) — base dir for pubkey-file +# CASAN_APPROVAL_JWT (optional RS256 IdP token) +# CASAN_IDP_PUBLIC_KEY (default central-governance/idp-public.pem) # Exit: 0 ok (prints "APPROVAL_OK role="), 3 deny (reason on stderr), 64 usage. SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" @@ -27,7 +30,7 @@ REVIEWERS_FILE="${CASAN_REVIEWERS_FILE:-$GOV_DIR/reviewers.registry}" REVIEWERS_DIR="${CASAN_REVIEWERS_DIR:-$GOV_DIR/reviewers}" ACTION="${1:-}"; ACTOR="${2:-}"; INPUT_FILE="${3:-}"; APPROVER="${4:-}"; SIG_FILE="${5:-}" -if [[ -z "$ACTION" || -z "$ACTOR" || -z "$INPUT_FILE" || -z "$APPROVER" || -z "$SIG_FILE" ]]; then +if [[ -z "$ACTION" || -z "$ACTOR" || -z "$INPUT_FILE" || -z "$APPROVER" ]]; then echo "Usage: approval-verify.sh " >&2 exit 64 fi @@ -35,7 +38,6 @@ fi deny() { echo "APPROVAL_DENIED reason=$1 approver=$APPROVER action=$ACTION" >&2; exit 3; } [[ -f "$INPUT_FILE" ]] || deny "input_file_missing" -[[ -f "$SIG_FILE" ]] || deny "signature_missing" [[ -f "$REVIEWERS_FILE" ]] || deny "reviewer_registry_missing" command -v openssl >/dev/null 2>&1 || deny "openssl_unavailable" @@ -44,15 +46,89 @@ hash_file() { else shasum -a 256 "$1" | awk '{print $1}'; fi } +INPUT_SHA="$(hash_file "$INPUT_FILE")" + +# Role authorization for this action (fallback to the "default" action policy). +ROLES="$(awk -v a="$ACTION" '$1=="action" && $2==a {print $3; exit}' "$REVIEWERS_FILE")" +[[ -n "$ROLES" ]] || ROLES="$(awk '$1=="action" && $2=="default" {print $3; exit}' "$REVIEWERS_FILE")" + +if [[ -n "${CASAN_APPROVAL_JWT:-}" ]]; then + IDP_PUB="${CASAN_IDP_PUBLIC_KEY:-$GOV_DIR/idp-public.pem}" + [[ -f "$IDP_PUB" ]] || deny "idp_pubkey_missing($IDP_PUB)" + JWT_OUT="$( + python3 - "$CASAN_APPROVAL_JWT" "$IDP_PUB" "$APPROVER" "$ROLES" "$ACTION" "$ACTOR" "$INPUT_SHA" <<'PY' +import base64 +import json +import os +import subprocess +import sys +import tempfile +import time + +jwt, pub, approver, roles_csv, action, actor, input_sha = sys.argv[1:8] + +def die(reason): + print(reason, file=sys.stderr) + sys.exit(3) + +def b64u_decode(part): + try: + return base64.urlsafe_b64decode(part + "=" * (-len(part) % 4)) + except Exception: + die("jwt_base64_invalid") + +parts = jwt.split(".") +if len(parts) != 3: + die("jwt_shape_invalid") +header = json.loads(b64u_decode(parts[0])) +claims = json.loads(b64u_decode(parts[1])) +if header.get("alg") != "RS256": + die("jwt_alg_not_allowed") +if int(claims.get("exp", 0)) <= int(time.time()): + die("jwt_expired") +if claims.get("sub") != approver: + die("jwt_sub_mismatch") +role = claims.get("role", "") +roles = [r for r in roles_csv.split(",") if r] +if role not in roles: + die(f"jwt_role_not_authorized(role={role} allowed={roles_csv})") +if claims.get("action") != action: + die("jwt_action_mismatch") +if claims.get("actor") != actor: + die("jwt_actor_mismatch") +if claims.get("input_sha256") != input_sha: + die("jwt_input_hash_mismatch") + +sig = b64u_decode(parts[2]) +signing_input = ".".join(parts[:2]).encode() +with tempfile.TemporaryDirectory() as td: + sig_path = os.path.join(td, "sig.bin") + msg_path = os.path.join(td, "msg.txt") + open(sig_path, "wb").write(sig) + open(msg_path, "wb").write(signing_input) + rc = subprocess.run( + ["openssl", "dgst", "-sha256", "-verify", pub, "-signature", sig_path, msg_path], + stdout=subprocess.DEVNULL, + stderr=subprocess.DEVNULL, + ).returncode +if rc != 0: + die("jwt_signature_invalid") +print(f"OK role={role}") +PY + )" || deny "oidc_jwt_invalid(${JWT_OUT:-see_stderr})" + ROLE="${JWT_OUT#OK role=}" + echo "APPROVAL_OK mechanism=oidc role=$ROLE approver=$APPROVER action=$ACTION" + exit 0 +fi + +[[ -n "$SIG_FILE" && -f "$SIG_FILE" ]] || deny "signature_missing" + # Reviewer lookup (first matching registered reviewer). REV_LINE="$(awk -v id="$APPROVER" '$1=="reviewer" && $2==id {print $3" "$4; exit}' "$REVIEWERS_FILE")" [[ -n "$REV_LINE" ]] || deny "approver_not_registered" ROLE="${REV_LINE%% *}" PUB_REL="${REV_LINE##* }" -# Role authorization for this action (fallback to the "default" action policy). -ROLES="$(awk -v a="$ACTION" '$1=="action" && $2==a {print $3; exit}' "$REVIEWERS_FILE")" -[[ -n "$ROLES" ]] || ROLES="$(awk '$1=="action" && $2=="default" {print $3; exit}' "$REVIEWERS_FILE")" case ",$ROLES," in *",$ROLE,"*) : ;; *) deny "approver_role_not_authorized(role=$ROLE action=$ACTION allowed=$ROLES)" ;; @@ -63,7 +139,6 @@ PUB="$PUB_REL"; [[ "$PUB" = /* ]] || PUB="$REVIEWERS_DIR/$PUB_REL" [[ -f "$PUB" ]] || deny "approver_pubkey_missing($PUB)" # Rebuild the exact signed assertion and verify. -INPUT_SHA="$(hash_file "$INPUT_FILE")" MSG="casan-approval|v1|$ACTION|$ACTOR|$INPUT_SHA|$APPROVER" TMP="$(mktemp)"; trap 'rm -f "$TMP"' EXIT printf '%s' "$MSG" > "$TMP" diff --git a/AINative_OKR_CASAN5/.specify/scripts/bash/evidence-pack-build.py b/AINative_OKR_CASAN5/.specify/scripts/bash/evidence-pack-build.py index 7c08c61..595f405 100755 --- a/AINative_OKR_CASAN5/.specify/scripts/bash/evidence-pack-build.py +++ b/AINative_OKR_CASAN5/.specify/scripts/bash/evidence-pack-build.py @@ -18,6 +18,7 @@ skipped (missing evidence => not certified, with the reason recorded). import hashlib import json import os +import subprocess import sys @@ -75,6 +76,25 @@ def main(): "judge_gate_tests": os.path.exists(os.path.join(root, ".specify/tests/phase3-judge-gate-tests.sh")), "note": "judge-gate fail-before/fix cycle proven by phase3-judge-gate-tests.sh", } + traceability_out = os.path.join(pack_dir, "traceability-matrix.json") + traceability_script = os.path.join(root, ".specify/scripts/bash/traceability-matrix.py") + traceability_rc = 1 + if os.path.isfile(traceability_script): + traceability_rc = subprocess.run( + [ + sys.executable, + traceability_script, + "--requirements", + os.path.join(root, "docs/input/okr-requirement.md"), + "--map", + os.path.join(root, ".specify/traceability-map.json"), + "--out", + traceability_out, + "--gate", + ], + stdout=subprocess.DEVNULL, + stderr=subprocess.DEVNULL, + ).returncode # H4 security sec_rows = read_jsonl(os.path.join(logs, "audit", "security.jsonl")) @@ -151,6 +171,8 @@ def main(): reasons.append("unresolved_cost_spike") if len(sec_rows) == 0: reasons.append("h4_security_not_exercised") + if traceability_rc != 0: + reasons.append("traceability_gate_failed") fp_ok = bool(fp) and fp.get("within_budget") is True if fp is None: reasons.append("benign_fp_report_missing(gate_skipped)") @@ -162,7 +184,7 @@ def main(): run_summary = { "run_id": run_id, "pack_version": "1.0-mvp", "certified": certified, "certification_reasons": reasons or ["all_required_gates_passed"], - "required_gates": ["H4-security", "H5-audit-chain", "H5-telemetry", "H6-cost", "benign-fp-budget"], + "required_gates": ["H3-traceability", "H4-security", "H5-audit-chain", "H5-telemetry", "H6-cost", "benign-fp-budget"], "harness_reports": sorted(reports.keys()), } with open(os.path.join(pack_dir, "run-summary.json"), "w", encoding="utf-8") as f: @@ -182,6 +204,7 @@ def main(): f"{total_tokens} provider tokens", f"- H2 tool audit: {reports['h2-tool-audit.json']['records']} records, " f"chain_ok={reports['h2-tool-audit.json']['chain_ok']}", + f"- H3 traceability: ok={traceability_rc == 0}", f"- Red-team: {reports['redteam-result.json']['vectors_defined']} vectors " f"(block_rate={reports['redteam-result.json']['adversarial_block_rate_pct']}%)", "", diff --git a/AINative_OKR_CASAN5/.specify/scripts/bash/governance-check.sh b/AINative_OKR_CASAN5/.specify/scripts/bash/governance-check.sh index bc8ce6d..a80486c 100755 --- a/AINative_OKR_CASAN5/.specify/scripts/bash/governance-check.sh +++ b/AINative_OKR_CASAN5/.specify/scripts/bash/governance-check.sh @@ -92,16 +92,20 @@ if [[ "$RISK_LEVEL" == "high" ]]; then # Approval-identity mode (V20): an env-var approver is NOT enough — the # reviewer must cryptographically SIGN this exact request and their role must # be authorized for the action. SoD (actor != approver) still enforced. - if [[ "$APPROVAL_DECISION" == "approve" && -n "$APPROVER" && -n "${CASAN_APPROVAL_SIG:-}" ]]; then + if [[ "$APPROVAL_DECISION" == "approve" && -n "$APPROVER" && ( -n "${CASAN_APPROVAL_SIG:-}" || -n "${CASAN_APPROVAL_JWT:-}" ) ]]; then if [[ "$APPROVER" == "$ACTOR" ]]; then APPROVAL_STATUS="separation_of_duties_violation" DECISION="denied" REASONS+=("separation-of-duties:actor-equals-approver") else AV_RC=0 - AV_OUT="$(bash "$SCRIPT_DIR/approval-verify.sh" "$ACTION_NAME" "$ACTOR" "$INPUT_FILE" "$APPROVER" "$CASAN_APPROVAL_SIG" 2>/dev/null)" || AV_RC=$? + AV_OUT="$(bash "$SCRIPT_DIR/approval-verify.sh" "$ACTION_NAME" "$ACTOR" "$INPUT_FILE" "$APPROVER" "${CASAN_APPROVAL_SIG:-"-"}" 2>/dev/null)" || AV_RC=$? if [[ "$AV_RC" -eq 0 ]]; then - APPROVAL_STATUS="human_approved_signed" + if printf '%s' "$AV_OUT" | grep -q "mechanism=oidc"; then + APPROVAL_STATUS="human_approved_oidc" + else + APPROVAL_STATUS="human_approved_signed" + fi DECISION="approved" REASONS+=("signed-approval:${AV_OUT#APPROVAL_OK }") else diff --git a/AINative_OKR_CASAN5/.specify/scripts/bash/model-call.py b/AINative_OKR_CASAN5/.specify/scripts/bash/model-call.py index 3655e9f..fccc9e1 100755 --- a/AINative_OKR_CASAN5/.specify/scripts/bash/model-call.py +++ b/AINative_OKR_CASAN5/.specify/scripts/bash/model-call.py @@ -23,6 +23,7 @@ Usage: import argparse import json import os +import subprocess import sys import time import urllib.request @@ -89,6 +90,17 @@ def call_ollama(model_name, prompt, role): host = os.environ.get("CASAN_OLLAMA_HOST", OLLAMA_HOST) if host != OLLAMA_HOST: fail(f"endpoint_not_allowed ollama host={host} (only {OLLAMA_HOST})") + digest_gate = os.path.join(os.path.dirname(__file__), "model-digest-check.sh") + if os.path.isfile(digest_gate): + check = subprocess.run( + ["bash", digest_gate, "verify", model_name], + text=True, + stdout=subprocess.PIPE, + stderr=subprocess.PIPE, + ) + if check.returncode != 0: + msg = (check.stderr or check.stdout or "model_digest_check_failed").strip() + fail(msg) url = f"http://{host}/api/generate" body = { "model": model_name, diff --git a/AINative_OKR_CASAN5/.specify/scripts/bash/model-digest-check.sh b/AINative_OKR_CASAN5/.specify/scripts/bash/model-digest-check.sh index 4d2fdb5..7014d5a 100644 --- a/AINative_OKR_CASAN5/.specify/scripts/bash/model-digest-check.sh +++ b/AINative_OKR_CASAN5/.specify/scripts/bash/model-digest-check.sh @@ -16,8 +16,9 @@ set -uo pipefail # model-digest-check.sh verify [model] # compare live digest to the pinned one # model-digest-check.sh show [model] # Env: CASAN_MODEL (default ornith:9b) · CASAN_MODEL_DIGEST (override) · -# CASAN_MODEL_DIGEST_PIN (pin file) · OLLAMA_HOST (default 127.0.0.1:11434) -# Exit: 0 match/pinned · 2 MISMATCH (swap detected) · 3 unpinned/undeterminable. +# CASAN_MODEL_DIGEST_PIN (pin file) · CASAN_MODEL_DIGEST_MODE=block|warn · +# OLLAMA_HOST (default 127.0.0.1:11434) +# Exit: 0 match/pinned/warned · 2 MISMATCH in block mode · 3 unpinned/undeterminable. SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" PROJECT_ROOT="$(cd "$SCRIPT_DIR/../../.." && pwd)" @@ -60,6 +61,10 @@ case "$CMD" in echo "MODEL_DIGEST_OK model=$MODEL digest=${D:0:24}…" exit 0 fi + if [[ "${CASAN_MODEL_DIGEST_MODE:-block}" == "warn" ]]; then + echo "MODEL_DIGEST_WARN model=$MODEL pinned=${P:0:24}… live=${D:0:24}… — model may have been swapped/poisoned" >&2 + exit 0 + fi echo "MODEL_DIGEST_MISMATCH model=$MODEL pinned=${P:0:24}… live=${D:0:24}… — model may have been swapped/poisoned" >&2 exit 2 ;; diff --git a/AINative_OKR_CASAN5/.specify/scripts/bash/traceability-matrix.py b/AINative_OKR_CASAN5/.specify/scripts/bash/traceability-matrix.py new file mode 100755 index 0000000..ae1d31e --- /dev/null +++ b/AINative_OKR_CASAN5/.specify/scripts/bash/traceability-matrix.py @@ -0,0 +1,115 @@ +#!/usr/bin/env python3 +"""CASAN Plan-10 traceability matrix generator/gate. + +Parses FR-* requirements from docs/input/okr-requirement.md and checks each +requirement has at least one existing code file and one existing test file. +""" +import argparse +import json +import os +import re +import sys +from datetime import datetime, timezone + + +FR_RE = re.compile(r"\|\s*(FR-\d+)\s*\|\s*([^|]+?)\s*\|") + + +def project_root() -> str: + return os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..", "..")) + + +def parse_requirements(path: str): + seen = {} + with open(path, encoding="utf-8") as f: + for line in f: + match = FR_RE.search(line) + if not match: + continue + fr_id, name = match.groups() + if fr_id not in seen: + seen[fr_id] = {"id": fr_id, "name": " ".join(name.split())} + return [seen[k] for k in sorted(seen)] + + +def existing_files(root: str, values): + present, missing = [], [] + for rel in values or []: + if os.path.isfile(os.path.join(root, rel)): + present.append(rel) + else: + missing.append(rel) + return present, missing + + +def main() -> int: + root = project_root() + ap = argparse.ArgumentParser() + ap.add_argument("--requirements", default=os.path.join(root, "docs/input/okr-requirement.md")) + ap.add_argument("--map", default=os.path.join(root, ".specify/traceability-map.json")) + ap.add_argument("--out", default=os.path.join(root, "docs/output/casan/traceability-matrix.json")) + ap.add_argument("--gate", action="store_true") + args = ap.parse_args() + + reqs = parse_requirements(args.requirements) + with open(args.map, encoding="utf-8") as f: + mapping = json.load(f) + + rows = [] + failures = [] + for req in reqs: + entry = mapping.get(req["id"], {}) + code, missing_code = existing_files(root, entry.get("code", [])) + tests, missing_tests = existing_files(root, entry.get("tests", [])) + status = "PASS" if code and tests and not missing_code and not missing_tests else "FAIL" + row = { + "id": req["id"], + "name": req["name"], + "status": status, + "code": code, + "tests": tests, + "missing_code": missing_code, + "missing_tests": missing_tests, + } + rows.append(row) + if status != "PASS": + failures.append(row) + + orphan_mappings = sorted(set(mapping) - {r["id"] for r in reqs}) + out = { + "generated_at": datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"), + "requirements_source": os.path.relpath(args.requirements, root), + "mapping_source": os.path.relpath(args.map, root), + "summary": { + "requirements": len(reqs), + "passed": sum(1 for r in rows if r["status"] == "PASS"), + "failed": len(failures), + "orphan_mappings": orphan_mappings, + }, + "matrix": rows, + } + os.makedirs(os.path.dirname(args.out), exist_ok=True) + with open(args.out, "w", encoding="utf-8") as f: + json.dump(out, f, indent=2, ensure_ascii=False) + f.write("\n") + + if failures: + for row in failures: + print( + f"TRACEABILITY_FAIL {row['id']} code={len(row['code'])} tests={len(row['tests'])} " + f"missing_code={len(row['missing_code'])} missing_tests={len(row['missing_tests'])}", + file=sys.stderr, + ) + if orphan_mappings: + print(f"TRACEABILITY_WARN orphan_mappings={','.join(orphan_mappings)}", file=sys.stderr) + print( + f"TRACEABILITY_MATRIX requirements={len(reqs)} pass={out['summary']['passed']} " + f"fail={len(failures)} out={args.out}" + ) + if args.gate and failures: + return 1 + return 0 + + +if __name__ == "__main__": + raise SystemExit(main()) diff --git a/AINative_OKR_CASAN5/.specify/security/model-digest.pin b/AINative_OKR_CASAN5/.specify/security/model-digest.pin new file mode 100644 index 0000000..641d645 --- /dev/null +++ b/AINative_OKR_CASAN5/.specify/security/model-digest.pin @@ -0,0 +1 @@ +ornith:9b a75697c145891910e312c95e4a9fc1ccb8653e5ef543b23b0403a4665b82fd91 diff --git a/AINative_OKR_CASAN5/.specify/tests/generate-agentops-dashboard.py b/AINative_OKR_CASAN5/.specify/tests/generate-agentops-dashboard.py index dfd6ca2..89c0946 100755 --- a/AINative_OKR_CASAN5/.specify/tests/generate-agentops-dashboard.py +++ b/AINative_OKR_CASAN5/.specify/tests/generate-agentops-dashboard.py @@ -114,7 +114,7 @@ tr:nth-child(even) td {{ background: #f7f9fc; }} CASAN Level 4 — chứng minh bằng tấn công Average {avg_score}/100 Harness thấp nhất {lowest_score} - 175 test · 0 fail + 218 core tests · 0 fail Recall model 0.85 > regex 0.00 @@ -132,7 +132,7 @@ tr:nth-child(even) td {{ background: #f7f9fc; }}

Bảo mật & Governance đã kiểm chứng (test đối kháng thật)

- Kiểm thử đối kháng 175 / 0 fail + Kiểm thử đối kháng 218 / 0 fail H4 recall model 0.85 > regex 0.00 Benign FP 0.00% · block 100.00% Audit hash-chain + ký KMS (rotate/non-exportable) diff --git a/AINative_OKR_CASAN5/.specify/tests/phase-h5-approval-tests.sh b/AINative_OKR_CASAN5/.specify/tests/phase-h5-approval-tests.sh index 68891ac..abf1823 100755 --- a/AINative_OKR_CASAN5/.specify/tests/phase-h5-approval-tests.sh +++ b/AINative_OKR_CASAN5/.specify/tests/phase-h5-approval-tests.sh @@ -94,6 +94,38 @@ expect_rc 0 "non-strict env approval unchanged (backward compatible)" \ env CASAN_ACTOR=alice CASAN_APPROVAL_DECISION=approve CASAN_APPROVER=bob \ bash "$S/governance-check.sh" "$REQ" "$WORK/out.txt" deploy +echo "===== H5 OIDC approval via mock IdP JWT (CASAN_APPROVAL_STRICT=1) =====" +openssl genrsa -out "$WORK/idp.priv.pem" 2048 2>/dev/null +openssl rsa -in "$WORK/idp.priv.pem" -pubout -out "$WORK/idp.pub.pem" 2>/dev/null +openssl genrsa -out "$WORK/fake-idp.priv.pem" 2048 2>/dev/null + +mint_jwt() { + python3 "$S/approval-jwt-mint.py" --key "$1" --sub "$2" --role "$3" \ + --action deploy --actor alice --input "$REQ" --exp-offset "$4" +} + +# 9. Valid IdP-signed JWT by an authorized role -> APPROVED +OIDC_JWT="$(mint_jwt "$WORK/idp.priv.pem" oidc-ops ops 300)" +OUT="$(gc CASAN_APPROVER=oidc-ops CASAN_APPROVAL_JWT="$OIDC_JWT" CASAN_IDP_PUBLIC_KEY="$WORK/idp.pub.pem" 2>/dev/null)"; RC=$? +{ [[ "$RC" -eq 0 ]] && printf '%s' "$OUT" | grep -q "human_approved_oidc"; } \ + && pass "valid IdP JWT by authorized role -> APPROVED" \ + || fail "valid IdP JWT rejected (rc=$RC out=$OUT)" + +# 10. Expired JWT -> DENY +EXPIRED_JWT="$(mint_jwt "$WORK/idp.priv.pem" oidc-ops ops -60)" +expect_rc 2 "expired IdP JWT is denied" \ + gc CASAN_APPROVER=oidc-ops CASAN_APPROVAL_JWT="$EXPIRED_JWT" CASAN_IDP_PUBLIC_KEY="$WORK/idp.pub.pem" + +# 11. Role not authorized for deploy -> DENY +WRONG_ROLE_JWT="$(mint_jwt "$WORK/idp.priv.pem" oidc-tech tech_lead 300)" +expect_rc 2 "IdP JWT with unauthorized role is denied" \ + gc CASAN_APPROVER=oidc-tech CASAN_APPROVAL_JWT="$WRONG_ROLE_JWT" CASAN_IDP_PUBLIC_KEY="$WORK/idp.pub.pem" + +# 12. JWT signed by a different key than the trusted IdP pubkey -> DENY +FORGED_JWT="$(mint_jwt "$WORK/fake-idp.priv.pem" oidc-ops ops 300)" +expect_rc 2 "JWT with forged IdP signature is denied" \ + gc CASAN_APPROVER=oidc-ops CASAN_APPROVAL_JWT="$FORGED_JWT" CASAN_IDP_PUBLIC_KEY="$WORK/idp.pub.pem" + echo "" echo "===== H5 APPROVAL-IDENTITY SUMMARY: PASS=$PASS FAIL=$FAIL =====" [[ "$FAIL" -eq 0 ]] || exit 1 diff --git a/AINative_OKR_CASAN5/.specify/tests/phase10-traceability-tests.sh b/AINative_OKR_CASAN5/.specify/tests/phase10-traceability-tests.sh new file mode 100755 index 0000000..e871064 --- /dev/null +++ b/AINative_OKR_CASAN5/.specify/tests/phase10-traceability-tests.sh @@ -0,0 +1,54 @@ +#!/usr/bin/env bash +set -uo pipefail + +# CASAN Plan-10 — Traceability REQ→code→test MVP tests. +# +# Proves the matrix is generated from the real requirement document and the gate +# fails when any FR loses test coverage. + +SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" +PROJECT_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)" +S="$PROJECT_ROOT/.specify/scripts/bash" +WORK="$(mktemp -d)" +trap 'rm -rf "$WORK"' EXIT + +PASS=0; FAIL=0 +pass() { echo "PASS: $1"; PASS=$((PASS + 1)); } +fail() { echo "FAIL: $1"; FAIL=$((FAIL + 1)); } + +REQ="$PROJECT_ROOT/docs/input/okr-requirement.md" +MAP="$PROJECT_ROOT/.specify/traceability-map.json" +OUT="$WORK/traceability-matrix.json" + +echo "===== Plan-10 traceability matrix =====" +if python3 "$S/traceability-matrix.py" --requirements "$REQ" --map "$MAP" --out "$OUT" --gate >/dev/null 2>"$WORK/pass.err"; then + pass "complete FR→code→test matrix passes gate" +else + cat "$WORK/pass.err" + fail "complete traceability matrix rejected" +fi + +COUNT="$(python3 -c "import json;d=json.load(open('$OUT'));print(d['summary']['requirements'])" 2>/dev/null)" +FAILED="$(python3 -c "import json;d=json.load(open('$OUT'));print(d['summary']['failed'])" 2>/dev/null)" +[[ "$COUNT" == "5" && "$FAILED" == "0" ]] \ + && pass "matrix captures all 5 FRs with zero failures" \ + || fail "matrix summary unexpected (requirements=$COUNT failed=$FAILED)" + +BROKEN="$WORK/traceability-map-broken.json" +python3 - "$MAP" "$BROKEN" <<'PY' +import json, sys +d = json.load(open(sys.argv[1])) +d["FR-04"]["tests"] = [] +json.dump(d, open(sys.argv[2], "w"), indent=2) +PY +set +e +python3 "$S/traceability-matrix.py" --requirements "$REQ" --map "$BROKEN" --out "$WORK/broken.json" --gate >/dev/null 2>"$WORK/broken.err" +RC=$? +set -e 2>/dev/null || true +[[ "$RC" -eq 1 ]] && grep -q "TRACEABILITY_FAIL FR-04" "$WORK/broken.err" \ + && pass "missing FR test coverage fails the gate" \ + || fail "missing test coverage did not fail as expected (rc=$RC)" + +echo "" +echo "===== TRACEABILITY SUMMARY: PASS=$PASS FAIL=$FAIL =====" +[[ "$FAIL" -eq 0 ]] || exit 1 diff --git a/AINative_OKR_CASAN5/.specify/tests/phase3-evidence-pack-tests.sh b/AINative_OKR_CASAN5/.specify/tests/phase3-evidence-pack-tests.sh index 5071e4c..9974dc8 100755 --- a/AINative_OKR_CASAN5/.specify/tests/phase3-evidence-pack-tests.sh +++ b/AINative_OKR_CASAN5/.specify/tests/phase3-evidence-pack-tests.sh @@ -41,10 +41,10 @@ MISSING=0 for f in run-summary.json h1-context-report.json h2-tool-audit.json h3-eval-scorecard.json \ h4-security-report.json h5-audit-chain-proof.json h6-cost-telemetry.json \ h7-orchestration-report.json redteam-result.json benign-fp-report.json \ - artifact-manifest.json decision-log.md; do + artifact-manifest.json traceability-matrix.json decision-log.md; do [[ -f "$PACKDIR/$f" ]] || { echo " missing $f"; MISSING=$((MISSING+1)); } done -[[ "$MISSING" -eq 0 ]] && pass "pack contains all 12 standard evidence files" || fail "pack missing $MISSING files" +[[ "$MISSING" -eq 0 ]] && pass "pack contains all 13 standard evidence files" || fail "pack missing $MISSING files" [[ -f "$PACKDIR/evidence-pack.sig" ]] && pass "pack is signed (evidence-pack.sig present)" || fail "pack signature missing" rc=0; bash "$EP" verify-pack "$RID" >/dev/null 2>&1 || rc=$? diff --git a/AINative_OKR_CASAN5/.specify/tests/phase3-model-router-tests.sh b/AINative_OKR_CASAN5/.specify/tests/phase3-model-router-tests.sh index 2f2bf8b..b6ad455 100755 --- a/AINative_OKR_CASAN5/.specify/tests/phase3-model-router-tests.sh +++ b/AINative_OKR_CASAN5/.specify/tests/phase3-model-router-tests.sh @@ -66,7 +66,28 @@ else pass "no API-key/secret pattern in .specify/logs" fi -# 6: cloud backend reports unavailable honestly while keys are unset. +# 6: model digest pinning catches a silent model swap, with warn mode available. +PIN="$WORK/model-digest.pin" +CASAN_MODEL_DIGEST_PIN="$PIN" CASAN_MODEL_DIGEST="digest-approved" bash "$SCRIPTS/model-digest-check.sh" pin ornith:9b >/dev/null 2>&1 +if CASAN_MODEL_DIGEST_PIN="$PIN" CASAN_MODEL_DIGEST="digest-approved" bash "$SCRIPTS/model-digest-check.sh" verify ornith:9b >/dev/null 2>&1; then + pass "model digest pin verifies approved digest" +else + fail "model digest pin rejected matching digest" +fi +set +e +CASAN_MODEL_DIGEST_PIN="$PIN" CASAN_MODEL_DIGEST="digest-swapped" bash "$SCRIPTS/model-digest-check.sh" verify ornith:9b >/dev/null 2>"$WORK/digest.err" +DRC=$? +CASAN_MODEL_DIGEST_PIN="$PIN" CASAN_MODEL_DIGEST="digest-swapped" CASAN_MODEL_DIGEST_MODE=warn bash "$SCRIPTS/model-digest-check.sh" verify ornith:9b >/dev/null 2>"$WORK/digest-warn.err" +WRC=$? +set -e 2>/dev/null || true +[[ "$DRC" -eq 2 ]] && grep -q "MODEL_DIGEST_MISMATCH" "$WORK/digest.err" \ + && pass "model digest mismatch blocks by default" \ + || fail "model digest mismatch did not block (rc=$DRC)" +[[ "$WRC" -eq 0 ]] && grep -q "MODEL_DIGEST_WARN" "$WORK/digest-warn.err" \ + && pass "model digest mismatch can warn during rollout" \ + || fail "model digest warn mode did not warn cleanly (rc=$WRC)" + +# 7: cloud backend reports unavailable honestly while keys are unset. if [[ -z "${ANTHROPIC_API_KEY:-}" ]]; then set +e bash "$ROUTER" "$WORK/s.txt" "$WORK/cloud.json" --role classify --model anthropic:claude-opus-4-8 >/dev/null 2>"$WORK/cloud.err" @@ -76,7 +97,7 @@ else skip "cloud-unavailable test (ANTHROPIC_API_KEY is set)" fi -# 7: deliberate failing primary route -> fallback through the REAL router (not exit 9). +# 8: deliberate failing primary route -> fallback through the REAL router (not exit 9). if [[ "$TUNNEL_UP" -eq 1 ]]; then printf 'Return exactly: OK\n' > "$WORK/f.txt" set +e diff --git a/AINative_OKR_CASAN5/.specify/traceability-map.json b/AINative_OKR_CASAN5/.specify/traceability-map.json new file mode 100644 index 0000000..7e84354 --- /dev/null +++ b/AINative_OKR_CASAN5/.specify/traceability-map.json @@ -0,0 +1,69 @@ +{ + "FR-01": { + "name": "Login", + "code": [ + "backend/src/auth/auth.controller.ts", + "backend/src/auth/auth.service.ts", + "frontend/src/pages/Login.tsx", + "frontend/src/hooks/useAuth.tsx" + ], + "tests": [ + "backend/test/services.test.ts", + "backend/test/e2e.test.ts" + ] + }, + "FR-02": { + "name": "Create Objective", + "code": [ + "backend/src/objectives/objectives.controller.ts", + "backend/src/objectives/objectives.service.ts", + "frontend/src/pages/CreateObjective.tsx", + "frontend/src/schemas/objective.schema.ts" + ], + "tests": [ + "backend/test/services.test.ts", + "backend/test/e2e.test.ts", + "frontend/src/__tests__/okr.test.tsx" + ] + }, + "FR-03": { + "name": "Create Key Result", + "code": [ + "backend/src/key-results/key-results.controller.ts", + "backend/src/key-results/key-results.service.ts", + "backend/src/key-results/dto/create-key-result.dto.ts" + ], + "tests": [ + "backend/test/services.test.ts", + "backend/test/e2e.test.ts" + ] + }, + "FR-04": { + "name": "Update Progress", + "code": [ + "backend/src/key-results/key-results.controller.ts", + "backend/src/key-results/key-results.service.ts", + "backend/src/key-results/dto/update-progress.dto.ts", + "frontend/src/pages/KeyResultDetail.tsx" + ], + "tests": [ + "backend/test/services.test.ts", + "backend/test/e2e.test.ts", + "frontend/src/__tests__/okr.test.tsx" + ] + }, + "FR-05": { + "name": "Dashboard", + "code": [ + "backend/src/objectives/objectives.controller.ts", + "backend/src/objectives/objectives.service.ts", + "frontend/src/pages/Dashboard.tsx", + "frontend/src/hooks/useObjectives.ts" + ], + "tests": [ + "backend/test/services.test.ts", + "backend/test/e2e.test.ts", + "frontend/src/__tests__/okr.test.tsx" + ] + } +} diff --git a/AINative_OKR_CASAN5/docs/output/casan/traceability-matrix.json b/AINative_OKR_CASAN5/docs/output/casan/traceability-matrix.json new file mode 100644 index 0000000..8d3bd75 --- /dev/null +++ b/AINative_OKR_CASAN5/docs/output/casan/traceability-matrix.json @@ -0,0 +1,100 @@ +{ + "generated_at": "2026-07-06T02:24:50Z", + "requirements_source": "docs/input/okr-requirement.md", + "mapping_source": ".specify/traceability-map.json", + "summary": { + "requirements": 5, + "passed": 5, + "failed": 0, + "orphan_mappings": [] + }, + "matrix": [ + { + "id": "FR-01", + "name": "Login", + "status": "PASS", + "code": [ + "backend/src/auth/auth.controller.ts", + "backend/src/auth/auth.service.ts", + "frontend/src/pages/Login.tsx", + "frontend/src/hooks/useAuth.tsx" + ], + "tests": [ + "backend/test/services.test.ts", + "backend/test/e2e.test.ts" + ], + "missing_code": [], + "missing_tests": [] + }, + { + "id": "FR-02", + "name": "Create Objective", + "status": "PASS", + "code": [ + "backend/src/objectives/objectives.controller.ts", + "backend/src/objectives/objectives.service.ts", + "frontend/src/pages/CreateObjective.tsx", + "frontend/src/schemas/objective.schema.ts" + ], + "tests": [ + "backend/test/services.test.ts", + "backend/test/e2e.test.ts", + "frontend/src/__tests__/okr.test.tsx" + ], + "missing_code": [], + "missing_tests": [] + }, + { + "id": "FR-03", + "name": "Create Key Result", + "status": "PASS", + "code": [ + "backend/src/key-results/key-results.controller.ts", + "backend/src/key-results/key-results.service.ts", + "backend/src/key-results/dto/create-key-result.dto.ts" + ], + "tests": [ + "backend/test/services.test.ts", + "backend/test/e2e.test.ts" + ], + "missing_code": [], + "missing_tests": [] + }, + { + "id": "FR-04", + "name": "Update Progress", + "status": "PASS", + "code": [ + "backend/src/key-results/key-results.controller.ts", + "backend/src/key-results/key-results.service.ts", + "backend/src/key-results/dto/update-progress.dto.ts", + "frontend/src/pages/KeyResultDetail.tsx" + ], + "tests": [ + "backend/test/services.test.ts", + "backend/test/e2e.test.ts", + "frontend/src/__tests__/okr.test.tsx" + ], + "missing_code": [], + "missing_tests": [] + }, + { + "id": "FR-05", + "name": "Dashboard", + "status": "PASS", + "code": [ + "backend/src/objectives/objectives.controller.ts", + "backend/src/objectives/objectives.service.ts", + "frontend/src/pages/Dashboard.tsx", + "frontend/src/hooks/useObjectives.ts" + ], + "tests": [ + "backend/test/services.test.ts", + "backend/test/e2e.test.ts", + "frontend/src/__tests__/okr.test.tsx" + ], + "missing_code": [], + "missing_tests": [] + } + ] +} diff --git a/casan-next-plans/CASAN_BACKLOG_STATUS.md b/casan-next-plans/CASAN_BACKLOG_STATUS.md index fb2f3e9..0cec4bd 100644 --- a/casan-next-plans/CASAN_BACKLOG_STATUS.md +++ b/casan-next-plans/CASAN_BACKLOG_STATUS.md @@ -4,14 +4,15 @@ > bước tiếp theo cụ thể + cờ phụ-thuộc-hạ-tầng, để **bất kỳ AI/người nào tiếp quản > cũng làm tiếp được ngay**. Cập nhật mỗi khi hoàn thành một mục. > -> Cập nhật lần cuối: 2026-07-06 · Nhánh: `feat/plan07-track-a-hardening` (đã merge vào `main` @ `c943273`). -> Test hiện tại: **211 PASS / 0 FAIL** trên 12 suite. Điểm công tâm ~81.6/100, harness thấp nhất 80. +> Cập nhật lần cuối: 2026-07-06 · Nhánh làm tiếp từ handoff Claude. +> Test hiện tại: **218 PASS / 0 FAIL** trên 13 core harness suite; `phase3-model-router` riêng **10 PASS / 0 FAIL**; frontend Vitest **16 PASS / 0 FAIL**. Backend `npm test` còn bị chặn bởi test-infra cũ (`schema.prisma` MySQL nhưng `setup-sqlite.mjs` chạy SQLite). +> Điểm công tâm vẫn quanh **~81/100**, harness thấp nhất 80; TIER 2 infra thật vẫn là trần Strong. > Nguồn liên quan: `CASAN_HARDENING_STATUS.md` (chi tiết control) · `evidence/scoring-run-report.md` (điểm). ## Quy ước - **Nhãn:** ✅ done+test · 🟡 MVP done (bản prod cần hạ tầng) · 🟦 deliverable-now (làm được offline) · 🔌 needs-infra (cần key/dịch vụ ngoài) · 📋 handoff-only (platform lớn, cần nhiều phiên). - **Nguyên tắc bàn giao:** mỗi mục 🔌/📋 phải có "Bước tiếp theo" đủ cụ thể để người khác bắt tay ngay. -- **Bất biến an toàn:** giữ 211/0; mọi control mới phải có test đối kháng fail-able; không hardcode verdict, không bypass. +- **Bất biến an toàn:** giữ 218/0 core harness; mọi control mới phải có test đối kháng fail-able; không hardcode verdict, không bypass. --- @@ -19,9 +20,9 @@ | # | Hạng mục | Plan | Trạng thái | Bước tiếp theo | |---|---|---|:--:|---| -| T1.1 | **Model-digest pinning** (chống model bị tráo/poison) | 07 B4 / V16 | 🟦 → (làm trong phiên này) | Ghi digest model đã duyệt vào `security/model-digest.pin`; `model-router.sh`/1 gate so digest hiện tại; đổi digest bất ngờ → cảnh báo/lock. Test: pin→match=OK, đổi→WARN/BLOCK. | -| T1.2 | **IdP/OIDC cho approval** (thay registry pubkey tĩnh) | 07 C4 / V20 | 🟦 (mock IdP) | Verify approver bằng **JWT ký bởi IdP** (RS256): `approval-verify.sh` chấp nhận `CASAN_APPROVAL_JWT` — verify chữ ký bằng JWKS/pubkey, check claim `sub`(=approver)/`role`/`exp`. MVP dùng khoá IdP-giả (như Vault dev); prod trỏ JWKS thật. Test: JWT hợp lệ→APPROVED, hết hạn/sai role/chữ ký giả→DENY. | -| T1.3 | **Traceability REQ→code→test** (khác biệt nhất) | 10 | 🟦 → (làm trong phiên này) | Sinh ma trận: parse FR-xx từ `docs/input/okr-requirement.md` → map sang file code + test; gate: mọi FR phải có ≥1 code + ≥1 test, cảnh báo FR mồ côi. Test: FR đủ→PASS, FR thiếu test→FAIL. | +| T1.1 | **Model-digest pinning** (chống model bị tráo/poison) | 07 B4 / V16 | ✅ done+test | `.specify/security/model-digest.pin` pin `ornith:9b`; `model-call.py` gọi `model-digest-check.sh verify` trước Ollama; mismatch mặc định BLOCK, rollout mode WARN. Test: `phase3-model-router` pin→OK, đổi→WARN/BLOCK. | +| T1.2 | **IdP/OIDC cho approval** (thay registry pubkey tĩnh) | 07 C4 / V20 | ✅ MVP done+test | `approval-verify.sh` chấp nhận `CASAN_APPROVAL_JWT` RS256 ký bởi mock IdP, verify chữ ký bằng pubkey, check `sub`/`role`/`exp` + bind `action`/`actor`/`input_sha256`. Test: JWT hợp lệ→APPROVED, hết hạn/sai role/chữ ký giả→DENY. Prod còn cần IdP/JWKS thật. | +| T1.3 | **Traceability REQ→code→test** (khác biệt nhất) | 10 | ✅ MVP done+test | `traceability-matrix.py` parse FR-xx từ requirement, dùng `.specify/traceability-map.json`, gate mọi FR có ≥1 code + ≥1 test; Evidence Pack thêm `traceability-matrix.json`. Test: FR đủ→PASS, FR thiếu test→FAIL. | ## TIER 2 — Plan-07 gaps cần hạ tầng (MVP đã có, cần bản prod) 🔌 @@ -38,7 +39,7 @@ | Plan | Trạng thái | Lõi cần làm (bước tiếp theo cho AI kế) | |---|:--:|---| -| **10 Traceability + H3 Eval** | 🟦 MVP phiên này (xem T1.3) | Sau MVP: nối traceability vào Evidence Pack (thêm `traceability-matrix.json` vào pack); mở rộng H3 eval-set độc lập (nhiều model). | +| **10 Traceability + H3 Eval** | ✅ MVP done+test | Đã nối traceability vào Evidence Pack. Sau MVP: line/symbol-level traceability + H3 eval-set độc lập (nhiều model). | | **02 LLM source-gen** | 📋 chưa bắt đầu | Thay template bằng LLM thật sinh source qua `model-router.sh`; đi qua wrapper H4→H7. Phụ thuộc 03. Bước 1: định contract prompt→file cho 1 module (objectives), gate bằng H3 judge + traceability. | | **03 Cloud patch** | 📋 chưa bắt đầu | Bỏ stub trong `model-call.py` cho OpenAI/Anthropic; test bằng key thật (🔌). Bước 1: env `CASAN_MODEL_PRIMARY=openai:…`, xác thực round-trip + ghi provider-usage thật. | | **04 Self-improve** | 📋 chưa bắt đầu | Khép vòng `casan improve`: đọc metrics/drift/hallucination → đề xuất vá → chạy lại gate. Phụ thuộc 02, 05. Bước 1: script đọc `metrics.jsonl` + `drift-report.json` → sinh backlog vá tự động. | @@ -46,7 +47,7 @@ | **06 Onboard dự án 2** | 📋 chưa bắt đầu | Chứng minh reuse: cắm 1 repo khác + golden/corpus/input, đăng ký qua `verify-harness-reuse.sh`, không sửa gate. Phụ thuộc 01. | | **08 Context compression** | 📋 chưa bắt đầu | Nén prompt/context giảm token (H1.5). Làm SAU khi core ổn (nén thêm bề mặt rủi ro → cần quét lại). Bước 1: đo token baseline mỗi step, thử tóm tắt có kiểm chứng (H3 judge so sánh). | | **12 Domain Pack SDK** | 📋 chưa bắt đầu | Onboard bằng khai báo (golden/corpus/policy theo domain). Phụ thuộc 01, 06. | -| **01 Restructure** | 📋 chưa bắt đầu | Tái cấu trúc thư mục Phase 0→6. Nền cho 06/12. Rủi ro cao (đụng nhiều path) → làm trên nhánh riêng, giữ 211/0. | +| **01 Restructure** | 📋 chưa bắt đầu | Tái cấu trúc thư mục Phase 0→6. Nền cho 06/12. Rủi ro cao (đụng nhiều path) → làm trên nhánh riêng, giữ 218/0. | | Future B1–B6 | 💤 vision | `CASAN_PLAN_FUTURE_PHASES.md` — approval workflow nâng cao · state machine · model benchmark · governed memory · auto-remediation · platform KPI. | --- @@ -62,7 +63,8 @@ Harness thấp nhất = 80 (H5, H6). Để CẢ pipeline vào Strong cần đón cd AINative_OKR_CASAN5 for s in run-casan4-harness adversarial-harness phase1-track-a phase2-track-c \ phase3-evidence-pack phase-h5-approval phase-h5-infra phase-h6-agentops \ - phase-c7-incident phase-h4-multilingual phase-c6-sandbox phase-h4-split-inject; do + phase-c7-incident phase-h4-multilingual phase-c6-sandbox phase-h4-split-inject \ + phase10-traceability; do bash .specify/tests/$s-tests.sh >/dev/null 2>&1 && echo "$s OK" || echo "$s FAIL"; done # KMS live cần Vault dev; C6 live cần Docker (không có thì SKIP, không fail). ``` diff --git a/casan-next-plans/CASAN_HARDENING_STATUS.md b/casan-next-plans/CASAN_HARDENING_STATUS.md index 025489b..dee2556 100644 --- a/casan-next-plans/CASAN_HARDENING_STATUS.md +++ b/casan-next-plans/CASAN_HARDENING_STATUS.md @@ -34,13 +34,13 @@ | Control | Where | Test | |---|---|---| -| `casan pack` / `casan verify-pack` (mapped to `evidence-pack.sh`): standard 12-file pack, hash manifest, signed head, tamper-evident verify, certified-run gate (no false certification, no silent skip) | `evidence-pack.sh`, `evidence-pack-build.py`, `evidence-pack-verify.py` | phase3-evidence-pack | +| `casan pack` / `casan verify-pack` (mapped to `evidence-pack.sh`): standard 13-file pack incl. `traceability-matrix.json`, hash manifest, signed head, tamper-evident verify, certified-run gate (no false certification, no silent skip) | `evidence-pack.sh`, `evidence-pack-build.py`, `evidence-pack-verify.py` | phase3-evidence-pack | ### Phase 4 — H5 governance hardening (raises the lowest harness) — mixed | ID | Control | Status | Where | Test | |---|---|---|---|---| -| C4 | **Approval-identity**: high-risk approval trusted only when a REGISTERED reviewer cryptographically SIGNS the request and their role is authorized — env-var approver no longer enough (SoD still enforced) | [implemented+tested] | `approval-sign.sh`, `approval-verify.sh`, `reviewers.registry`, `governance-check.sh` (`CASAN_APPROVAL_STRICT=1`) | phase-h5-approval (8) | +| C4 | **Approval-identity + mock IdP/OIDC**: high-risk approval trusted only when a REGISTERED reviewer signs the request OR an IdP-signed RS256 JWT proves approver identity/role/expiry and binds to this request — env-var approver no longer enough (SoD still enforced) | [implemented+tested] (mock IdP; production JWKS still planned) | `approval-sign.sh`, `approval-jwt-mint.py`, `approval-verify.sh`, `reviewers.registry`, `governance-check.sh` (`CASAN_APPROVAL_STRICT=1`) | phase-h5-approval (12) | | B3 | **KMS key management**: sign audit/telemetry head via Vault Transit (key `exportable:false` → never leaves KMS) + key **rotation**; validated live | [implemented+tested] (live when Vault present; skip-aware otherwise) | `vault-kms.sh` (`rotate`, `assert-nonexportable`), `sign-audit-head.sh` | phase-h5-infra (KMS) | | C5 | **External WORM audit**: ship audit head to a hash-linked append-only ledger; detect local rollback (`AUDIT_GAP_DETECTED`) and ledger tamper (`AUDIT_LEDGER_TAMPERED`) | [implemented+tested] (local ledger MVP) | `worm-ledger.py`, `audit-ship.sh`, `verify-audit-gap.sh` | phase-h5-infra (WORM) | @@ -61,6 +61,8 @@ | B1 | **Multilingual VI/JA injection (V2)**: VI/JA block-patterns (matched on raw UTF-8, anchored on the injection object) catch injections English regex missed, with 0 false positives on the benign VI/JA corpus | [implemented+tested] | `prompt-filter.yaml` (PI-VI-*, PI-JA-*) | phase-h4-multilingual (7) | | C6 | **TRUE runtime isolation (V22)**: container sandbox (`--network=none --read-only --pids-limit --cap-drop=ALL`, workspace-only mount) — the kernel neutralises host-file reads / egress / out-of-workspace writes; upgrades the static scaffold | [implemented+tested] (live via Docker; skip-aware) | `sandbox-container.sh`, `sandbox-run.sh` (`CASAN_SANDBOX_MODE=container`) | phase-c6-sandbox (6) | | B2 | **Split + classifier injection (V5,V6)**: `context-assemble-scan.sh` scans the concatenated context so a payload split across benign pieces is caught on assembly; verdict-steering patterns (PI-CLS-*) block content that tries to hijack the evaluator | [implemented+tested] | `context-assemble-scan.sh`, `prompt-filter.yaml` (PI-CLS-*) | phase-h4-split-inject (8) | +| B4 | **Model-digest pinning (V16)**: approved Ollama model digest pinned; router verifies live digest before generation/classify/judge; mismatch blocks by default with warn mode for rollout | [implemented+tested] | `model-digest-check.sh`, `model-call.py`, `.specify/security/model-digest.pin` | phase3-model-router (digest cases) | +| Plan-10 | **Traceability REQ→code→test MVP**: parse `FR-*`, require code+test mapping per FR, generate matrix, and include it in Evidence Pack | [implemented+tested] | `traceability-matrix.py`, `.specify/traceability-map.json`, `docs/output/casan/traceability-matrix.json` | phase10-traceability (3) | ## 2. Test inventory (all suites) @@ -71,14 +73,15 @@ | `phase1-track-a-tests.sh` | 25 | Track A hardening | | `phase2-track-c-tests.sh` | 29 | Track C-MVP | | `phase3-evidence-pack-tests.sh` | 7 | Evidence Pack MVP | -| `phase-h5-approval-tests.sh` | 8 | Approval-identity (C4) | +| `phase-h5-approval-tests.sh` | 12 | Approval-identity (C4) + mock IdP/OIDC JWT | | `phase-h5-infra-tests.sh` | 7 | KMS (B3, live/skip-aware) + WORM (C5) | | `phase-h6-agentops-tests.sh` | 20 | live alerting (D1) + provider-API/reconcile (D2) + hosted dashboard (D3) + window breaker (D4); against live local HTTP endpoints | | `phase-c7-incident-tests.sh` | 15 | **New** — incident severity + scoped kill-switch (C7) + wrapper enforcement | | `phase-h4-multilingual-tests.sh` | 7 | **New** — VI/JA injection block + benign VI/JA 0-FP (B1) | | `phase-c6-sandbox-tests.sh` | 6 | **New** — TRUE container isolation (C6, live via Docker / skip-aware) | | `phase-h4-split-inject-tests.sh` | 8 | **New** — split-injection assembly scan + classifier-inject (B2) | -| **Total** | **211** | Baseline 79 preserved; +132 new hardening checks. Last full run 2026-07-05 @ head of `feat/plan07-track-a-hardening`, 0 fail (KMS + container isolation validated live via Vault dev + Docker). | +| `phase10-traceability-tests.sh` | 3 | **New** — Plan-10 FR→code→test matrix + fail-able missing-test gate | +| **Total** | **218** | Baseline 79 preserved; +139 new hardening/traceability checks. Last full harness run 2026-07-06, 0 fail. Direct `phase3-model-router-tests.sh` adds 10/0 for model-router/digest/cloud/fallback checks. | Run order note: `run-casan4-harness-tests.sh` does `rm -rf .specify/logs`, so run it **first** and never concurrently with the other suites. @@ -87,13 +90,13 @@ Run order note: `run-casan4-harness-tests.sh` does `rm -rf .specify/logs`, so ru | Area | Status | Plan ref | |---|---|---| -| Model-digest pinning | [planned] — sliding-window circuit breaker (V15) is now done (Phase 5 D4) | Plan-07 B4 (V16) | +| Production model provenance beyond local Ollama digest | [partial] — local model digest pinning is implemented+tested; cloud provider model attestations/SBOM-style provenance still planned | Plan-07 B4 (V16) | | Live alerting to a managed channel (Slack/PagerDuty + on-call rota) | [partial] — webhook dispatch + dedup + dead-letter done; managed channel & escalation are config away, incident workflow is C7 | Plan-07 C7 / Phase 5 D1 | | Hosted telemetry dashboard | [partial] — HTTP-served dashboard + stale-aware `/healthz` done locally; deployed host (nginx/container, auth) planned | Phase 5 D3 | | Provider billing-API telemetry | [partial] — API fetch + schema gate + local-vs-provider reconciliation done against a live local endpoint; real OpenAI/Anthropic usage-API calls (needs keys) planned | Phase 5 D2 | | True runtime isolation | [partial] — real container isolation done + validated live via Docker (C6 phase-6); nsjail/rootless + a hardened base image for CI still planned | Plan-07 C6 (V22) | | KMS key management (rotation, non-exportable) | [partial] — Vault Transit path implemented + validated live; not yet the default (local-key fallback), no HSM/short-lived IdP tokens | Plan-07 B3 | -| Reviewer approval workflow | [partial] — cryptographic **approval-identity** done (signed reviewer + role); live **IdP (OIDC/JWT)** + policy versioning/diff still planned | Plan-07 C4 (V20) | +| Reviewer approval workflow | [partial] — cryptographic **approval-identity** done (signed reviewer + role) + mock **IdP/OIDC JWT** done; live IdP/JWKS + policy versioning/diff still planned | Plan-07 C4 (V20) | | External append-only (WORM) audit | [partial] — hash-linked local ledger + rollback/tamper detection done; true WORM store (S3 Object Lock/QLDB) + trusted timestamp planned | Plan-07 C5 (V21) | | Live CVE/OSV scanning wired in | [partial] — availability detected; local denylist authoritative offline | Plan-07 C2 follow-up | @@ -103,14 +106,14 @@ Track A + Track C-MVP + Evidence Pack + H5/H6 hardening + the deep-gap closers (C7 incident/kill-switch, VI/JA multilingual, true container isolation, split & classifier injection) raise H4/H5/H6 from "PoC/demo (~3.0/5)" to **early internal-production hardening**, with executable adversarial tests for every -control (**211 checks, 0 fail** — last full run 2026-07-05; KMS + container -isolation validated live via Vault dev + Docker). Fair maturity score +control (**218 core checks, 0 fail** — last full harness run 2026-07-06; KMS + container +isolation validated live/skip-aware via Vault dev + Docker). Fair maturity score (`00_SUBMISSION_PACKAGE/evidence/scoring-run-report.md`): **H4 80→83** (multilingual + split/classifier closed), **H2 80→82** (real sandbox isolation), C7 incident dimension closed; **H5 and H6 remain at 80** (their remaining gaps are infra), so the **lowest harness stays 80** — CASAN **Level 4**, proven by attack. This is **not** full production readiness: crossing the whole pipeline into "Strong (81+)" still needs the -H5/H6 infra items — live IdP (OIDC/JWT), a true WORM store (S3 Object Lock), +H5/H6 infra items — live IdP/JWKS, a true WORM store (S3 Object Lock), KMS-by-default + HSM, a deployed dashboard host + managed alert channel/on-call, and real billing-API telemetry — the [partial]/[planned] rows above and in `CASAN_PLAN_07_PRODUCTION_HARDENING.md`. diff --git a/casan-next-plans/CASAN_PLAN_00_INDEX.md b/casan-next-plans/CASAN_PLAN_00_INDEX.md index 8b1f913..9e4560a 100644 --- a/casan-next-plans/CASAN_PLAN_00_INDEX.md +++ b/casan-next-plans/CASAN_PLAN_00_INDEX.md @@ -19,7 +19,7 @@ > **Thứ tự trong Plan-07:** **Track A** (quick wins) → **Track C-MVP** = C1 Tool-authz + C2 Supply-chain + C3 Data-exfil + C6 Sandbox (**minimum bar** trước khi cho agent ghi code/chạy test production-like) → Track B + C-Governance/Ops (production nghiêm túc). | 08 | `CASAN_PLAN_08_CONTEXT_COMPRESSION.md` | Nén prompt/context giảm token (H1.5 cross-cutting dưới H1/H6) | 01, 07, 03 | Trung (tối ưu, sau Plan-07 core) | | 09 | `CASAN_PLAN_09_EVIDENCE_PACK.md` | Evidence Pack & Certification (proof pack mỗi run) | nhẹ (gom H1–H7) | **Cao (đúng phương châm)** | -| 10 | `CASAN_PLAN_10_TRACEABILITY_EVAL.md` | Traceability REQ→code→test + H3 Evaluation (gộp ý "11") | 02, 09 | Cao (khác biệt nhất) | +| 10 | `CASAN_PLAN_10_TRACEABILITY_EVAL.md` | Traceability REQ→code→test + H3 Evaluation (gộp ý "11") | 02, 09 | Cao (khác biệt nhất) | | 12 | `CASAN_PLAN_12_DOMAIN_PACK.md` | Domain Pack SDK — onboard bằng khai báo | 01, 06 | Trung–Cao (reuse thật) | | — | `CASAN_PLAN_FUTURE_PHASES.md` | Backlog phase sau (B1–B6: approval workflow · state machine · model benchmark · governed memory · auto-remediation · platform KPI) | các plan nền | Vision (chưa làm) | @@ -64,9 +64,9 @@ flowchart LR | 04 Self-improve | ⬜ chưa bắt đầu | | | 05 CI/CD | ⬜ chưa bắt đầu | | | 06 Onboard | ⬜ chưa bắt đầu | | -| 07 Production hardening | 🟡 Track A ✅ + C-MVP ✅ + H5-hardening ✅ (B + C7 + IdP/WORM-store chưa) | **Track A 25/25 · Track C-MVP 29/29 · H5-hardening (C4 approval-identity 8/8 + KMS live + C5 WORM = infra 7/7)**. Điểm công tâm: H5 76→80, harness thấp nhất giờ H6=79, ~80.7/100. Track B + C7 + live IdP/WORM-store còn [planned]. Chi tiết: `CASAN_HARDENING_STATUS.md` · `evidence/scoring-run-report.md` | +| 07 Production hardening | 🟡 Track A ✅ + C-MVP ✅ + H5/H6/deep-gap hardening ✅ (infra prod còn) | **Track A 25/25 · Track C-MVP 29/29 · H5 approval/OIDC 12/12 · infra 7/7 · H6 20/20 · C7/H4/C6 deep-gap closers done**. Điểm công tâm ~81/100; harness thấp nhất 80. Live IdP/JWKS, WORM-store thật, KMS default/HSM, managed alert/billing/dashboard deploy còn [planned]. Chi tiết: `CASAN_HARDENING_STATUS.md` · `evidence/scoring-run-report.md` | | 08 Context compression | ⬜ chưa bắt đầu | H1.5 cross-cutting; làm **sau** Plan-07 core (nén thêm bề mặt rủi ro) | -| 09 Evidence Pack | 🟢 MVP ✅ (7/7) | `casan pack`/`verify-pack` → `evidence-pack.sh`: 12-file pack + manifest + signed head + certified-run gate, tamper-evident. Chi tiết: `CASAN_HARDENING_STATUS.md` | -| 10 Traceability + H3 Eval | ⬜ chưa bắt đầu | Khác biệt nhất; lõi = ma trận REQ→code→test | +| 09 Evidence Pack | 🟢 MVP ✅ (7/7) | `casan pack`/`verify-pack` → `evidence-pack.sh`: 13-file pack incl. traceability matrix + manifest + signed head + certified-run gate, tamper-evident. Chi tiết: `CASAN_HARDENING_STATUS.md` | +| 10 Traceability + H3 Eval | 🟢 MVP ✅ (3/3) | `traceability-matrix.py --gate`: parse FR-xx → code/test matrix; Evidence Pack chứa `traceability-matrix.json`. Còn line/symbol-level + eval-set độc lập. | | 12 Domain Pack SDK | ⬜ chưa bắt đầu | Cần Plan-01 xong trước | | Future phases (B1–B6) | 💤 vision | Backlog `CASAN_PLAN_FUTURE_PHASES.md` — chưa xây | diff --git a/casan-next-plans/CASAN_PLAN_07_PRODUCTION_HARDENING.md b/casan-next-plans/CASAN_PLAN_07_PRODUCTION_HARDENING.md index d36208b..775b34e 100644 --- a/casan-next-plans/CASAN_PLAN_07_PRODUCTION_HARDENING.md +++ b/casan-next-plans/CASAN_PLAN_07_PRODUCTION_HARDENING.md @@ -19,7 +19,7 @@ ## 2. Thang điểm sẵn sàng production (0–5, cao = tốt) -> ✅ **CẬP NHẬT 2026-07-05 — Track A + C-MVP + Evidence Pack + H5-hardening + H6-hardening ĐÃ LÀM + TEST (175 checks, 0 fail; KMS đã validate LIVE qua Vault 2026-07-04).** +> ✅ **CẬP NHẬT 2026-07-06 — Track A + C-MVP + Evidence Pack + H5/H6/deep-gap hardening + Plan-10 traceability ĐÃ LÀM + TEST (218 core checks, 0 fail; model-router riêng 10/0).** > Bảng dưới có cột **Baseline → Nay**. Điểm chấm CÔNG TÂM (0–100, theo `casan_harness_assessment.md`): > **H4 = 80 · H5 = 76→80 ⬆ · H6 = 79→80 ⬆ · trung bình 7 harness ~80.9/100 · không còn harness nào dưới 80 → CASAN Level 4 (vững ngưỡng)**. > Nguồn: `00_SUBMISSION_PACKAGE/evidence/scoring-run-report.md`. Chi tiết implemented-vs-planned: `CASAN_HARDENING_STATUS.md`. @@ -33,7 +33,7 @@ | Quan sát (observability) | 3 | **4** | telemetry toàn vẹn (ký) · **alerting LIVE** (webhook + dedup + dead-letter, end-to-end từ step fail) · **dashboard hosted** (`/healthz` stale-aware) · **provider-API reconcile** (bắt under-reporting) · window breaker (V15) | dashboard deploy thật + auth · kênh alert managed (Slack/PagerDuty + on-call) · billing-API thật | | Đa domain/i18n | 2 | **2.5** | benign corpus VI/JA/EN đo được (FP 0%) | detection vẫn chủ yếu EN (Track B) | | Quản lý khóa | 2 | **4** | **KMS live** (Vault Transit) — ký qua KMS, **rotate**, khoá **non-exportable** (đã chạy thật) | KMS chưa mặc định (fallback local) · HSM · IdP token ngắn hạn | -| Phủ kiểm thử | 4 | **4.5** | **175 test** (35+44+25+29+7+8+7+20) đối kháng, 0 fail | thêm ca đa ngôn ngữ khi làm Track B | +| Phủ kiểm thử | 4 | **4.5** | **218 core tests** (35+44+25+29+7+12+7+20+15+7+6+8+3) đối kháng, 0 fail | line/symbol-level traceability + CI release gate | **Điểm trung bình (H4/H5/H6 mở rộng) ~3.0 → ~4.0/5; chấm công tâm per-harness H4=80 · H5=80 · H6=80 (~4.0/5). Không còn harness nào dưới 80.** @@ -45,12 +45,12 @@ | Supply-chain (dependency sinh ra) | 1 | **3.5** | `supply-chain-gate.sh` (V18) — typosquat/postinstall/denylist + dep-diff | CVE/OSV scanner live chưa nối | | Data-governance / anti-exfil | 2 | **4** | `data-exfil-guard.sh` (V19) — secret→cloud BLOCK, PII→audit mask | — | | Evidence Pack (Plan-09) | — | **4** | `casan pack/verify-pack` — manifest ký, tamper-evident, certified-gate | KMS sign · hosted store | -| Policy governance / approval | 2 | **4** | **approval-identity** (C4) — reviewer KÝ request + role authz, hết env-var; SoD giữ (đã test 8/8) | live IdP (OIDC/JWT) · policy versioning/diff [planned] | +| Policy governance / approval | 2 | **4** | **approval-identity + mock IdP/OIDC** (C4) — reviewer KÝ request hoặc JWT RS256 + role authz, hết env-var; SoD giữ (đã test 12/12) | live IdP/JWKS · policy versioning/diff [planned] | | Runtime sandbox | 1 | **2.5** | `sandbox-run.sh` scaffold (V22) — chặn ssh/egress/forkbomb/write-outside + ulimit | **cô lập kernel thật** (container/nsjail) [planned] | | External append-only audit | 1 | **3.5** | **WORM ledger** (C5) — ship head hash-link ngoài, bắt rollback (`AUDIT_GAP_DETECTED`) + tamper | WORM store thật (S3 Object Lock) · trusted timestamp [planned] | | Incident response | 1 | **1** | — | severity/owner/kill-switch (C7) [planned] | -**→ C-MVP (C1+C2+C3 + Evidence Pack) ~3.8/5 + H5-hardening (C4 approval-identity, KMS live, C5 WORM) + H6-hardening (D1 alerting live, D2 provider-API reconcile, D3 dashboard hosted, D4 window breaker V15) [đã làm + test thật]. Còn: sandbox isolation thật, C7 incident, live IdP/WORM-store, dashboard deploy + kênh alert managed, billing-API thật [planned].** Trio H4/H5/H6 nay **~4.0/5 (H4=80·H5=80·H6=80)**; harness thấp nhất nhích **76 (H5) → 79 (H6) → 80 (đồng đều)**. Production toàn diện vẫn cần các mục [planned] ở trên. +**→ C-MVP (C1+C2+C3 + Evidence Pack) ~3.8/5 + H5-hardening (C4 approval-identity + mock OIDC, KMS live, C5 WORM) + H6-hardening (D1 alerting live, D2 provider-API reconcile, D3 dashboard hosted, D4 window breaker V15) + deep-gap closers + Plan-10 traceability [đã làm + test thật]. Còn: live IdP/JWKS, WORM-store thật, KMS default/HSM, dashboard deploy + kênh alert managed, billing-API thật [planned].** Trio H4/H5/H6 nay **~4.0/5 (H4=80·H5=80·H6=80)**; harness thấp nhất nhích **76 (H5) → 79 (H6) → 80 (đồng đều)**. Production toàn diện vẫn cần các mục [planned] ở trên. ## 3. Bảng đường lọt (tóm tắt từ threat-model) diff --git a/casan-next-plans/CASAN_PLAN_10_TRACEABILITY_EVAL.md b/casan-next-plans/CASAN_PLAN_10_TRACEABILITY_EVAL.md new file mode 100644 index 0000000..b4096b9 --- /dev/null +++ b/casan-next-plans/CASAN_PLAN_10_TRACEABILITY_EVAL.md @@ -0,0 +1,36 @@ +# CASAN PLAN 10 — Traceability REQ→Code→Test + H3 Eval + +> Status 2026-07-06: **MVP implemented + tested**. Scope is deterministic +> traceability for the OKR sample app; broader H3 eval-set expansion remains a +> platform follow-up. + +## MVP delivered + +| Capability | Where | Verification | +|---|---|---| +| Parse `FR-*` from `docs/input/okr-requirement.md` | `.specify/scripts/bash/traceability-matrix.py` | `phase10-traceability-tests.sh` | +| Declarative FR→code→test map | `.specify/traceability-map.json` | gate checks every mapped file exists | +| Gate every FR has >=1 code file and >=1 test file | `traceability-matrix.py --gate` | missing FR test coverage fails | +| Evidence Pack includes traceability matrix | `evidence-pack-build.py` | `phase3-evidence-pack-tests.sh` expects 13 files | + +## Current result + +```bash +python3 .specify/scripts/bash/traceability-matrix.py --gate +# TRACEABILITY_MATRIX requirements=5 pass=5 fail=0 ... + +bash .specify/tests/phase10-traceability-tests.sh +# TRACEABILITY SUMMARY: PASS=3 FAIL=0 +``` + +Generated artifact: + +- `AINative_OKR_CASAN5/docs/output/casan/traceability-matrix.json` + +## Next steps + +| Priority | Work | Done when | +|---|---|---| +| P1 | Add traceability line references or symbol references, not only file paths | matrix can point to precise code/test locations | +| P2 | Expand H3 eval-set beyond the OKR sample app | independent eval corpus runs through judge gates | +| P3 | Enforce traceability in release CI | CI blocks a new FR without code+test coverage | diff --git a/casan-next-plans/CASAN_PLAN_QA.md b/casan-next-plans/CASAN_PLAN_QA.md index c586fba..04bc777 100644 --- a/casan-next-plans/CASAN_PLAN_QA.md +++ b/casan-next-plans/CASAN_PLAN_QA.md @@ -18,7 +18,7 @@ Không — vì chúng tôi **cố ý không mở hết**. Chỉ 3 plan mới đ Vì (a) rủi ro vỡ bản demo đang chạy; (b) một số plan **phụ thuộc nhau** (ví dụ Domain Pack cần restructure xong; Model benchmark cần nối model thật xong mới có số). Làm sai thứ tự = tốn công mà không có bằng chứng. **A4. Roadmap này có làm mất tính trung thực khi trình bày không?** -Không, nếu nói đúng nhãn: *"phần đã làm & đo là H4/H5/H6 + hardening; Evidence Pack/Traceability/Domain Pack là bước kế tiếp đã có kế hoạch; state machine/governed memory là tầm nhìn dài hạn chưa xây."* Ranh giới rõ ràng = điểm cộng độ chín. +Không, nếu nói đúng nhãn: *"phần đã làm & đo là H4/H5/H6 + hardening; Evidence Pack và Traceability MVP đã có test; Domain Pack là bước kế tiếp đã có kế hoạch; state machine/governed memory là tầm nhìn dài hạn chưa xây."* Ranh giới rõ ràng = điểm cộng độ chín. **A5. Nếu chỉ được chọn 3 plan, chọn gì và vì sao?** **09 Evidence Pack · 10 Traceability · 12 Domain Pack.** Ba cái này nâng CASAN từ "harness bảo vệ AI" lên "**nền tảng AI-SDLC có bằng chứng, đo chất lượng, tái dùng đa domain**" — đúng 3 trục thi.