Complete CASAN backlog tier 1 controls
This commit is contained in:
@@ -6,20 +6,23 @@
|
||||
## 1. Môi trường chạy
|
||||
| | |
|
||||
|---|---|
|
||||
| Thời điểm | **2026-07-05 (UTC)** — chạy tuần tự lại toàn bộ sau H6-hardening |
|
||||
| Git | `feat/plan07-track-a-hardening` (sau `2af67ef` + H6-hardening) |
|
||||
| Thời điểm | **2026-07-06 (JST)** — chạy tuần tự lại sau B4 digest pinning + C4 mock IdP/OIDC + Plan-10 traceability |
|
||||
| Git | working tree handoff continuation (commit pending) |
|
||||
| Model | Ollama `ornith:9b` @127.0.0.1:11434 — live |
|
||||
| KMS | Vault @127.0.0.1:8200 — **down lần chạy này** (KMS SKIP; đã validate **live** 2026-07-04 @ `00aabfa`) |
|
||||
| Runtime | node v24.12.0 · python 3.9.0 · openssl 3.6.3 · macOS |
|
||||
|
||||
## 2. Bằng chứng thô (100% thật, exit code on-screen — chạy tuần tự lại lần này)
|
||||
- **175 test PASS / 0 FAIL** trên **8 suite**:
|
||||
- **218 test PASS / 0 FAIL** trên **13 core harness suite**:
|
||||
run-casan4 **35** · adversarial **44** · phase1-track-a **25** · phase2-track-c **29** ·
|
||||
phase3-evidence-pack **7** · phase-h5-approval **8** · phase-h5-infra **7** · **phase-h6-agentops 20 (mới)**.
|
||||
phase3-evidence-pack **7** · phase-h5-approval **12** · phase-h5-infra **7** · phase-h6-agentops **20** ·
|
||||
phase-c7-incident **15** · phase-h4-multilingual **7** · phase-c6-sandbox **6** · phase-h4-split-inject **8** · phase10-traceability **3**.
|
||||
- Direct model-router suite: **10 PASS / 0 FAIL** (includes model-digest pin OK, mismatch BLOCK, mismatch WARN rollout mode).
|
||||
- Frontend Vitest: **16 PASS / 0 FAIL**. Backend `npm test` is blocked by pre-existing app-test infra mismatch (`schema.prisma` provider MySQL but `setup-sqlite.mjs` applies the MySQL migration to SQLite).
|
||||
- `security-gate` aggregate: **verdict PASS=11 FAIL=0 SKIP=0** (run 2026-07-04).
|
||||
- H4 recall: model 0.85 > regex 0.00 (GATE PASS). Benign-FP: **fp_rate 0.00% · block_rate 100.00%** (95 mẫu benign EN/VI/JA + 12 vector).
|
||||
- **H5 (run 2026-07-04, Vault live):**
|
||||
- Approval-identity: env-var approver → BLOCK; reviewer **ký** request + đúng role → APPROVED; chữ ký giả / sai role / replay / tự-duyệt → BLOCK (8/8).
|
||||
- Approval-identity: env-var approver → BLOCK; reviewer **ký** request + đúng role → APPROVED; mock IdP JWT hợp lệ → APPROVED; JWT hết hạn / sai role / chữ ký giả → BLOCK (12/12).
|
||||
- KMS **live** (Vault Transit): sign→verify (v1) → **rotate** → sign→verify (v2) → **khoá NON-exportable** (export bị từ chối).
|
||||
- WORM: ship anchor → in-sync; rollback log local → **AUDIT_GAP_DETECTED**; sửa ledger → **AUDIT_LEDGER_TAMPERED**.
|
||||
- **H6 mới (chạy thật lần này — mọi check chạy LIVE qua endpoint HTTP local, cùng chuẩn "live" như Vault dev):**
|
||||
@@ -35,9 +38,9 @@
|
||||
|----|---------|:--:|:--:|:--:|---|---|---|
|
||||
| H1 | Context | 90 | 84 | **84** | Good+ | pipeline-context pointer store, log-levels, context-validate | chưa nén/RAG context lớn (Plan-08 planned) |
|
||||
| H2 | Tool | 75 | 80 | **80** | Good+ | registry, idempotency, rate-limit, action-gate, supply-chain, tool-audit | sandbox mới scaffold (chưa cô lập thật) |
|
||||
| H3 | Evaluation | 85 | 82 | **82** | Strong- | LLM-judge multi-gate live, golden dataset, auto-retry, judge-gate 5/0 | judge chỉ 1 model local, chưa eval-set độc lập quy mô |
|
||||
| H3 | Evaluation | 85 | 82 | **82** | Strong- | LLM-judge multi-gate live, golden dataset, auto-retry, judge-gate 5/0, Plan-10 FR→code→test matrix 3/0 | judge chỉ 1 model local, chưa eval-set độc lập quy mô/line-level traceability |
|
||||
| H4 | Security | 20 | 80 | **80** | Good(đỉnh) | injection (direct/paraphrase→model/obfus/unicode/base64), indirect-artifact, secret in+out, PII, tool-output scan, strict fail-closed, recall 0.85, FP 0% | multilingual VI/JA & split-injection [planned]; sandbox scaffold; 1 model |
|
||||
| H5 | Governance | 25 | 76 | **80** ⬆ | Good(đỉnh) | hash-chain + RSA head, SoD, least-privilege, rate-limit, secrets-scan, no-bypass, tool-audit, telemetry-integrity **+ approval-identity ký-danh-tính + KMS live (rotate/non-exportable) + WORM audit ngoài (gap/tamper)** | **IdP live** (registry pubkey tĩnh, chưa OIDC/JWT); **WORM store thật** (ledger local, chưa S3-Object-Lock); KMS chưa mặc định (fallback khoá local) |
|
||||
| H5 | Governance | 25 | 76 | **80** ⬆ | Good(đỉnh) | hash-chain + RSA head, SoD, least-privilege, rate-limit, secrets-scan, no-bypass, tool-audit, telemetry-integrity **+ approval-identity ký-danh-tính + mock IdP/OIDC JWT + KMS live (rotate/non-exportable) + WORM audit ngoài (gap/tamper)** | **IdP live/JWKS thật**; **WORM store thật** (ledger local, chưa S3-Object-Lock); KMS chưa mặc định (fallback khoá local) |
|
||||
| H6 | AgentOps | 30 | 79 | **80** ⬆ | Good(đỉnh) | cost-spike 4 chế độ (rel+abs+cumulative+cold-start), drift, hallucination-rate, telemetry token thật + ký toàn vẹn **+ alerting LIVE (webhook + dedup + dead-letter, end-to-end từ step fail) + provider-telemetry API (fetch + đối soát bắt under-reporting) + dashboard hosted (/healthz stale-aware) + window circuit-breaker (V15)** | dashboard host thật (deploy nginx/container + auth); kênh alert managed (Slack/PagerDuty + on-call, C7 incident); billing-API thật (OpenAI/Anthropic, cần key) |
|
||||
| H7 | Orchestration | 80 | 80 | **80** | Good(đỉnh) | Boss DAG, BACK-TO-PLAN/retry, rollback thật, model-fallback thật, drift | chưa transaction-rollback xuyên nhiều step |
|
||||
|
||||
@@ -52,7 +55,7 @@
|
||||
- (5/5 gate)×100 chỉ đo **độ phủ control**, không đo **độ trưởng thành/vận-hành-thật** — report này tách bạch: mục 2 = coverage/pass thật, mục 3 = trưởng thành công tâm.
|
||||
|
||||
## 5. Ranh giới trung thực
|
||||
- CASAN **Level 4 chứng minh bằng tấn công (175 test)**. Level 5 các control hiện thực + test cục bộ; production Level 5 cần **IdP thật**, **WORM store (S3 Object Lock)**, KMS mặc định + HSM, **billing-API thật** (usage endpoint OpenAI/Anthropic), **dashboard deploy thật** (nginx/container + auth), **kênh alert managed + on-call (C7)**, sandbox isolation thật.
|
||||
- CASAN **Level 4 chứng minh bằng tấn công (218 core harness tests)**. Level 5 các control hiện thực + test cục bộ; production Level 5 cần **IdP thật/JWKS**, **WORM store (S3 Object Lock)**, KMS mặc định + HSM, **billing-API thật** (usage endpoint OpenAI/Anthropic), **dashboard deploy thật** (nginx/container + auth), **kênh alert managed + on-call**, sandbox isolation rootless/nsjail/base image CI.
|
||||
- KMS đã chạy **live qua Vault dev** ở lần chấm 2026-07-04 (đường Transit thật, khoá non-exportable); lần chạy 2026-07-05 Vault down → suite KMS **SKIP đúng thiết kế** (không tính là fail). Production thay bằng Vault/AWS-KMS/CloudHSM.
|
||||
- H6 "live" nghĩa là: webhook, provider-usage API, dashboard `/healthz` đều là **endpoint HTTP thật chạy local** (cùng chuẩn Vault-dev) — chưa phải dịch vụ hosted/managed bên ngoài.
|
||||
- Model = Ollama ornith:9b **local**; đường cloud (OpenAI/Anthropic) đã hiện thực trong `model-call.py` nhưng **chưa test bằng key thật**.
|
||||
@@ -66,12 +69,18 @@ bash .specify/tests/adversarial-harness-tests.sh # 44/0
|
||||
bash .specify/tests/phase1-track-a-tests.sh # 25/0
|
||||
bash .specify/tests/phase2-track-c-tests.sh # 29/0
|
||||
bash .specify/tests/phase3-evidence-pack-tests.sh # 7/0
|
||||
bash .specify/tests/phase-h5-approval-tests.sh # 8/0
|
||||
bash .specify/tests/phase-h5-approval-tests.sh # 12/0
|
||||
# KMS live: bật Vault dev trước để phần KMS chạy thật (không SKIP)
|
||||
docker run -d -p 8200:8200 -e VAULT_DEV_ROOT_TOKEN_ID=root hashicorp/vault
|
||||
VAULT_ADDR=http://127.0.0.1:8200 VAULT_TOKEN=root \
|
||||
bash .specify/tests/phase-h5-infra-tests.sh # 7/0 (KMS live + WORM)
|
||||
bash .specify/tests/phase-h6-agentops-tests.sh # 20/0 (alerting live + provider-API + hosted dashboard + V15)
|
||||
bash .specify/tests/phase-c7-incident-tests.sh # 15/0
|
||||
bash .specify/tests/phase-h4-multilingual-tests.sh # 7/0
|
||||
bash .specify/tests/phase-c6-sandbox-tests.sh # 6/0
|
||||
bash .specify/tests/phase-h4-split-inject-tests.sh # 8/0
|
||||
bash .specify/tests/phase10-traceability-tests.sh # 3/0
|
||||
bash .specify/tests/phase3-model-router-tests.sh # 10/0 (direct model-router/digest suite)
|
||||
bash .specify/scripts/bash/security-gate.sh # verdict PASS=11 FAIL=0
|
||||
```
|
||||
> Mục 3 là **đánh giá trưởng thành theo rubric** (người chấm, neo vào bằng chứng + gap thật), không phải output tự động của scorecard.sh (vốn chỉ đo coverage). Suite model cần Ollama live; suite KMS cần Vault live để chạy (không có thì SKIP, không tính là fail). Suite H6 **tự dựng** webhook sink / mock provider-API / dashboard server trên cổng ephemeral local — deterministic, không cần model.
|
||||
|
||||
@@ -0,0 +1,69 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Mint a mock IdP RS256 approval JWT for CASAN tests/dev.
|
||||
|
||||
The token is bound to the same high-risk request that governance-check verifies:
|
||||
sub=<approver>, role=<IdP role>, action, actor, and input_sha256.
|
||||
"""
|
||||
import argparse
|
||||
import base64
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
import time
|
||||
|
||||
|
||||
def b64u(data: bytes) -> str:
|
||||
return base64.urlsafe_b64encode(data).decode().rstrip("=")
|
||||
|
||||
|
||||
def main() -> int:
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--key", required=True)
|
||||
ap.add_argument("--sub", required=True)
|
||||
ap.add_argument("--role", required=True)
|
||||
ap.add_argument("--action", required=True)
|
||||
ap.add_argument("--actor", required=True)
|
||||
ap.add_argument("--input", required=True)
|
||||
ap.add_argument("--exp-offset", type=int, default=300)
|
||||
args = ap.parse_args()
|
||||
|
||||
with open(args.input, "rb") as f:
|
||||
input_sha = hashlib.sha256(f.read()).hexdigest()
|
||||
now = int(time.time())
|
||||
header = {"alg": "RS256", "typ": "JWT"}
|
||||
claims = {
|
||||
"iss": "casan-mock-idp",
|
||||
"sub": args.sub,
|
||||
"role": args.role,
|
||||
"action": args.action,
|
||||
"actor": args.actor,
|
||||
"input_sha256": input_sha,
|
||||
"iat": now,
|
||||
"exp": now + args.exp_offset,
|
||||
}
|
||||
signing_input = ".".join([
|
||||
b64u(json.dumps(header, separators=(",", ":"), sort_keys=True).encode()),
|
||||
b64u(json.dumps(claims, separators=(",", ":"), sort_keys=True).encode()),
|
||||
])
|
||||
with tempfile.TemporaryDirectory() as td:
|
||||
msg = os.path.join(td, "msg.txt")
|
||||
sig = os.path.join(td, "sig.bin")
|
||||
open(msg, "wb").write(signing_input.encode())
|
||||
rc = subprocess.run(
|
||||
["openssl", "dgst", "-sha256", "-sign", args.key, "-out", sig, msg],
|
||||
stdout=subprocess.DEVNULL,
|
||||
stderr=subprocess.DEVNULL,
|
||||
).returncode
|
||||
if rc != 0:
|
||||
print("approval-jwt-mint: signing failed", file=sys.stderr)
|
||||
return 1
|
||||
token = signing_input + "." + b64u(open(sig, "rb").read())
|
||||
print(token)
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -1,7 +1,7 @@
|
||||
#!/usr/bin/env bash
|
||||
set -uo pipefail
|
||||
|
||||
# CASAN H5 — Signed-approval verifier (Approval-identity MVP · C4 / V20).
|
||||
# CASAN H5 — Signed/JWT approval verifier (Approval-identity MVP · C4 / V20).
|
||||
#
|
||||
# Problem: high-risk approval used to trust a plain env var (CASAN_APPROVER=bob) —
|
||||
# anyone who can set the env can "approve". This binds an approval to a REGISTERED
|
||||
@@ -13,11 +13,14 @@ set -uo pipefail
|
||||
#
|
||||
# Usage:
|
||||
# approval-verify.sh <action> <actor> <input-file> <approver-id> <sig-file>
|
||||
# CASAN_APPROVAL_JWT=<rs256-jwt> approval-verify.sh <action> <actor> <input-file> <approver-id> -
|
||||
# Registry (line format, no yaml dep):
|
||||
# reviewer <id> <role> <pubkey-file>
|
||||
# action <action-name|default> <comma,roles>
|
||||
# Env: CASAN_REVIEWERS_FILE (default governance/reviewers.registry)
|
||||
# CASAN_REVIEWERS_DIR (default governance/reviewers) — base dir for pubkey-file
|
||||
# CASAN_APPROVAL_JWT (optional RS256 IdP token)
|
||||
# CASAN_IDP_PUBLIC_KEY (default central-governance/idp-public.pem)
|
||||
# Exit: 0 ok (prints "APPROVAL_OK role=<role>"), 3 deny (reason on stderr), 64 usage.
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
@@ -27,7 +30,7 @@ REVIEWERS_FILE="${CASAN_REVIEWERS_FILE:-$GOV_DIR/reviewers.registry}"
|
||||
REVIEWERS_DIR="${CASAN_REVIEWERS_DIR:-$GOV_DIR/reviewers}"
|
||||
|
||||
ACTION="${1:-}"; ACTOR="${2:-}"; INPUT_FILE="${3:-}"; APPROVER="${4:-}"; SIG_FILE="${5:-}"
|
||||
if [[ -z "$ACTION" || -z "$ACTOR" || -z "$INPUT_FILE" || -z "$APPROVER" || -z "$SIG_FILE" ]]; then
|
||||
if [[ -z "$ACTION" || -z "$ACTOR" || -z "$INPUT_FILE" || -z "$APPROVER" ]]; then
|
||||
echo "Usage: approval-verify.sh <action> <actor> <input-file> <approver-id> <sig-file>" >&2
|
||||
exit 64
|
||||
fi
|
||||
@@ -35,7 +38,6 @@ fi
|
||||
deny() { echo "APPROVAL_DENIED reason=$1 approver=$APPROVER action=$ACTION" >&2; exit 3; }
|
||||
|
||||
[[ -f "$INPUT_FILE" ]] || deny "input_file_missing"
|
||||
[[ -f "$SIG_FILE" ]] || deny "signature_missing"
|
||||
[[ -f "$REVIEWERS_FILE" ]] || deny "reviewer_registry_missing"
|
||||
command -v openssl >/dev/null 2>&1 || deny "openssl_unavailable"
|
||||
|
||||
@@ -44,15 +46,89 @@ hash_file() {
|
||||
else shasum -a 256 "$1" | awk '{print $1}'; fi
|
||||
}
|
||||
|
||||
INPUT_SHA="$(hash_file "$INPUT_FILE")"
|
||||
|
||||
# Role authorization for this action (fallback to the "default" action policy).
|
||||
ROLES="$(awk -v a="$ACTION" '$1=="action" && $2==a {print $3; exit}' "$REVIEWERS_FILE")"
|
||||
[[ -n "$ROLES" ]] || ROLES="$(awk '$1=="action" && $2=="default" {print $3; exit}' "$REVIEWERS_FILE")"
|
||||
|
||||
if [[ -n "${CASAN_APPROVAL_JWT:-}" ]]; then
|
||||
IDP_PUB="${CASAN_IDP_PUBLIC_KEY:-$GOV_DIR/idp-public.pem}"
|
||||
[[ -f "$IDP_PUB" ]] || deny "idp_pubkey_missing($IDP_PUB)"
|
||||
JWT_OUT="$(
|
||||
python3 - "$CASAN_APPROVAL_JWT" "$IDP_PUB" "$APPROVER" "$ROLES" "$ACTION" "$ACTOR" "$INPUT_SHA" <<'PY'
|
||||
import base64
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
import time
|
||||
|
||||
jwt, pub, approver, roles_csv, action, actor, input_sha = sys.argv[1:8]
|
||||
|
||||
def die(reason):
|
||||
print(reason, file=sys.stderr)
|
||||
sys.exit(3)
|
||||
|
||||
def b64u_decode(part):
|
||||
try:
|
||||
return base64.urlsafe_b64decode(part + "=" * (-len(part) % 4))
|
||||
except Exception:
|
||||
die("jwt_base64_invalid")
|
||||
|
||||
parts = jwt.split(".")
|
||||
if len(parts) != 3:
|
||||
die("jwt_shape_invalid")
|
||||
header = json.loads(b64u_decode(parts[0]))
|
||||
claims = json.loads(b64u_decode(parts[1]))
|
||||
if header.get("alg") != "RS256":
|
||||
die("jwt_alg_not_allowed")
|
||||
if int(claims.get("exp", 0)) <= int(time.time()):
|
||||
die("jwt_expired")
|
||||
if claims.get("sub") != approver:
|
||||
die("jwt_sub_mismatch")
|
||||
role = claims.get("role", "")
|
||||
roles = [r for r in roles_csv.split(",") if r]
|
||||
if role not in roles:
|
||||
die(f"jwt_role_not_authorized(role={role} allowed={roles_csv})")
|
||||
if claims.get("action") != action:
|
||||
die("jwt_action_mismatch")
|
||||
if claims.get("actor") != actor:
|
||||
die("jwt_actor_mismatch")
|
||||
if claims.get("input_sha256") != input_sha:
|
||||
die("jwt_input_hash_mismatch")
|
||||
|
||||
sig = b64u_decode(parts[2])
|
||||
signing_input = ".".join(parts[:2]).encode()
|
||||
with tempfile.TemporaryDirectory() as td:
|
||||
sig_path = os.path.join(td, "sig.bin")
|
||||
msg_path = os.path.join(td, "msg.txt")
|
||||
open(sig_path, "wb").write(sig)
|
||||
open(msg_path, "wb").write(signing_input)
|
||||
rc = subprocess.run(
|
||||
["openssl", "dgst", "-sha256", "-verify", pub, "-signature", sig_path, msg_path],
|
||||
stdout=subprocess.DEVNULL,
|
||||
stderr=subprocess.DEVNULL,
|
||||
).returncode
|
||||
if rc != 0:
|
||||
die("jwt_signature_invalid")
|
||||
print(f"OK role={role}")
|
||||
PY
|
||||
)" || deny "oidc_jwt_invalid(${JWT_OUT:-see_stderr})"
|
||||
ROLE="${JWT_OUT#OK role=}"
|
||||
echo "APPROVAL_OK mechanism=oidc role=$ROLE approver=$APPROVER action=$ACTION"
|
||||
exit 0
|
||||
fi
|
||||
|
||||
[[ -n "$SIG_FILE" && -f "$SIG_FILE" ]] || deny "signature_missing"
|
||||
|
||||
# Reviewer lookup (first matching registered reviewer).
|
||||
REV_LINE="$(awk -v id="$APPROVER" '$1=="reviewer" && $2==id {print $3" "$4; exit}' "$REVIEWERS_FILE")"
|
||||
[[ -n "$REV_LINE" ]] || deny "approver_not_registered"
|
||||
ROLE="${REV_LINE%% *}"
|
||||
PUB_REL="${REV_LINE##* }"
|
||||
|
||||
# Role authorization for this action (fallback to the "default" action policy).
|
||||
ROLES="$(awk -v a="$ACTION" '$1=="action" && $2==a {print $3; exit}' "$REVIEWERS_FILE")"
|
||||
[[ -n "$ROLES" ]] || ROLES="$(awk '$1=="action" && $2=="default" {print $3; exit}' "$REVIEWERS_FILE")"
|
||||
case ",$ROLES," in
|
||||
*",$ROLE,"*) : ;;
|
||||
*) deny "approver_role_not_authorized(role=$ROLE action=$ACTION allowed=$ROLES)" ;;
|
||||
@@ -63,7 +139,6 @@ PUB="$PUB_REL"; [[ "$PUB" = /* ]] || PUB="$REVIEWERS_DIR/$PUB_REL"
|
||||
[[ -f "$PUB" ]] || deny "approver_pubkey_missing($PUB)"
|
||||
|
||||
# Rebuild the exact signed assertion and verify.
|
||||
INPUT_SHA="$(hash_file "$INPUT_FILE")"
|
||||
MSG="casan-approval|v1|$ACTION|$ACTOR|$INPUT_SHA|$APPROVER"
|
||||
TMP="$(mktemp)"; trap 'rm -f "$TMP"' EXIT
|
||||
printf '%s' "$MSG" > "$TMP"
|
||||
|
||||
@@ -18,6 +18,7 @@ skipped (missing evidence => not certified, with the reason recorded).
|
||||
import hashlib
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
|
||||
|
||||
@@ -75,6 +76,25 @@ def main():
|
||||
"judge_gate_tests": os.path.exists(os.path.join(root, ".specify/tests/phase3-judge-gate-tests.sh")),
|
||||
"note": "judge-gate fail-before/fix cycle proven by phase3-judge-gate-tests.sh",
|
||||
}
|
||||
traceability_out = os.path.join(pack_dir, "traceability-matrix.json")
|
||||
traceability_script = os.path.join(root, ".specify/scripts/bash/traceability-matrix.py")
|
||||
traceability_rc = 1
|
||||
if os.path.isfile(traceability_script):
|
||||
traceability_rc = subprocess.run(
|
||||
[
|
||||
sys.executable,
|
||||
traceability_script,
|
||||
"--requirements",
|
||||
os.path.join(root, "docs/input/okr-requirement.md"),
|
||||
"--map",
|
||||
os.path.join(root, ".specify/traceability-map.json"),
|
||||
"--out",
|
||||
traceability_out,
|
||||
"--gate",
|
||||
],
|
||||
stdout=subprocess.DEVNULL,
|
||||
stderr=subprocess.DEVNULL,
|
||||
).returncode
|
||||
|
||||
# H4 security
|
||||
sec_rows = read_jsonl(os.path.join(logs, "audit", "security.jsonl"))
|
||||
@@ -151,6 +171,8 @@ def main():
|
||||
reasons.append("unresolved_cost_spike")
|
||||
if len(sec_rows) == 0:
|
||||
reasons.append("h4_security_not_exercised")
|
||||
if traceability_rc != 0:
|
||||
reasons.append("traceability_gate_failed")
|
||||
fp_ok = bool(fp) and fp.get("within_budget") is True
|
||||
if fp is None:
|
||||
reasons.append("benign_fp_report_missing(gate_skipped)")
|
||||
@@ -162,7 +184,7 @@ def main():
|
||||
run_summary = {
|
||||
"run_id": run_id, "pack_version": "1.0-mvp",
|
||||
"certified": certified, "certification_reasons": reasons or ["all_required_gates_passed"],
|
||||
"required_gates": ["H4-security", "H5-audit-chain", "H5-telemetry", "H6-cost", "benign-fp-budget"],
|
||||
"required_gates": ["H3-traceability", "H4-security", "H5-audit-chain", "H5-telemetry", "H6-cost", "benign-fp-budget"],
|
||||
"harness_reports": sorted(reports.keys()),
|
||||
}
|
||||
with open(os.path.join(pack_dir, "run-summary.json"), "w", encoding="utf-8") as f:
|
||||
@@ -182,6 +204,7 @@ def main():
|
||||
f"{total_tokens} provider tokens",
|
||||
f"- H2 tool audit: {reports['h2-tool-audit.json']['records']} records, "
|
||||
f"chain_ok={reports['h2-tool-audit.json']['chain_ok']}",
|
||||
f"- H3 traceability: ok={traceability_rc == 0}",
|
||||
f"- Red-team: {reports['redteam-result.json']['vectors_defined']} vectors "
|
||||
f"(block_rate={reports['redteam-result.json']['adversarial_block_rate_pct']}%)",
|
||||
"",
|
||||
|
||||
@@ -92,16 +92,20 @@ if [[ "$RISK_LEVEL" == "high" ]]; then
|
||||
# Approval-identity mode (V20): an env-var approver is NOT enough — the
|
||||
# reviewer must cryptographically SIGN this exact request and their role must
|
||||
# be authorized for the action. SoD (actor != approver) still enforced.
|
||||
if [[ "$APPROVAL_DECISION" == "approve" && -n "$APPROVER" && -n "${CASAN_APPROVAL_SIG:-}" ]]; then
|
||||
if [[ "$APPROVAL_DECISION" == "approve" && -n "$APPROVER" && ( -n "${CASAN_APPROVAL_SIG:-}" || -n "${CASAN_APPROVAL_JWT:-}" ) ]]; then
|
||||
if [[ "$APPROVER" == "$ACTOR" ]]; then
|
||||
APPROVAL_STATUS="separation_of_duties_violation"
|
||||
DECISION="denied"
|
||||
REASONS+=("separation-of-duties:actor-equals-approver")
|
||||
else
|
||||
AV_RC=0
|
||||
AV_OUT="$(bash "$SCRIPT_DIR/approval-verify.sh" "$ACTION_NAME" "$ACTOR" "$INPUT_FILE" "$APPROVER" "$CASAN_APPROVAL_SIG" 2>/dev/null)" || AV_RC=$?
|
||||
AV_OUT="$(bash "$SCRIPT_DIR/approval-verify.sh" "$ACTION_NAME" "$ACTOR" "$INPUT_FILE" "$APPROVER" "${CASAN_APPROVAL_SIG:-"-"}" 2>/dev/null)" || AV_RC=$?
|
||||
if [[ "$AV_RC" -eq 0 ]]; then
|
||||
if printf '%s' "$AV_OUT" | grep -q "mechanism=oidc"; then
|
||||
APPROVAL_STATUS="human_approved_oidc"
|
||||
else
|
||||
APPROVAL_STATUS="human_approved_signed"
|
||||
fi
|
||||
DECISION="approved"
|
||||
REASONS+=("signed-approval:${AV_OUT#APPROVAL_OK }")
|
||||
else
|
||||
|
||||
@@ -23,6 +23,7 @@ Usage:
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import subprocess
|
||||
import sys
|
||||
import time
|
||||
import urllib.request
|
||||
@@ -89,6 +90,17 @@ def call_ollama(model_name, prompt, role):
|
||||
host = os.environ.get("CASAN_OLLAMA_HOST", OLLAMA_HOST)
|
||||
if host != OLLAMA_HOST:
|
||||
fail(f"endpoint_not_allowed ollama host={host} (only {OLLAMA_HOST})")
|
||||
digest_gate = os.path.join(os.path.dirname(__file__), "model-digest-check.sh")
|
||||
if os.path.isfile(digest_gate):
|
||||
check = subprocess.run(
|
||||
["bash", digest_gate, "verify", model_name],
|
||||
text=True,
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.PIPE,
|
||||
)
|
||||
if check.returncode != 0:
|
||||
msg = (check.stderr or check.stdout or "model_digest_check_failed").strip()
|
||||
fail(msg)
|
||||
url = f"http://{host}/api/generate"
|
||||
body = {
|
||||
"model": model_name,
|
||||
|
||||
@@ -16,8 +16,9 @@ set -uo pipefail
|
||||
# model-digest-check.sh verify [model] # compare live digest to the pinned one
|
||||
# model-digest-check.sh show [model]
|
||||
# Env: CASAN_MODEL (default ornith:9b) · CASAN_MODEL_DIGEST (override) ·
|
||||
# CASAN_MODEL_DIGEST_PIN (pin file) · OLLAMA_HOST (default 127.0.0.1:11434)
|
||||
# Exit: 0 match/pinned · 2 MISMATCH (swap detected) · 3 unpinned/undeterminable.
|
||||
# CASAN_MODEL_DIGEST_PIN (pin file) · CASAN_MODEL_DIGEST_MODE=block|warn ·
|
||||
# OLLAMA_HOST (default 127.0.0.1:11434)
|
||||
# Exit: 0 match/pinned/warned · 2 MISMATCH in block mode · 3 unpinned/undeterminable.
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
PROJECT_ROOT="$(cd "$SCRIPT_DIR/../../.." && pwd)"
|
||||
@@ -60,6 +61,10 @@ case "$CMD" in
|
||||
echo "MODEL_DIGEST_OK model=$MODEL digest=${D:0:24}…"
|
||||
exit 0
|
||||
fi
|
||||
if [[ "${CASAN_MODEL_DIGEST_MODE:-block}" == "warn" ]]; then
|
||||
echo "MODEL_DIGEST_WARN model=$MODEL pinned=${P:0:24}… live=${D:0:24}… — model may have been swapped/poisoned" >&2
|
||||
exit 0
|
||||
fi
|
||||
echo "MODEL_DIGEST_MISMATCH model=$MODEL pinned=${P:0:24}… live=${D:0:24}… — model may have been swapped/poisoned" >&2
|
||||
exit 2
|
||||
;;
|
||||
|
||||
@@ -0,0 +1,115 @@
|
||||
#!/usr/bin/env python3
|
||||
"""CASAN Plan-10 traceability matrix generator/gate.
|
||||
|
||||
Parses FR-* requirements from docs/input/okr-requirement.md and checks each
|
||||
requirement has at least one existing code file and one existing test file.
|
||||
"""
|
||||
import argparse
|
||||
import json
|
||||
import os
|
||||
import re
|
||||
import sys
|
||||
from datetime import datetime, timezone
|
||||
|
||||
|
||||
FR_RE = re.compile(r"\|\s*(FR-\d+)\s*\|\s*([^|]+?)\s*\|")
|
||||
|
||||
|
||||
def project_root() -> str:
|
||||
return os.path.abspath(os.path.join(os.path.dirname(__file__), "..", "..", ".."))
|
||||
|
||||
|
||||
def parse_requirements(path: str):
|
||||
seen = {}
|
||||
with open(path, encoding="utf-8") as f:
|
||||
for line in f:
|
||||
match = FR_RE.search(line)
|
||||
if not match:
|
||||
continue
|
||||
fr_id, name = match.groups()
|
||||
if fr_id not in seen:
|
||||
seen[fr_id] = {"id": fr_id, "name": " ".join(name.split())}
|
||||
return [seen[k] for k in sorted(seen)]
|
||||
|
||||
|
||||
def existing_files(root: str, values):
|
||||
present, missing = [], []
|
||||
for rel in values or []:
|
||||
if os.path.isfile(os.path.join(root, rel)):
|
||||
present.append(rel)
|
||||
else:
|
||||
missing.append(rel)
|
||||
return present, missing
|
||||
|
||||
|
||||
def main() -> int:
|
||||
root = project_root()
|
||||
ap = argparse.ArgumentParser()
|
||||
ap.add_argument("--requirements", default=os.path.join(root, "docs/input/okr-requirement.md"))
|
||||
ap.add_argument("--map", default=os.path.join(root, ".specify/traceability-map.json"))
|
||||
ap.add_argument("--out", default=os.path.join(root, "docs/output/casan/traceability-matrix.json"))
|
||||
ap.add_argument("--gate", action="store_true")
|
||||
args = ap.parse_args()
|
||||
|
||||
reqs = parse_requirements(args.requirements)
|
||||
with open(args.map, encoding="utf-8") as f:
|
||||
mapping = json.load(f)
|
||||
|
||||
rows = []
|
||||
failures = []
|
||||
for req in reqs:
|
||||
entry = mapping.get(req["id"], {})
|
||||
code, missing_code = existing_files(root, entry.get("code", []))
|
||||
tests, missing_tests = existing_files(root, entry.get("tests", []))
|
||||
status = "PASS" if code and tests and not missing_code and not missing_tests else "FAIL"
|
||||
row = {
|
||||
"id": req["id"],
|
||||
"name": req["name"],
|
||||
"status": status,
|
||||
"code": code,
|
||||
"tests": tests,
|
||||
"missing_code": missing_code,
|
||||
"missing_tests": missing_tests,
|
||||
}
|
||||
rows.append(row)
|
||||
if status != "PASS":
|
||||
failures.append(row)
|
||||
|
||||
orphan_mappings = sorted(set(mapping) - {r["id"] for r in reqs})
|
||||
out = {
|
||||
"generated_at": datetime.now(timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ"),
|
||||
"requirements_source": os.path.relpath(args.requirements, root),
|
||||
"mapping_source": os.path.relpath(args.map, root),
|
||||
"summary": {
|
||||
"requirements": len(reqs),
|
||||
"passed": sum(1 for r in rows if r["status"] == "PASS"),
|
||||
"failed": len(failures),
|
||||
"orphan_mappings": orphan_mappings,
|
||||
},
|
||||
"matrix": rows,
|
||||
}
|
||||
os.makedirs(os.path.dirname(args.out), exist_ok=True)
|
||||
with open(args.out, "w", encoding="utf-8") as f:
|
||||
json.dump(out, f, indent=2, ensure_ascii=False)
|
||||
f.write("\n")
|
||||
|
||||
if failures:
|
||||
for row in failures:
|
||||
print(
|
||||
f"TRACEABILITY_FAIL {row['id']} code={len(row['code'])} tests={len(row['tests'])} "
|
||||
f"missing_code={len(row['missing_code'])} missing_tests={len(row['missing_tests'])}",
|
||||
file=sys.stderr,
|
||||
)
|
||||
if orphan_mappings:
|
||||
print(f"TRACEABILITY_WARN orphan_mappings={','.join(orphan_mappings)}", file=sys.stderr)
|
||||
print(
|
||||
f"TRACEABILITY_MATRIX requirements={len(reqs)} pass={out['summary']['passed']} "
|
||||
f"fail={len(failures)} out={args.out}"
|
||||
)
|
||||
if args.gate and failures:
|
||||
return 1
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1 @@
|
||||
ornith:9b a75697c145891910e312c95e4a9fc1ccb8653e5ef543b23b0403a4665b82fd91
|
||||
@@ -114,7 +114,7 @@ tr:nth-child(even) td {{ background: #f7f9fc; }}
|
||||
<span class="badge lv">CASAN Level 4 — chứng minh bằng tấn công</span>
|
||||
<span class="badge">Average {avg_score}/100</span>
|
||||
<span class="badge">Harness thấp nhất {lowest_score}</span>
|
||||
<span class="badge">175 test · 0 fail</span>
|
||||
<span class="badge">218 core tests · 0 fail</span>
|
||||
<span class="badge">Recall model 0.85 > regex 0.00</span>
|
||||
</div>
|
||||
|
||||
@@ -132,7 +132,7 @@ tr:nth-child(even) td {{ background: #f7f9fc; }}
|
||||
|
||||
<h2>Bảo mật & Governance đã kiểm chứng (test đối kháng thật)</h2>
|
||||
<div class="panel"><div class="chips">
|
||||
<span class="chip"><span class="d"></span>Kiểm thử đối kháng <b>175 / 0 fail</b></span>
|
||||
<span class="chip"><span class="d"></span>Kiểm thử đối kháng <b>218 / 0 fail</b></span>
|
||||
<span class="chip"><span class="d"></span>H4 recall model <b>0.85</b> > regex 0.00</span>
|
||||
<span class="chip"><span class="d"></span>Benign FP <b>0.00%</b> · block <b>100.00%</b></span>
|
||||
<span class="chip"><span class="d"></span>Audit hash-chain + ký KMS (rotate/non-exportable)</span>
|
||||
|
||||
@@ -94,6 +94,38 @@ expect_rc 0 "non-strict env approval unchanged (backward compatible)" \
|
||||
env CASAN_ACTOR=alice CASAN_APPROVAL_DECISION=approve CASAN_APPROVER=bob \
|
||||
bash "$S/governance-check.sh" "$REQ" "$WORK/out.txt" deploy
|
||||
|
||||
echo "===== H5 OIDC approval via mock IdP JWT (CASAN_APPROVAL_STRICT=1) ====="
|
||||
openssl genrsa -out "$WORK/idp.priv.pem" 2048 2>/dev/null
|
||||
openssl rsa -in "$WORK/idp.priv.pem" -pubout -out "$WORK/idp.pub.pem" 2>/dev/null
|
||||
openssl genrsa -out "$WORK/fake-idp.priv.pem" 2048 2>/dev/null
|
||||
|
||||
mint_jwt() {
|
||||
python3 "$S/approval-jwt-mint.py" --key "$1" --sub "$2" --role "$3" \
|
||||
--action deploy --actor alice --input "$REQ" --exp-offset "$4"
|
||||
}
|
||||
|
||||
# 9. Valid IdP-signed JWT by an authorized role -> APPROVED
|
||||
OIDC_JWT="$(mint_jwt "$WORK/idp.priv.pem" oidc-ops ops 300)"
|
||||
OUT="$(gc CASAN_APPROVER=oidc-ops CASAN_APPROVAL_JWT="$OIDC_JWT" CASAN_IDP_PUBLIC_KEY="$WORK/idp.pub.pem" 2>/dev/null)"; RC=$?
|
||||
{ [[ "$RC" -eq 0 ]] && printf '%s' "$OUT" | grep -q "human_approved_oidc"; } \
|
||||
&& pass "valid IdP JWT by authorized role -> APPROVED" \
|
||||
|| fail "valid IdP JWT rejected (rc=$RC out=$OUT)"
|
||||
|
||||
# 10. Expired JWT -> DENY
|
||||
EXPIRED_JWT="$(mint_jwt "$WORK/idp.priv.pem" oidc-ops ops -60)"
|
||||
expect_rc 2 "expired IdP JWT is denied" \
|
||||
gc CASAN_APPROVER=oidc-ops CASAN_APPROVAL_JWT="$EXPIRED_JWT" CASAN_IDP_PUBLIC_KEY="$WORK/idp.pub.pem"
|
||||
|
||||
# 11. Role not authorized for deploy -> DENY
|
||||
WRONG_ROLE_JWT="$(mint_jwt "$WORK/idp.priv.pem" oidc-tech tech_lead 300)"
|
||||
expect_rc 2 "IdP JWT with unauthorized role is denied" \
|
||||
gc CASAN_APPROVER=oidc-tech CASAN_APPROVAL_JWT="$WRONG_ROLE_JWT" CASAN_IDP_PUBLIC_KEY="$WORK/idp.pub.pem"
|
||||
|
||||
# 12. JWT signed by a different key than the trusted IdP pubkey -> DENY
|
||||
FORGED_JWT="$(mint_jwt "$WORK/fake-idp.priv.pem" oidc-ops ops 300)"
|
||||
expect_rc 2 "JWT with forged IdP signature is denied" \
|
||||
gc CASAN_APPROVER=oidc-ops CASAN_APPROVAL_JWT="$FORGED_JWT" CASAN_IDP_PUBLIC_KEY="$WORK/idp.pub.pem"
|
||||
|
||||
echo ""
|
||||
echo "===== H5 APPROVAL-IDENTITY SUMMARY: PASS=$PASS FAIL=$FAIL ====="
|
||||
[[ "$FAIL" -eq 0 ]] || exit 1
|
||||
|
||||
@@ -0,0 +1,54 @@
|
||||
#!/usr/bin/env bash
|
||||
set -uo pipefail
|
||||
|
||||
# CASAN Plan-10 — Traceability REQ→code→test MVP tests.
|
||||
#
|
||||
# Proves the matrix is generated from the real requirement document and the gate
|
||||
# fails when any FR loses test coverage.
|
||||
|
||||
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
|
||||
PROJECT_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)"
|
||||
S="$PROJECT_ROOT/.specify/scripts/bash"
|
||||
WORK="$(mktemp -d)"
|
||||
trap 'rm -rf "$WORK"' EXIT
|
||||
|
||||
PASS=0; FAIL=0
|
||||
pass() { echo "PASS: $1"; PASS=$((PASS + 1)); }
|
||||
fail() { echo "FAIL: $1"; FAIL=$((FAIL + 1)); }
|
||||
|
||||
REQ="$PROJECT_ROOT/docs/input/okr-requirement.md"
|
||||
MAP="$PROJECT_ROOT/.specify/traceability-map.json"
|
||||
OUT="$WORK/traceability-matrix.json"
|
||||
|
||||
echo "===== Plan-10 traceability matrix ====="
|
||||
if python3 "$S/traceability-matrix.py" --requirements "$REQ" --map "$MAP" --out "$OUT" --gate >/dev/null 2>"$WORK/pass.err"; then
|
||||
pass "complete FR→code→test matrix passes gate"
|
||||
else
|
||||
cat "$WORK/pass.err"
|
||||
fail "complete traceability matrix rejected"
|
||||
fi
|
||||
|
||||
COUNT="$(python3 -c "import json;d=json.load(open('$OUT'));print(d['summary']['requirements'])" 2>/dev/null)"
|
||||
FAILED="$(python3 -c "import json;d=json.load(open('$OUT'));print(d['summary']['failed'])" 2>/dev/null)"
|
||||
[[ "$COUNT" == "5" && "$FAILED" == "0" ]] \
|
||||
&& pass "matrix captures all 5 FRs with zero failures" \
|
||||
|| fail "matrix summary unexpected (requirements=$COUNT failed=$FAILED)"
|
||||
|
||||
BROKEN="$WORK/traceability-map-broken.json"
|
||||
python3 - "$MAP" "$BROKEN" <<'PY'
|
||||
import json, sys
|
||||
d = json.load(open(sys.argv[1]))
|
||||
d["FR-04"]["tests"] = []
|
||||
json.dump(d, open(sys.argv[2], "w"), indent=2)
|
||||
PY
|
||||
set +e
|
||||
python3 "$S/traceability-matrix.py" --requirements "$REQ" --map "$BROKEN" --out "$WORK/broken.json" --gate >/dev/null 2>"$WORK/broken.err"
|
||||
RC=$?
|
||||
set -e 2>/dev/null || true
|
||||
[[ "$RC" -eq 1 ]] && grep -q "TRACEABILITY_FAIL FR-04" "$WORK/broken.err" \
|
||||
&& pass "missing FR test coverage fails the gate" \
|
||||
|| fail "missing test coverage did not fail as expected (rc=$RC)"
|
||||
|
||||
echo ""
|
||||
echo "===== TRACEABILITY SUMMARY: PASS=$PASS FAIL=$FAIL ====="
|
||||
[[ "$FAIL" -eq 0 ]] || exit 1
|
||||
@@ -41,10 +41,10 @@ MISSING=0
|
||||
for f in run-summary.json h1-context-report.json h2-tool-audit.json h3-eval-scorecard.json \
|
||||
h4-security-report.json h5-audit-chain-proof.json h6-cost-telemetry.json \
|
||||
h7-orchestration-report.json redteam-result.json benign-fp-report.json \
|
||||
artifact-manifest.json decision-log.md; do
|
||||
artifact-manifest.json traceability-matrix.json decision-log.md; do
|
||||
[[ -f "$PACKDIR/$f" ]] || { echo " missing $f"; MISSING=$((MISSING+1)); }
|
||||
done
|
||||
[[ "$MISSING" -eq 0 ]] && pass "pack contains all 12 standard evidence files" || fail "pack missing $MISSING files"
|
||||
[[ "$MISSING" -eq 0 ]] && pass "pack contains all 13 standard evidence files" || fail "pack missing $MISSING files"
|
||||
[[ -f "$PACKDIR/evidence-pack.sig" ]] && pass "pack is signed (evidence-pack.sig present)" || fail "pack signature missing"
|
||||
|
||||
rc=0; bash "$EP" verify-pack "$RID" >/dev/null 2>&1 || rc=$?
|
||||
|
||||
@@ -66,7 +66,28 @@ else
|
||||
pass "no API-key/secret pattern in .specify/logs"
|
||||
fi
|
||||
|
||||
# 6: cloud backend reports unavailable honestly while keys are unset.
|
||||
# 6: model digest pinning catches a silent model swap, with warn mode available.
|
||||
PIN="$WORK/model-digest.pin"
|
||||
CASAN_MODEL_DIGEST_PIN="$PIN" CASAN_MODEL_DIGEST="digest-approved" bash "$SCRIPTS/model-digest-check.sh" pin ornith:9b >/dev/null 2>&1
|
||||
if CASAN_MODEL_DIGEST_PIN="$PIN" CASAN_MODEL_DIGEST="digest-approved" bash "$SCRIPTS/model-digest-check.sh" verify ornith:9b >/dev/null 2>&1; then
|
||||
pass "model digest pin verifies approved digest"
|
||||
else
|
||||
fail "model digest pin rejected matching digest"
|
||||
fi
|
||||
set +e
|
||||
CASAN_MODEL_DIGEST_PIN="$PIN" CASAN_MODEL_DIGEST="digest-swapped" bash "$SCRIPTS/model-digest-check.sh" verify ornith:9b >/dev/null 2>"$WORK/digest.err"
|
||||
DRC=$?
|
||||
CASAN_MODEL_DIGEST_PIN="$PIN" CASAN_MODEL_DIGEST="digest-swapped" CASAN_MODEL_DIGEST_MODE=warn bash "$SCRIPTS/model-digest-check.sh" verify ornith:9b >/dev/null 2>"$WORK/digest-warn.err"
|
||||
WRC=$?
|
||||
set -e 2>/dev/null || true
|
||||
[[ "$DRC" -eq 2 ]] && grep -q "MODEL_DIGEST_MISMATCH" "$WORK/digest.err" \
|
||||
&& pass "model digest mismatch blocks by default" \
|
||||
|| fail "model digest mismatch did not block (rc=$DRC)"
|
||||
[[ "$WRC" -eq 0 ]] && grep -q "MODEL_DIGEST_WARN" "$WORK/digest-warn.err" \
|
||||
&& pass "model digest mismatch can warn during rollout" \
|
||||
|| fail "model digest warn mode did not warn cleanly (rc=$WRC)"
|
||||
|
||||
# 7: cloud backend reports unavailable honestly while keys are unset.
|
||||
if [[ -z "${ANTHROPIC_API_KEY:-}" ]]; then
|
||||
set +e
|
||||
bash "$ROUTER" "$WORK/s.txt" "$WORK/cloud.json" --role classify --model anthropic:claude-opus-4-8 >/dev/null 2>"$WORK/cloud.err"
|
||||
@@ -76,7 +97,7 @@ else
|
||||
skip "cloud-unavailable test (ANTHROPIC_API_KEY is set)"
|
||||
fi
|
||||
|
||||
# 7: deliberate failing primary route -> fallback through the REAL router (not exit 9).
|
||||
# 8: deliberate failing primary route -> fallback through the REAL router (not exit 9).
|
||||
if [[ "$TUNNEL_UP" -eq 1 ]]; then
|
||||
printf 'Return exactly: OK\n' > "$WORK/f.txt"
|
||||
set +e
|
||||
|
||||
@@ -0,0 +1,69 @@
|
||||
{
|
||||
"FR-01": {
|
||||
"name": "Login",
|
||||
"code": [
|
||||
"backend/src/auth/auth.controller.ts",
|
||||
"backend/src/auth/auth.service.ts",
|
||||
"frontend/src/pages/Login.tsx",
|
||||
"frontend/src/hooks/useAuth.tsx"
|
||||
],
|
||||
"tests": [
|
||||
"backend/test/services.test.ts",
|
||||
"backend/test/e2e.test.ts"
|
||||
]
|
||||
},
|
||||
"FR-02": {
|
||||
"name": "Create Objective",
|
||||
"code": [
|
||||
"backend/src/objectives/objectives.controller.ts",
|
||||
"backend/src/objectives/objectives.service.ts",
|
||||
"frontend/src/pages/CreateObjective.tsx",
|
||||
"frontend/src/schemas/objective.schema.ts"
|
||||
],
|
||||
"tests": [
|
||||
"backend/test/services.test.ts",
|
||||
"backend/test/e2e.test.ts",
|
||||
"frontend/src/__tests__/okr.test.tsx"
|
||||
]
|
||||
},
|
||||
"FR-03": {
|
||||
"name": "Create Key Result",
|
||||
"code": [
|
||||
"backend/src/key-results/key-results.controller.ts",
|
||||
"backend/src/key-results/key-results.service.ts",
|
||||
"backend/src/key-results/dto/create-key-result.dto.ts"
|
||||
],
|
||||
"tests": [
|
||||
"backend/test/services.test.ts",
|
||||
"backend/test/e2e.test.ts"
|
||||
]
|
||||
},
|
||||
"FR-04": {
|
||||
"name": "Update Progress",
|
||||
"code": [
|
||||
"backend/src/key-results/key-results.controller.ts",
|
||||
"backend/src/key-results/key-results.service.ts",
|
||||
"backend/src/key-results/dto/update-progress.dto.ts",
|
||||
"frontend/src/pages/KeyResultDetail.tsx"
|
||||
],
|
||||
"tests": [
|
||||
"backend/test/services.test.ts",
|
||||
"backend/test/e2e.test.ts",
|
||||
"frontend/src/__tests__/okr.test.tsx"
|
||||
]
|
||||
},
|
||||
"FR-05": {
|
||||
"name": "Dashboard",
|
||||
"code": [
|
||||
"backend/src/objectives/objectives.controller.ts",
|
||||
"backend/src/objectives/objectives.service.ts",
|
||||
"frontend/src/pages/Dashboard.tsx",
|
||||
"frontend/src/hooks/useObjectives.ts"
|
||||
],
|
||||
"tests": [
|
||||
"backend/test/services.test.ts",
|
||||
"backend/test/e2e.test.ts",
|
||||
"frontend/src/__tests__/okr.test.tsx"
|
||||
]
|
||||
}
|
||||
}
|
||||
@@ -0,0 +1,100 @@
|
||||
{
|
||||
"generated_at": "2026-07-06T02:24:50Z",
|
||||
"requirements_source": "docs/input/okr-requirement.md",
|
||||
"mapping_source": ".specify/traceability-map.json",
|
||||
"summary": {
|
||||
"requirements": 5,
|
||||
"passed": 5,
|
||||
"failed": 0,
|
||||
"orphan_mappings": []
|
||||
},
|
||||
"matrix": [
|
||||
{
|
||||
"id": "FR-01",
|
||||
"name": "Login",
|
||||
"status": "PASS",
|
||||
"code": [
|
||||
"backend/src/auth/auth.controller.ts",
|
||||
"backend/src/auth/auth.service.ts",
|
||||
"frontend/src/pages/Login.tsx",
|
||||
"frontend/src/hooks/useAuth.tsx"
|
||||
],
|
||||
"tests": [
|
||||
"backend/test/services.test.ts",
|
||||
"backend/test/e2e.test.ts"
|
||||
],
|
||||
"missing_code": [],
|
||||
"missing_tests": []
|
||||
},
|
||||
{
|
||||
"id": "FR-02",
|
||||
"name": "Create Objective",
|
||||
"status": "PASS",
|
||||
"code": [
|
||||
"backend/src/objectives/objectives.controller.ts",
|
||||
"backend/src/objectives/objectives.service.ts",
|
||||
"frontend/src/pages/CreateObjective.tsx",
|
||||
"frontend/src/schemas/objective.schema.ts"
|
||||
],
|
||||
"tests": [
|
||||
"backend/test/services.test.ts",
|
||||
"backend/test/e2e.test.ts",
|
||||
"frontend/src/__tests__/okr.test.tsx"
|
||||
],
|
||||
"missing_code": [],
|
||||
"missing_tests": []
|
||||
},
|
||||
{
|
||||
"id": "FR-03",
|
||||
"name": "Create Key Result",
|
||||
"status": "PASS",
|
||||
"code": [
|
||||
"backend/src/key-results/key-results.controller.ts",
|
||||
"backend/src/key-results/key-results.service.ts",
|
||||
"backend/src/key-results/dto/create-key-result.dto.ts"
|
||||
],
|
||||
"tests": [
|
||||
"backend/test/services.test.ts",
|
||||
"backend/test/e2e.test.ts"
|
||||
],
|
||||
"missing_code": [],
|
||||
"missing_tests": []
|
||||
},
|
||||
{
|
||||
"id": "FR-04",
|
||||
"name": "Update Progress",
|
||||
"status": "PASS",
|
||||
"code": [
|
||||
"backend/src/key-results/key-results.controller.ts",
|
||||
"backend/src/key-results/key-results.service.ts",
|
||||
"backend/src/key-results/dto/update-progress.dto.ts",
|
||||
"frontend/src/pages/KeyResultDetail.tsx"
|
||||
],
|
||||
"tests": [
|
||||
"backend/test/services.test.ts",
|
||||
"backend/test/e2e.test.ts",
|
||||
"frontend/src/__tests__/okr.test.tsx"
|
||||
],
|
||||
"missing_code": [],
|
||||
"missing_tests": []
|
||||
},
|
||||
{
|
||||
"id": "FR-05",
|
||||
"name": "Dashboard",
|
||||
"status": "PASS",
|
||||
"code": [
|
||||
"backend/src/objectives/objectives.controller.ts",
|
||||
"backend/src/objectives/objectives.service.ts",
|
||||
"frontend/src/pages/Dashboard.tsx",
|
||||
"frontend/src/hooks/useObjectives.ts"
|
||||
],
|
||||
"tests": [
|
||||
"backend/test/services.test.ts",
|
||||
"backend/test/e2e.test.ts",
|
||||
"frontend/src/__tests__/okr.test.tsx"
|
||||
],
|
||||
"missing_code": [],
|
||||
"missing_tests": []
|
||||
}
|
||||
]
|
||||
}
|
||||
@@ -4,14 +4,15 @@
|
||||
> bước tiếp theo cụ thể + cờ phụ-thuộc-hạ-tầng, để **bất kỳ AI/người nào tiếp quản
|
||||
> cũng làm tiếp được ngay**. Cập nhật mỗi khi hoàn thành một mục.
|
||||
>
|
||||
> Cập nhật lần cuối: 2026-07-06 · Nhánh: `feat/plan07-track-a-hardening` (đã merge vào `main` @ `c943273`).
|
||||
> Test hiện tại: **211 PASS / 0 FAIL** trên 12 suite. Điểm công tâm ~81.6/100, harness thấp nhất 80.
|
||||
> Cập nhật lần cuối: 2026-07-06 · Nhánh làm tiếp từ handoff Claude.
|
||||
> Test hiện tại: **218 PASS / 0 FAIL** trên 13 core harness suite; `phase3-model-router` riêng **10 PASS / 0 FAIL**; frontend Vitest **16 PASS / 0 FAIL**. Backend `npm test` còn bị chặn bởi test-infra cũ (`schema.prisma` MySQL nhưng `setup-sqlite.mjs` chạy SQLite).
|
||||
> Điểm công tâm vẫn quanh **~81/100**, harness thấp nhất 80; TIER 2 infra thật vẫn là trần Strong.
|
||||
> Nguồn liên quan: `CASAN_HARDENING_STATUS.md` (chi tiết control) · `evidence/scoring-run-report.md` (điểm).
|
||||
|
||||
## Quy ước
|
||||
- **Nhãn:** ✅ done+test · 🟡 MVP done (bản prod cần hạ tầng) · 🟦 deliverable-now (làm được offline) · 🔌 needs-infra (cần key/dịch vụ ngoài) · 📋 handoff-only (platform lớn, cần nhiều phiên).
|
||||
- **Nguyên tắc bàn giao:** mỗi mục 🔌/📋 phải có "Bước tiếp theo" đủ cụ thể để người khác bắt tay ngay.
|
||||
- **Bất biến an toàn:** giữ 211/0; mọi control mới phải có test đối kháng fail-able; không hardcode verdict, không bypass.
|
||||
- **Bất biến an toàn:** giữ 218/0 core harness; mọi control mới phải có test đối kháng fail-able; không hardcode verdict, không bypass.
|
||||
|
||||
---
|
||||
|
||||
@@ -19,9 +20,9 @@
|
||||
|
||||
| # | Hạng mục | Plan | Trạng thái | Bước tiếp theo |
|
||||
|---|---|---|:--:|---|
|
||||
| T1.1 | **Model-digest pinning** (chống model bị tráo/poison) | 07 B4 / V16 | 🟦 → (làm trong phiên này) | Ghi digest model đã duyệt vào `security/model-digest.pin`; `model-router.sh`/1 gate so digest hiện tại; đổi digest bất ngờ → cảnh báo/lock. Test: pin→match=OK, đổi→WARN/BLOCK. |
|
||||
| T1.2 | **IdP/OIDC cho approval** (thay registry pubkey tĩnh) | 07 C4 / V20 | 🟦 (mock IdP) | Verify approver bằng **JWT ký bởi IdP** (RS256): `approval-verify.sh` chấp nhận `CASAN_APPROVAL_JWT` — verify chữ ký bằng JWKS/pubkey, check claim `sub`(=approver)/`role`/`exp`. MVP dùng khoá IdP-giả (như Vault dev); prod trỏ JWKS thật. Test: JWT hợp lệ→APPROVED, hết hạn/sai role/chữ ký giả→DENY. |
|
||||
| T1.3 | **Traceability REQ→code→test** (khác biệt nhất) | 10 | 🟦 → (làm trong phiên này) | Sinh ma trận: parse FR-xx từ `docs/input/okr-requirement.md` → map sang file code + test; gate: mọi FR phải có ≥1 code + ≥1 test, cảnh báo FR mồ côi. Test: FR đủ→PASS, FR thiếu test→FAIL. |
|
||||
| T1.1 | **Model-digest pinning** (chống model bị tráo/poison) | 07 B4 / V16 | ✅ done+test | `.specify/security/model-digest.pin` pin `ornith:9b`; `model-call.py` gọi `model-digest-check.sh verify` trước Ollama; mismatch mặc định BLOCK, rollout mode WARN. Test: `phase3-model-router` pin→OK, đổi→WARN/BLOCK. |
|
||||
| T1.2 | **IdP/OIDC cho approval** (thay registry pubkey tĩnh) | 07 C4 / V20 | ✅ MVP done+test | `approval-verify.sh` chấp nhận `CASAN_APPROVAL_JWT` RS256 ký bởi mock IdP, verify chữ ký bằng pubkey, check `sub`/`role`/`exp` + bind `action`/`actor`/`input_sha256`. Test: JWT hợp lệ→APPROVED, hết hạn/sai role/chữ ký giả→DENY. Prod còn cần IdP/JWKS thật. |
|
||||
| T1.3 | **Traceability REQ→code→test** (khác biệt nhất) | 10 | ✅ MVP done+test | `traceability-matrix.py` parse FR-xx từ requirement, dùng `.specify/traceability-map.json`, gate mọi FR có ≥1 code + ≥1 test; Evidence Pack thêm `traceability-matrix.json`. Test: FR đủ→PASS, FR thiếu test→FAIL. |
|
||||
|
||||
## TIER 2 — Plan-07 gaps cần hạ tầng (MVP đã có, cần bản prod) 🔌
|
||||
|
||||
@@ -38,7 +39,7 @@
|
||||
|
||||
| Plan | Trạng thái | Lõi cần làm (bước tiếp theo cho AI kế) |
|
||||
|---|:--:|---|
|
||||
| **10 Traceability + H3 Eval** | 🟦 MVP phiên này (xem T1.3) | Sau MVP: nối traceability vào Evidence Pack (thêm `traceability-matrix.json` vào pack); mở rộng H3 eval-set độc lập (nhiều model). |
|
||||
| **10 Traceability + H3 Eval** | ✅ MVP done+test | Đã nối traceability vào Evidence Pack. Sau MVP: line/symbol-level traceability + H3 eval-set độc lập (nhiều model). |
|
||||
| **02 LLM source-gen** | 📋 chưa bắt đầu | Thay template bằng LLM thật sinh source qua `model-router.sh`; đi qua wrapper H4→H7. Phụ thuộc 03. Bước 1: định contract prompt→file cho 1 module (objectives), gate bằng H3 judge + traceability. |
|
||||
| **03 Cloud patch** | 📋 chưa bắt đầu | Bỏ stub trong `model-call.py` cho OpenAI/Anthropic; test bằng key thật (🔌). Bước 1: env `CASAN_MODEL_PRIMARY=openai:…`, xác thực round-trip + ghi provider-usage thật. |
|
||||
| **04 Self-improve** | 📋 chưa bắt đầu | Khép vòng `casan improve`: đọc metrics/drift/hallucination → đề xuất vá → chạy lại gate. Phụ thuộc 02, 05. Bước 1: script đọc `metrics.jsonl` + `drift-report.json` → sinh backlog vá tự động. |
|
||||
@@ -46,7 +47,7 @@
|
||||
| **06 Onboard dự án 2** | 📋 chưa bắt đầu | Chứng minh reuse: cắm 1 repo khác + golden/corpus/input, đăng ký qua `verify-harness-reuse.sh`, không sửa gate. Phụ thuộc 01. |
|
||||
| **08 Context compression** | 📋 chưa bắt đầu | Nén prompt/context giảm token (H1.5). Làm SAU khi core ổn (nén thêm bề mặt rủi ro → cần quét lại). Bước 1: đo token baseline mỗi step, thử tóm tắt có kiểm chứng (H3 judge so sánh). |
|
||||
| **12 Domain Pack SDK** | 📋 chưa bắt đầu | Onboard bằng khai báo (golden/corpus/policy theo domain). Phụ thuộc 01, 06. |
|
||||
| **01 Restructure** | 📋 chưa bắt đầu | Tái cấu trúc thư mục Phase 0→6. Nền cho 06/12. Rủi ro cao (đụng nhiều path) → làm trên nhánh riêng, giữ 211/0. |
|
||||
| **01 Restructure** | 📋 chưa bắt đầu | Tái cấu trúc thư mục Phase 0→6. Nền cho 06/12. Rủi ro cao (đụng nhiều path) → làm trên nhánh riêng, giữ 218/0. |
|
||||
| Future B1–B6 | 💤 vision | `CASAN_PLAN_FUTURE_PHASES.md` — approval workflow nâng cao · state machine · model benchmark · governed memory · auto-remediation · platform KPI. |
|
||||
|
||||
---
|
||||
@@ -62,7 +63,8 @@ Harness thấp nhất = 80 (H5, H6). Để CẢ pipeline vào Strong cần đón
|
||||
cd AINative_OKR_CASAN5
|
||||
for s in run-casan4-harness adversarial-harness phase1-track-a phase2-track-c \
|
||||
phase3-evidence-pack phase-h5-approval phase-h5-infra phase-h6-agentops \
|
||||
phase-c7-incident phase-h4-multilingual phase-c6-sandbox phase-h4-split-inject; do
|
||||
phase-c7-incident phase-h4-multilingual phase-c6-sandbox phase-h4-split-inject \
|
||||
phase10-traceability; do
|
||||
bash .specify/tests/$s-tests.sh >/dev/null 2>&1 && echo "$s OK" || echo "$s FAIL"; done
|
||||
# KMS live cần Vault dev; C6 live cần Docker (không có thì SKIP, không fail).
|
||||
```
|
||||
|
||||
@@ -34,13 +34,13 @@
|
||||
|
||||
| Control | Where | Test |
|
||||
|---|---|---|
|
||||
| `casan pack` / `casan verify-pack` (mapped to `evidence-pack.sh`): standard 12-file pack, hash manifest, signed head, tamper-evident verify, certified-run gate (no false certification, no silent skip) | `evidence-pack.sh`, `evidence-pack-build.py`, `evidence-pack-verify.py` | phase3-evidence-pack |
|
||||
| `casan pack` / `casan verify-pack` (mapped to `evidence-pack.sh`): standard 13-file pack incl. `traceability-matrix.json`, hash manifest, signed head, tamper-evident verify, certified-run gate (no false certification, no silent skip) | `evidence-pack.sh`, `evidence-pack-build.py`, `evidence-pack-verify.py` | phase3-evidence-pack |
|
||||
|
||||
### Phase 4 — H5 governance hardening (raises the lowest harness) — mixed
|
||||
|
||||
| ID | Control | Status | Where | Test |
|
||||
|---|---|---|---|---|
|
||||
| C4 | **Approval-identity**: high-risk approval trusted only when a REGISTERED reviewer cryptographically SIGNS the request and their role is authorized — env-var approver no longer enough (SoD still enforced) | [implemented+tested] | `approval-sign.sh`, `approval-verify.sh`, `reviewers.registry`, `governance-check.sh` (`CASAN_APPROVAL_STRICT=1`) | phase-h5-approval (8) |
|
||||
| C4 | **Approval-identity + mock IdP/OIDC**: high-risk approval trusted only when a REGISTERED reviewer signs the request OR an IdP-signed RS256 JWT proves approver identity/role/expiry and binds to this request — env-var approver no longer enough (SoD still enforced) | [implemented+tested] (mock IdP; production JWKS still planned) | `approval-sign.sh`, `approval-jwt-mint.py`, `approval-verify.sh`, `reviewers.registry`, `governance-check.sh` (`CASAN_APPROVAL_STRICT=1`) | phase-h5-approval (12) |
|
||||
| B3 | **KMS key management**: sign audit/telemetry head via Vault Transit (key `exportable:false` → never leaves KMS) + key **rotation**; validated live | [implemented+tested] (live when Vault present; skip-aware otherwise) | `vault-kms.sh` (`rotate`, `assert-nonexportable`), `sign-audit-head.sh` | phase-h5-infra (KMS) |
|
||||
| C5 | **External WORM audit**: ship audit head to a hash-linked append-only ledger; detect local rollback (`AUDIT_GAP_DETECTED`) and ledger tamper (`AUDIT_LEDGER_TAMPERED`) | [implemented+tested] (local ledger MVP) | `worm-ledger.py`, `audit-ship.sh`, `verify-audit-gap.sh` | phase-h5-infra (WORM) |
|
||||
|
||||
@@ -61,6 +61,8 @@
|
||||
| B1 | **Multilingual VI/JA injection (V2)**: VI/JA block-patterns (matched on raw UTF-8, anchored on the injection object) catch injections English regex missed, with 0 false positives on the benign VI/JA corpus | [implemented+tested] | `prompt-filter.yaml` (PI-VI-*, PI-JA-*) | phase-h4-multilingual (7) |
|
||||
| C6 | **TRUE runtime isolation (V22)**: container sandbox (`--network=none --read-only --pids-limit --cap-drop=ALL`, workspace-only mount) — the kernel neutralises host-file reads / egress / out-of-workspace writes; upgrades the static scaffold | [implemented+tested] (live via Docker; skip-aware) | `sandbox-container.sh`, `sandbox-run.sh` (`CASAN_SANDBOX_MODE=container`) | phase-c6-sandbox (6) |
|
||||
| B2 | **Split + classifier injection (V5,V6)**: `context-assemble-scan.sh` scans the concatenated context so a payload split across benign pieces is caught on assembly; verdict-steering patterns (PI-CLS-*) block content that tries to hijack the evaluator | [implemented+tested] | `context-assemble-scan.sh`, `prompt-filter.yaml` (PI-CLS-*) | phase-h4-split-inject (8) |
|
||||
| B4 | **Model-digest pinning (V16)**: approved Ollama model digest pinned; router verifies live digest before generation/classify/judge; mismatch blocks by default with warn mode for rollout | [implemented+tested] | `model-digest-check.sh`, `model-call.py`, `.specify/security/model-digest.pin` | phase3-model-router (digest cases) |
|
||||
| Plan-10 | **Traceability REQ→code→test MVP**: parse `FR-*`, require code+test mapping per FR, generate matrix, and include it in Evidence Pack | [implemented+tested] | `traceability-matrix.py`, `.specify/traceability-map.json`, `docs/output/casan/traceability-matrix.json` | phase10-traceability (3) |
|
||||
|
||||
## 2. Test inventory (all suites)
|
||||
|
||||
@@ -71,14 +73,15 @@
|
||||
| `phase1-track-a-tests.sh` | 25 | Track A hardening |
|
||||
| `phase2-track-c-tests.sh` | 29 | Track C-MVP |
|
||||
| `phase3-evidence-pack-tests.sh` | 7 | Evidence Pack MVP |
|
||||
| `phase-h5-approval-tests.sh` | 8 | Approval-identity (C4) |
|
||||
| `phase-h5-approval-tests.sh` | 12 | Approval-identity (C4) + mock IdP/OIDC JWT |
|
||||
| `phase-h5-infra-tests.sh` | 7 | KMS (B3, live/skip-aware) + WORM (C5) |
|
||||
| `phase-h6-agentops-tests.sh` | 20 | live alerting (D1) + provider-API/reconcile (D2) + hosted dashboard (D3) + window breaker (D4); against live local HTTP endpoints |
|
||||
| `phase-c7-incident-tests.sh` | 15 | **New** — incident severity + scoped kill-switch (C7) + wrapper enforcement |
|
||||
| `phase-h4-multilingual-tests.sh` | 7 | **New** — VI/JA injection block + benign VI/JA 0-FP (B1) |
|
||||
| `phase-c6-sandbox-tests.sh` | 6 | **New** — TRUE container isolation (C6, live via Docker / skip-aware) |
|
||||
| `phase-h4-split-inject-tests.sh` | 8 | **New** — split-injection assembly scan + classifier-inject (B2) |
|
||||
| **Total** | **211** | Baseline 79 preserved; +132 new hardening checks. Last full run 2026-07-05 @ head of `feat/plan07-track-a-hardening`, 0 fail (KMS + container isolation validated live via Vault dev + Docker). |
|
||||
| `phase10-traceability-tests.sh` | 3 | **New** — Plan-10 FR→code→test matrix + fail-able missing-test gate |
|
||||
| **Total** | **218** | Baseline 79 preserved; +139 new hardening/traceability checks. Last full harness run 2026-07-06, 0 fail. Direct `phase3-model-router-tests.sh` adds 10/0 for model-router/digest/cloud/fallback checks. |
|
||||
|
||||
Run order note: `run-casan4-harness-tests.sh` does `rm -rf .specify/logs`, so run it
|
||||
**first** and never concurrently with the other suites.
|
||||
@@ -87,13 +90,13 @@ Run order note: `run-casan4-harness-tests.sh` does `rm -rf .specify/logs`, so ru
|
||||
|
||||
| Area | Status | Plan ref |
|
||||
|---|---|---|
|
||||
| Model-digest pinning | [planned] — sliding-window circuit breaker (V15) is now done (Phase 5 D4) | Plan-07 B4 (V16) |
|
||||
| Production model provenance beyond local Ollama digest | [partial] — local model digest pinning is implemented+tested; cloud provider model attestations/SBOM-style provenance still planned | Plan-07 B4 (V16) |
|
||||
| Live alerting to a managed channel (Slack/PagerDuty + on-call rota) | [partial] — webhook dispatch + dedup + dead-letter done; managed channel & escalation are config away, incident workflow is C7 | Plan-07 C7 / Phase 5 D1 |
|
||||
| Hosted telemetry dashboard | [partial] — HTTP-served dashboard + stale-aware `/healthz` done locally; deployed host (nginx/container, auth) planned | Phase 5 D3 |
|
||||
| Provider billing-API telemetry | [partial] — API fetch + schema gate + local-vs-provider reconciliation done against a live local endpoint; real OpenAI/Anthropic usage-API calls (needs keys) planned | Phase 5 D2 |
|
||||
| True runtime isolation | [partial] — real container isolation done + validated live via Docker (C6 phase-6); nsjail/rootless + a hardened base image for CI still planned | Plan-07 C6 (V22) |
|
||||
| KMS key management (rotation, non-exportable) | [partial] — Vault Transit path implemented + validated live; not yet the default (local-key fallback), no HSM/short-lived IdP tokens | Plan-07 B3 |
|
||||
| Reviewer approval workflow | [partial] — cryptographic **approval-identity** done (signed reviewer + role); live **IdP (OIDC/JWT)** + policy versioning/diff still planned | Plan-07 C4 (V20) |
|
||||
| Reviewer approval workflow | [partial] — cryptographic **approval-identity** done (signed reviewer + role) + mock **IdP/OIDC JWT** done; live IdP/JWKS + policy versioning/diff still planned | Plan-07 C4 (V20) |
|
||||
| External append-only (WORM) audit | [partial] — hash-linked local ledger + rollback/tamper detection done; true WORM store (S3 Object Lock/QLDB) + trusted timestamp planned | Plan-07 C5 (V21) |
|
||||
| Live CVE/OSV scanning wired in | [partial] — availability detected; local denylist authoritative offline | Plan-07 C2 follow-up |
|
||||
|
||||
@@ -103,14 +106,14 @@ Track A + Track C-MVP + Evidence Pack + H5/H6 hardening + the deep-gap closers
|
||||
(C7 incident/kill-switch, VI/JA multilingual, true container isolation, split &
|
||||
classifier injection) raise H4/H5/H6 from "PoC/demo (~3.0/5)" to **early
|
||||
internal-production hardening**, with executable adversarial tests for every
|
||||
control (**211 checks, 0 fail** — last full run 2026-07-05; KMS + container
|
||||
isolation validated live via Vault dev + Docker). Fair maturity score
|
||||
control (**218 core checks, 0 fail** — last full harness run 2026-07-06; KMS + container
|
||||
isolation validated live/skip-aware via Vault dev + Docker). Fair maturity score
|
||||
(`00_SUBMISSION_PACKAGE/evidence/scoring-run-report.md`): **H4 80→83** (multilingual
|
||||
+ split/classifier closed), **H2 80→82** (real sandbox isolation), C7 incident
|
||||
dimension closed; **H5 and H6 remain at 80** (their remaining gaps are infra), so the
|
||||
**lowest harness stays 80** — CASAN **Level 4**, proven by attack. This is **not** full
|
||||
production readiness: crossing the whole pipeline into "Strong (81+)" still needs the
|
||||
H5/H6 infra items — live IdP (OIDC/JWT), a true WORM store (S3 Object Lock),
|
||||
H5/H6 infra items — live IdP/JWKS, a true WORM store (S3 Object Lock),
|
||||
KMS-by-default + HSM, a deployed dashboard host + managed alert channel/on-call, and
|
||||
real billing-API telemetry — the [partial]/[planned] rows above and in
|
||||
`CASAN_PLAN_07_PRODUCTION_HARDENING.md`.
|
||||
|
||||
@@ -64,9 +64,9 @@ flowchart LR
|
||||
| 04 Self-improve | ⬜ chưa bắt đầu | |
|
||||
| 05 CI/CD | ⬜ chưa bắt đầu | |
|
||||
| 06 Onboard | ⬜ chưa bắt đầu | |
|
||||
| 07 Production hardening | 🟡 Track A ✅ + C-MVP ✅ + H5-hardening ✅ (B + C7 + IdP/WORM-store chưa) | **Track A 25/25 · Track C-MVP 29/29 · H5-hardening (C4 approval-identity 8/8 + KMS live + C5 WORM = infra 7/7)**. Điểm công tâm: H5 76→80, harness thấp nhất giờ H6=79, ~80.7/100. Track B + C7 + live IdP/WORM-store còn [planned]. Chi tiết: `CASAN_HARDENING_STATUS.md` · `evidence/scoring-run-report.md` |
|
||||
| 07 Production hardening | 🟡 Track A ✅ + C-MVP ✅ + H5/H6/deep-gap hardening ✅ (infra prod còn) | **Track A 25/25 · Track C-MVP 29/29 · H5 approval/OIDC 12/12 · infra 7/7 · H6 20/20 · C7/H4/C6 deep-gap closers done**. Điểm công tâm ~81/100; harness thấp nhất 80. Live IdP/JWKS, WORM-store thật, KMS default/HSM, managed alert/billing/dashboard deploy còn [planned]. Chi tiết: `CASAN_HARDENING_STATUS.md` · `evidence/scoring-run-report.md` |
|
||||
| 08 Context compression | ⬜ chưa bắt đầu | H1.5 cross-cutting; làm **sau** Plan-07 core (nén thêm bề mặt rủi ro) |
|
||||
| 09 Evidence Pack | 🟢 MVP ✅ (7/7) | `casan pack`/`verify-pack` → `evidence-pack.sh`: 12-file pack + manifest + signed head + certified-run gate, tamper-evident. Chi tiết: `CASAN_HARDENING_STATUS.md` |
|
||||
| 10 Traceability + H3 Eval | ⬜ chưa bắt đầu | Khác biệt nhất; lõi = ma trận REQ→code→test |
|
||||
| 09 Evidence Pack | 🟢 MVP ✅ (7/7) | `casan pack`/`verify-pack` → `evidence-pack.sh`: 13-file pack incl. traceability matrix + manifest + signed head + certified-run gate, tamper-evident. Chi tiết: `CASAN_HARDENING_STATUS.md` |
|
||||
| 10 Traceability + H3 Eval | 🟢 MVP ✅ (3/3) | `traceability-matrix.py --gate`: parse FR-xx → code/test matrix; Evidence Pack chứa `traceability-matrix.json`. Còn line/symbol-level + eval-set độc lập. |
|
||||
| 12 Domain Pack SDK | ⬜ chưa bắt đầu | Cần Plan-01 xong trước |
|
||||
| Future phases (B1–B6) | 💤 vision | Backlog `CASAN_PLAN_FUTURE_PHASES.md` — chưa xây |
|
||||
|
||||
@@ -19,7 +19,7 @@
|
||||
|
||||
## 2. Thang điểm sẵn sàng production (0–5, cao = tốt)
|
||||
|
||||
> ✅ **CẬP NHẬT 2026-07-05 — Track A + C-MVP + Evidence Pack + H5-hardening + H6-hardening ĐÃ LÀM + TEST (175 checks, 0 fail; KMS đã validate LIVE qua Vault 2026-07-04).**
|
||||
> ✅ **CẬP NHẬT 2026-07-06 — Track A + C-MVP + Evidence Pack + H5/H6/deep-gap hardening + Plan-10 traceability ĐÃ LÀM + TEST (218 core checks, 0 fail; model-router riêng 10/0).**
|
||||
> Bảng dưới có cột **Baseline → Nay**. Điểm chấm CÔNG TÂM (0–100, theo `casan_harness_assessment.md`):
|
||||
> **H4 = 80 · H5 = 76→80 ⬆ · H6 = 79→80 ⬆ · trung bình 7 harness ~80.9/100 · không còn harness nào dưới 80 → CASAN Level 4 (vững ngưỡng)**.
|
||||
> Nguồn: `00_SUBMISSION_PACKAGE/evidence/scoring-run-report.md`. Chi tiết implemented-vs-planned: `CASAN_HARDENING_STATUS.md`.
|
||||
@@ -33,7 +33,7 @@
|
||||
| Quan sát (observability) | 3 | **4** | telemetry toàn vẹn (ký) · **alerting LIVE** (webhook + dedup + dead-letter, end-to-end từ step fail) · **dashboard hosted** (`/healthz` stale-aware) · **provider-API reconcile** (bắt under-reporting) · window breaker (V15) | dashboard deploy thật + auth · kênh alert managed (Slack/PagerDuty + on-call) · billing-API thật |
|
||||
| Đa domain/i18n | 2 | **2.5** | benign corpus VI/JA/EN đo được (FP 0%) | detection vẫn chủ yếu EN (Track B) |
|
||||
| Quản lý khóa | 2 | **4** | **KMS live** (Vault Transit) — ký qua KMS, **rotate**, khoá **non-exportable** (đã chạy thật) | KMS chưa mặc định (fallback local) · HSM · IdP token ngắn hạn |
|
||||
| Phủ kiểm thử | 4 | **4.5** | **175 test** (35+44+25+29+7+8+7+20) đối kháng, 0 fail | thêm ca đa ngôn ngữ khi làm Track B |
|
||||
| Phủ kiểm thử | 4 | **4.5** | **218 core tests** (35+44+25+29+7+12+7+20+15+7+6+8+3) đối kháng, 0 fail | line/symbol-level traceability + CI release gate |
|
||||
|
||||
**Điểm trung bình (H4/H5/H6 mở rộng) ~3.0 → ~4.0/5; chấm công tâm per-harness H4=80 · H5=80 · H6=80 (~4.0/5). Không còn harness nào dưới 80.**
|
||||
|
||||
@@ -45,12 +45,12 @@
|
||||
| Supply-chain (dependency sinh ra) | 1 | **3.5** | `supply-chain-gate.sh` (V18) — typosquat/postinstall/denylist + dep-diff | CVE/OSV scanner live chưa nối |
|
||||
| Data-governance / anti-exfil | 2 | **4** | `data-exfil-guard.sh` (V19) — secret→cloud BLOCK, PII→audit mask | — |
|
||||
| Evidence Pack (Plan-09) | — | **4** | `casan pack/verify-pack` — manifest ký, tamper-evident, certified-gate | KMS sign · hosted store |
|
||||
| Policy governance / approval | 2 | **4** | **approval-identity** (C4) — reviewer KÝ request + role authz, hết env-var; SoD giữ (đã test 8/8) | live IdP (OIDC/JWT) · policy versioning/diff [planned] |
|
||||
| Policy governance / approval | 2 | **4** | **approval-identity + mock IdP/OIDC** (C4) — reviewer KÝ request hoặc JWT RS256 + role authz, hết env-var; SoD giữ (đã test 12/12) | live IdP/JWKS · policy versioning/diff [planned] |
|
||||
| Runtime sandbox | 1 | **2.5** | `sandbox-run.sh` scaffold (V22) — chặn ssh/egress/forkbomb/write-outside + ulimit | **cô lập kernel thật** (container/nsjail) [planned] |
|
||||
| External append-only audit | 1 | **3.5** | **WORM ledger** (C5) — ship head hash-link ngoài, bắt rollback (`AUDIT_GAP_DETECTED`) + tamper | WORM store thật (S3 Object Lock) · trusted timestamp [planned] |
|
||||
| Incident response | 1 | **1** | — | severity/owner/kill-switch (C7) [planned] |
|
||||
|
||||
**→ C-MVP (C1+C2+C3 + Evidence Pack) ~3.8/5 + H5-hardening (C4 approval-identity, KMS live, C5 WORM) + H6-hardening (D1 alerting live, D2 provider-API reconcile, D3 dashboard hosted, D4 window breaker V15) [đã làm + test thật]. Còn: sandbox isolation thật, C7 incident, live IdP/WORM-store, dashboard deploy + kênh alert managed, billing-API thật [planned].** Trio H4/H5/H6 nay **~4.0/5 (H4=80·H5=80·H6=80)**; harness thấp nhất nhích **76 (H5) → 79 (H6) → 80 (đồng đều)**. Production toàn diện vẫn cần các mục [planned] ở trên.
|
||||
**→ C-MVP (C1+C2+C3 + Evidence Pack) ~3.8/5 + H5-hardening (C4 approval-identity + mock OIDC, KMS live, C5 WORM) + H6-hardening (D1 alerting live, D2 provider-API reconcile, D3 dashboard hosted, D4 window breaker V15) + deep-gap closers + Plan-10 traceability [đã làm + test thật]. Còn: live IdP/JWKS, WORM-store thật, KMS default/HSM, dashboard deploy + kênh alert managed, billing-API thật [planned].** Trio H4/H5/H6 nay **~4.0/5 (H4=80·H5=80·H6=80)**; harness thấp nhất nhích **76 (H5) → 79 (H6) → 80 (đồng đều)**. Production toàn diện vẫn cần các mục [planned] ở trên.
|
||||
|
||||
## 3. Bảng đường lọt (tóm tắt từ threat-model)
|
||||
|
||||
|
||||
@@ -0,0 +1,36 @@
|
||||
# CASAN PLAN 10 — Traceability REQ→Code→Test + H3 Eval
|
||||
|
||||
> Status 2026-07-06: **MVP implemented + tested**. Scope is deterministic
|
||||
> traceability for the OKR sample app; broader H3 eval-set expansion remains a
|
||||
> platform follow-up.
|
||||
|
||||
## MVP delivered
|
||||
|
||||
| Capability | Where | Verification |
|
||||
|---|---|---|
|
||||
| Parse `FR-*` from `docs/input/okr-requirement.md` | `.specify/scripts/bash/traceability-matrix.py` | `phase10-traceability-tests.sh` |
|
||||
| Declarative FR→code→test map | `.specify/traceability-map.json` | gate checks every mapped file exists |
|
||||
| Gate every FR has >=1 code file and >=1 test file | `traceability-matrix.py --gate` | missing FR test coverage fails |
|
||||
| Evidence Pack includes traceability matrix | `evidence-pack-build.py` | `phase3-evidence-pack-tests.sh` expects 13 files |
|
||||
|
||||
## Current result
|
||||
|
||||
```bash
|
||||
python3 .specify/scripts/bash/traceability-matrix.py --gate
|
||||
# TRACEABILITY_MATRIX requirements=5 pass=5 fail=0 ...
|
||||
|
||||
bash .specify/tests/phase10-traceability-tests.sh
|
||||
# TRACEABILITY SUMMARY: PASS=3 FAIL=0
|
||||
```
|
||||
|
||||
Generated artifact:
|
||||
|
||||
- `AINative_OKR_CASAN5/docs/output/casan/traceability-matrix.json`
|
||||
|
||||
## Next steps
|
||||
|
||||
| Priority | Work | Done when |
|
||||
|---|---|---|
|
||||
| P1 | Add traceability line references or symbol references, not only file paths | matrix can point to precise code/test locations |
|
||||
| P2 | Expand H3 eval-set beyond the OKR sample app | independent eval corpus runs through judge gates |
|
||||
| P3 | Enforce traceability in release CI | CI blocks a new FR without code+test coverage |
|
||||
@@ -18,7 +18,7 @@ Không — vì chúng tôi **cố ý không mở hết**. Chỉ 3 plan mới đ
|
||||
Vì (a) rủi ro vỡ bản demo đang chạy; (b) một số plan **phụ thuộc nhau** (ví dụ Domain Pack cần restructure xong; Model benchmark cần nối model thật xong mới có số). Làm sai thứ tự = tốn công mà không có bằng chứng.
|
||||
|
||||
**A4. Roadmap này có làm mất tính trung thực khi trình bày không?**
|
||||
Không, nếu nói đúng nhãn: *"phần đã làm & đo là H4/H5/H6 + hardening; Evidence Pack/Traceability/Domain Pack là bước kế tiếp đã có kế hoạch; state machine/governed memory là tầm nhìn dài hạn chưa xây."* Ranh giới rõ ràng = điểm cộng độ chín.
|
||||
Không, nếu nói đúng nhãn: *"phần đã làm & đo là H4/H5/H6 + hardening; Evidence Pack và Traceability MVP đã có test; Domain Pack là bước kế tiếp đã có kế hoạch; state machine/governed memory là tầm nhìn dài hạn chưa xây."* Ranh giới rõ ràng = điểm cộng độ chín.
|
||||
|
||||
**A5. Nếu chỉ được chọn 3 plan, chọn gì và vì sao?**
|
||||
**09 Evidence Pack · 10 Traceability · 12 Domain Pack.** Ba cái này nâng CASAN từ "harness bảo vệ AI" lên "**nền tảng AI-SDLC có bằng chứng, đo chất lượng, tái dùng đa domain**" — đúng 3 trục thi.
|
||||
|
||||
Reference in New Issue
Block a user