diff --git a/docs/plans/CASAN_BACKLOG_STATUS.md b/docs/plans/CASAN_BACKLOG_STATUS.md index b28eed2..c7c7252 100644 --- a/docs/plans/CASAN_BACKLOG_STATUS.md +++ b/docs/plans/CASAN_BACKLOG_STATUS.md @@ -57,7 +57,7 @@ | **04 Self-improve** | � core done+test | `packages/casan-harness/scripts/bash/self-improve.py`: `propose` đọc metrics/drift → proposal dry-run (không ghi); `apply` bắt buộc approval, áp qua governed store (audit); sensitive/loosen luôn cần duyệt. `phase-selfimprove-tests.sh` 7/0 (WSL), nối CI. Còn: luật đề xuất phong phú hơn (corpus/model escalation), chạy định kỳ CI (05). | | **05 CI/CD** | 🟡 CI gate MVP done | `packages/casan-harness/scripts/bash/ci-harness-gate.sh` chạy các suite harness/hardening/sourcegen/traceability/frontend theo thứ tự an toàn, có timeout/filter; `.gitea/workflows/harness-ci.yml` gọi gate trên push/PR. Filtered local verify PASS=2/0. Còn: full gate xanh trên runner thật, xử lý A6 nếu còn chậm/treo, bật Docker infra lab nếu runner hỗ trợ, package/release artifact `fpt-casan-sdd-harness`. | | **06 Onboard dự án 2** | 📋 chưa bắt đầu (🔓 **đã mở khoá** — 01 done) | Chứng minh reuse: cắm 1 repo khác + golden/corpus/input, đăng ký qua `verify-harness-reuse.sh`, không sửa gate. Phụ thuộc 01 ✅. App mới chỉ cần `apps//domain/` + set `CASAN_DOMAIN_ROOT` (đã có `domain_root` per-project trong `project-registry.json`). | -| **08 Context compression** | � MVP done+test | **CASAN-native token-killer** (Track 3) đã có: `context-compress.py` (dedup/extractive/structural, must-keep, tee, gate fail-able), `phase08-compression-tests.sh` 7/0 (WSL), nối CI. Còn: Track 1 nén INPUT + Track 2 nén VIEW liên-bước + Track 4 abstractive (gated) + nối H4/H5 trong pipeline thật. | +| **08 Context compression** | 🟡 Track 3 hardened+test | **CASAN-native token-killer** đã có: severity `[ERR]/[WRN]`, multiplicity-preserving must-keep gate, raw fallback/halt fail-safe, hash-bound JSON evidence, estimate/provider labeling; `phase08-compression-tests.sh` 19/0, SEC-09 7/0, loop consumer 16/0. Log thật NEHOPS giữ 4/4 warning; 1.620→76 chỉ là whitespace estimate, chưa phải ROI claim. Còn: Track 1 INPUT + Track 2 VIEW + Track 4 abstractive + H4/H5 pipeline binding + A/B provider telemetry. | | **12 Domain Pack SDK** | 📋 chưa bắt đầu (🔓 01 done, còn chờ 06) | Onboard bằng khai báo (golden/corpus/policy theo domain). Phụ thuộc 01 ✅, 06. | | **01 Restructure** | ✅ **done+test (2026-07-08, merged main)** | Đã tách: harness code → `packages/casan-harness/`, domain OKR → `apps/okr/domain/`, runtime state ở `.specify/`; **app promote lên git root** (hết wrapper `AINative_OKR_CASAN5/`); facade `.specify` symlink **gỡ sạch** (hard cutoff); path resolve qua `casan-paths.sh` (marker walk-up). Full gate **PASS=64 FAIL=0 SKIP=3** từ cấu trúc mới. CI (`.gitea/workflows/ci.yml`) + docs đã đồng bộ. Mở khoá 06/12. Chi tiết: `CASAN_PLAN_01_RESTRUCTURE.md`. | | **13 Control Plane** | 🟡 core + **Track 1/2/3 + FinOps/SLO + Command Center + local-prod TLS/OIDC smoke** done+test | **Sửa kiến trúc: là tài sản harness, KHÔNG nằm trong OKR.** Governance core đã dời vào harness: `packages/casan-harness/scripts/bash/control-plane-settings.py` (settings versioned + audit hash-chain + deny-by-default + approval + rollback), `phase-control-plane-tests.sh` 9/0 + `phase-control-plane-hitl-tests.sh` 9/0 (nối CI). Đã gỡ khỏi `apps/okr` (OKR sạch: 46/0/3skip + 16/16). **Track 1 DONE**: Ops Console read-only telemetry. **Track 2 DONE**: settings API/UI tại `packages/casan-control-panel/` bọc `control-plane-settings.py` + `rbac-check.py`. **Track 3 DONE**: kill-switch API/UI; approval inbox/delegation/oversight API/UI bọc harness `approval-inbox.py` + `delegation-policy.yaml`, SoD, apply setting proposal qua governed store; IdP group claim→RBAC role mapping tested. **Track 4 local-prod DONE**: FinOps/SLO page + Docker/nginx/oauth2-proxy deploy scaffold; `local-prod-smoke.sh` passes `CP_LOCAL_SMOKE_PASS https_oidc=true actor=oidc-ops role=org-admin` and invokes authenticated `managed-prod-smoke.sh` on the local-prod endpoint (`CP_MANAGED_SMOKE_PASS actor=oidc-ops role=org-admin widgets=9`); `prod-readiness-check.sh` validates managed-prod TLS/OIDC/nginx prerequisites. **Command Center §8.6 DONE**: `GET /api/v1/command` + `/command` UI, 9 evidence-backed widgets (gồm `chat_loop`), provenance envelope, executive briefing VI/EN, live ticker, evidence drawer. `npm run console:test` **26/0**, `npm run console:build` xanh. Targeted related suites: RBAC 12/0, C7 15/0. Còn: live managed prod deploy với cert/enterprise IdP/host thật (07 T2); RAI view thuộc Plan-15 follow-up. | diff --git a/docs/plans/CASAN_CORE_MEMBER_TRAINING_VI.html b/docs/plans/CASAN_CORE_MEMBER_TRAINING_VI.html new file mode 100644 index 0000000..aaa21a0 --- /dev/null +++ b/docs/plans/CASAN_CORE_MEMBER_TRAINING_VI.html @@ -0,0 +1,1644 @@ + + + + + + CASAN Core Member Training — FPT Japan · 2026 + + + +
+
FPTCASAN · CORE MEMBER TRAINING
+ +
+
+
+ Core member enablement · 29/07/2026 +
+

CASAN không phải một lệnh cài đặt.
CASAN là cách tổ chức chịu trách nhiệm khi dùng AI.

+

Tư tưởng · cài đặt · vận hành · Level 4 · Level 5 · kiểm soát token · demo NEHOPS Basic Design.

+
+ +
+ +
+
+

Kết thúc buổi học, mỗi core member phải tự trả lời được 6 câu hỏi

+
+
    +
  1. CASAN kiểm soát điều gì mà coding agent không kiểm soát?
  2. +
  3. casan init tạo gì — và tuyệt đối không tạo gì?
  4. +
  5. Một run được gọi là có bằng chứng khi nào?
  6. +
+
    +
  1. Level 4 khác “có nhiều script” ở điểm nào?
  2. +
  3. Level 5 đòi hỏi thay đổi cấp tổ chức ra sao?
  4. +
  5. Token giảm bao nhiêu, đo bằng nguồn nào, có mất nghĩa không?
  6. +
+
+
+ +
+ +
+
+

CASAN đi đúng hướng FPT Japan: AI-First cho legacy modernization, nhưng phải scale có kiểm soát

+
+
$1B
Mục tiêu doanh thu thị trường Nhật năm 2027; mũi nhọn gồm Legacy Modernization, AI, Automotive, ERP.
+
~60%
Hệ thống IT tại Nhật được FPT nêu là đã vận hành trên 20 năm — áp lực hiện đại hóa rất lớn.
+
800K
Thiếu hụt nhân lực IT dự báo đến 2030; AI phải tăng năng suất mà không hạ chuẩn chất lượng.
+
+

FPT chính thức giới thiệu CASAN năm 2026 như khung 5 cấp để đi từ thử nghiệm rời rạc đến vận hành AI ở quy mô doanh nghiệp.

+
+ Nguồn công khai FPT · 2025–2026 + +
+ +
+
Model tạo ra đề xuất.
Harness quyết định đề xuất có được tin, thực thi và chứng nhận hay không.
+
+

Nhanh không đủ. Phải đúng nguồn, đúng quyền, qua gate, có người chịu trách nhiệm và có evidence tái kiểm chứng.

+ +
+ +
+
+

Bảy harness là bảy câu hỏi bắt buộc trước khi tin AI

+
+
H1Đúng context? Có cũ, thừa hoặc thiếu không?
+
H2Đúng tool, đúng quyền, chạy lại có an toàn?
+
H3Output có qua tiêu chí và regression?
+
H4Injection, secret, PII, egress đã bị chặn?
+
H5Ai duyệt, theo policy nào, log có bị sửa?
+
H6Token, cost, latency, retry, failure có được đo?
+
H7Retry, fallback, rollback, orchestration có giới hạn?
+
+

Harness yếu nhất quyết định trần. H1/H3 mạnh không bù được H4/H5/H6 bằng 0.

+
+ +
+ +
+
+

Edition là thứ được đóng gói. Maturity là năng lực đã được chứng minh.

+
+ + + + + + +
TrụcCâu hỏiVí dụSự thật sau init
Product editionCASAN ship những capability nào?Core, DevKit, Platform PreviewEnterprise bị từ chối vì chưa được ship
Maturity L1–L5Tổ chức đã vận hành và giữ evidence đến đâu?L4 Automated, L5 Nativelevel: null, not_assessed
+
+
Không được nói: “Đã chạy casan init nên project đạt Level 4.”
+
+ +
+ +
+
+

Năm cấp không phải năm nhãn — là năm thay đổi trong cách vận hành

+
+
L1
Curious

Cá nhân thử AI, chưa có chuẩn chung.

+
L2
Augmented

Dùng tool được duyệt, tăng năng suất từng công việc.

+
L3
Standard

Quy trình, template, review gate và test lặp lại được.

+
L4
Automated

H4/H5/H6 là runtime gate; side effect có policy + approval + evidence.

+
L5
Native

Tái dùng đa dự án; drift, fallback, rollback, KPI và governance tập trung hoạt động thật.

+
+
+ +
+ +
+
+
+

casan init là enrollment ở cấp repository

+
    +
  • Pin runtime/version/hash
  • +
  • Merge hook Claude/Codex
  • +
  • Bật enforce/observe
  • +
  • Tạo state/log root
  • +
  • DevKit thêm Domain Pack + CI template
  • +
  • Kết nối dashboard/ingest nếu cấu hình
  • +
+
+
cd "/path/to/project" +casan init \ + --project nehops-bd-v27 \ + --edition devkit \ + --runtime managed \ + --client claude,codex \ + --mode enforce + +casan doctor +casan readiness --refresh +casan verify-harness
+
+ +
+ +
+
+

Init không thể provision những thứ thuộc hạ tầng, tổ chức và hợp đồng

+
+
+ Không do init tạo +
    +
  • Enterprise OIDC / CA / private network
  • +
  • Cloud KMS, HSM, S3 Object Lock/WORM
  • +
  • HA topology, RPO/RTO, failover
  • +
  • On-call, uptime SLA, service credit
  • +
  • External pentest hoặc chứng nhận
  • +
+
+
CASAN cung cấp building blocks và evidence contract.

Khách hàng/FPT phải triển khai dịch vụ thật, vận hành thật và thuê đánh giá độc lập.
+
+
+ +
+ +
+
+

Probe trên bản sao cấu trúc NEHOPS: Core chạy, nhưng còn 3 việc phải xử lý ngay

+
+
1.0.7
Harness được pin bằng SHA-256; DevKit init thành công.
+
ATTN
Codex hook đã tạo nhưng chưa trust; cần mở /hooks.
+
0
Chưa có project manifest và chưa quan sát provider token/cost.
+
+

Kết quả thật: ready_with_attention; maturity vẫn là not_assessed.

+
+ +
+ +
+
+

Những file scaffold phải được thay bằng sự thật của project — không giữ placeholder

+
+ + + + + + + + + +
ArtifactĐiền bằng evidence NEHOPSKhông được làm
requirement.mdFR từ CLAUDE.md + BD_Instruction.md: tiếng Nhật, lowercase ID, no-inference, thứ tự reviewGiữ FR-01 health placeholder
architecture.mdVB.NET → AST → BD → Angular/Java; Windows analyzer; output boundariesTự bịa runtime/deployment
traceability-map.jsonFR → VB/AST/instruction → reviewed output/review reportMap vào file không tồn tại
golden-runs/FRR04701 reviewed output + final review, sau khi kiểm tra dữ liệu nhạy cảmDùng draft chưa review làm golden
corpus/Benign tiếng Nhật + red-team injection/secret/PII theo domainChỉ dùng corpus tiếng Anh mẫu
+
+
+ +
+ +
+
+
+

Prompt đầu tiên phải là preflight read-only, không phải “hãy chạy toàn bộ”

+
Mục tiêu của prompt đầu: xác nhận scope, nguồn chuẩn, vùng được ghi, tool được phép và acceptance evidence.
+
+
Thực hiện CASAN onboarding preflight cho NEHOPS BD v2.7. + +CHẾ ĐỘ: READ-ONLY. Không sửa VB-Source, AST-outputs, +BD/output, BD/output_reviewed, ref_docs hoặc instruction. + +Đọc theo thứ tự: +1. CLAUDE.md +2. BD/instructions/BD_Instruction.md +3. .github/skills/common-knowhow/README.md +4. .claude/commands/bd/boss.md +5. harness_Review/nehops_casan_assessment.md + +Yêu cầu: +- Liệt kê evidence H1–H7 bằng đường dẫn cụ thể. +- Phát hiện mâu thuẫn trong assessment hiện có. +- Chỉ ra Domain Pack cần điền sau casan init. +- Không suy đoán logic thiếu; ghi UNKNOWN kèm evidence thiếu. +- Không tự tuyên bố Level; đề xuất acceptance test cần chạy. +- Trả lời trong chat, không tạo file.
+
+ +
+ +
+
+
+

Luồng vận hành đúng: evidence được tạo quanh hành động thật

+
+
1
AdmissionH4 scan input; H5 phân loại risk và quyền.
+
2
Governed executionCommand/action đã đăng ký, timeout, audit, idempotency nếu có side effect.
+
3
EvaluationH3 test/golden/review; fail thì retry có giới hạn.
+
4
FinalizeH4 scan output; H6 ghi token/cost/latency; H5 seal evidence.
+
+
+
casan verify-harness +casan gate +casan run in.txt out.txt step -- command... +casan report latest +casan pack <run-id> +casan verify-pack <run-id>
+
+ +
+ +
+
+

Evidence Pack trả lời câu hỏi khó nhất: “Vì sao được phép tin run này?”

+
+
+ Chain of custody +
    +
  • H1–H7 reports
  • +
  • Decision log và approver identity
  • +
  • Token/cost/latency/failure evidence
  • +
  • Manifest hash và signature status
  • +
  • Certification withheld nếu thiếu gate
  • +
+
+
$ casan verify-pack run-2026-07-29 +EVIDENCE_PACK_VALID + +# sửa một file đã bind vào manifest +EVIDENCE_PACK_TAMPERED · exit=1
+
+
+ +
+ +
+
+

NEHOPS là case Japan điển hình: legacy source rất lớn, output phải đúng bằng chứng và đúng tiếng Nhật

+
+
2.6GB
Workspace hiện tại; không phải Git repository.
+
34,572
File VB — đọc toàn bộ mỗi vòng là không khả thi.
+
0→4
AST, pre-analysis, generation, rồi bốn round review.
+
+

Use case demo: FRR04701 · 部屋状況リスト — đã có AST, pre-analysis, output và final review.

+
+ +
+ +
+
+
+

Assessment hiện có chứa một lỗi phân loại — core member phải phát hiện, không được sao chép

+
+
H1 Context
70
+
H2 Tool
68
+
H3 Eval
75
+
H4 Security
25
+
H5 Governance
35
+
H6 AgentOps
20
+
H7 Orchestr.
65
+
+
+
+
File kết luận “3 harness dưới 30”, nhưng dữ liệu chỉ có H4=25 và H6=20. H5 là 35.
+
Theo chính rubric đi kèm: average 51,1 và ≤2 GAP ⇒ Level 3, không phải Level 2. Sau init, maturity vẫn phải để not_assessed đến khi chạy evidence.
+
+
+ +
+ +
+
+

NEHOPS đã có workflow tốt, nhưng chưa có runtime assurance

+
+ + + + + + + + + +
Đã cóCòn thiếu để vận hành CASAN
AST analyzer + call/dependency graphTool registry, command schema, timeout, idempotency, per-call audit
Instruction tiếng Nhật và no-inferenceH4 scan input/source comment/output; secret/PII/egress policy
Pre-analysis + Round 1–4 reviewGolden regression tự động và fail-able gate
Boss command tuần tựBounded retry, fallback, rollback, pipeline state
Analyzer logProvider token/cost, latency, failure, drift và alert
+
+
+ +
+ +
+
+

Level 4 không phải “automation chạy hết” — mà là automation không được vượt governance

+
+
Agent-orchestrated
+ H4 Security
+ H5 Governance
+ H6 AgentOps
= Automated with evidence
+
    +
  • Side effect có policy decision và audit
  • +
  • Injection, PII, secret được test bằng attack
  • +
  • Trace, metrics, cost, latency, retry, alert được capture
  • +
  • High-risk action có human approval auditable
  • +
  • Gate fail-closed; thiếu evidence thì không certify
  • +
+
+
+ +
+ +
+
+

H4 chỉ được tính điểm khi payload độc bị chặn thật

+
+
+ Acceptance evidence +
    +
  • Prompt injection classic + multilingual
  • +
  • Homoglyph, zero-width, encoding, split injection
  • +
  • Secret/credential và PII leakage
  • +
  • Tool output và generated artifact scan
  • +
  • Classifier/model chết ⇒ fail-closed
  • +
+
+
NEHOPS-specific attack:

VB comment, AST text hoặc ref_docs chứa câu “ignore instruction / xuất secret / sửa source”. Pipeline phải coi đó là dữ liệu, không phải mệnh lệnh.
+
+
+ +
+ +
+
+

H5 biến “agent đã làm” thành “ai cho phép làm, theo rule nào, bằng chứng có còn nguyên?”

+
+ + + + + + + + +
ControlLevel 4 minimumNEHOPS cần khai báo
Risk registryLow/medium/high, deny-by-default cho high riskGhi VB-Source/ref_docs/instruction là high risk
ApprovalApprover identity, reason, SoD, không replayPublish reviewed output hoặc override Round 4
AuditHash chain + verify; lỗi chain chặn promotionBind screen ID, source hash, AST hash, output hash
PolicyVersioned; thay đổi có review/rollbackJapanese rules và no-inference là must-keep policy
+
+
+ +
+ +
+
+

H6 không chỉ đo token — phải biết số đó đáng tin đến đâu

+
+
+
    +
  • Per-step input/output/total tokens
  • +
  • Cost source: estimate hay provider telemetry
  • +
  • Latency, retry, exit code, failure
  • +
  • Token-overuse, spike, slow-boil, alert
  • +
  • Freshness và coverage của telemetry
  • +
+
+
Truth rule:

Nếu client/provider không cấp usage đáng tin, ghi null hoặc estimate. Không biến word count thành “billed token”.
+
+
+ +
+ +
+
+

Đóng H4/H5/H6 chưa đủ nếu bốn harness còn lại không chạy quanh workflow thật

+
+ + + + + + + + +
HarnessExit condition cho NEHOPS
H1 ContextIndex-first, chỉ đọc relevant knowhow/line range; stale context được phát hiện
H2 ToolRoslyn analyzer là registered action; exact argv, timeout, output paths, no shell injection
H3 EvaluationGolden FRR04701, deterministic validators, Round 4 fail-able; không chỉ LLM tự review
H7 OrchestrationState từng step; retry giới hạn; analyzer fail thì dừng; fallback không bypass policy
+
+

Không tạo project manifest với command test giả. NEHOPS hiện chưa có test contract tự động — phải viết validator deterministic trước.

+
+ +
+ +
+
+

Để claim Level 4 trong môi trường enterprise Japan, thêm lớp production assurance

+
+ + + + + + + + + +
ControlRepo có gìĐiều kiện đóng thật
OIDCMock/local smoke + oauth2-proxy scaffoldEnterprise issuer/audience/JWKS, CA, group-role mapping, negative test
Tenant isolationPer-tenant state/RBAC/pre-auditIndependent review identity/app/db/network/KMS boundaries
KMS/WORMVault/S3 paths + preflight scriptsNon-exportable key, workload identity, Object Lock COMPLIANCE
HA/DRBackup/verify/restore workflowTopology, replication, RPO/RTO, restore + failover drill
SLA/reviewDraft + pentest scopeMeasured SLO, on-call/contract, independent signed report + retest
+
+
+ +
+ +
+
+

Chỉ tuyên bố Level 4 khi một run đầy đủ trả lời “Có” cho toàn bộ bảng này

+
+
    +
  • Attack injection/secret/PII bị chặn
  • +
  • High-risk denied by default
  • +
  • Approval có identity + reason + scope
  • +
  • Audit chain verify thành công
  • +
  • Trace/metrics/alerts hợp lệ và còn fresh
  • +
+
    +
  • Golden regression không drift ngoài ngưỡng
  • +
  • Tool/action không vượt allowlist
  • +
  • Retry/fallback không bypass governance
  • +
  • Evidence Pack verify được sau khi export
  • +
  • Không có Harness GAP hoặc evidence “not recorded”
  • +
+
+
Score là kết quả của evidence — không phải input do người đánh giá tự gõ.
+
+ +
+ +
+
+

CASAN đã đo token/cost, nhưng chưa có baseline tiết kiệm đáng tin cho NEHOPS

+
+
64
Runtime metric records trong repo CASAN: 85.202 token, cost estimate 0,08714 USD.
+
98
Provider usage records: 234.906 token, cost 0,16668 USD; gồm sample/local/test/API.
+
N/A
Không thể suy ra “% tiết kiệm” FPT/NEHOPS từ hai tập log hỗn hợp này.
+
+

Không cộng hai tập log thành tổng spend: chúng có thể là hai projection của cùng hoạt động và có dữ liệu test/demo.

+
+ +
+ +
+
+
+

Log thật giảm 2.446 → 24 “token” — nhưng kết quả phải bị từ chối

+
+
99,0%
Word-count bị loại bởi structural compressor trên analyzer-20260512.log.
+
0
Dòng [WRN] VBSource root not found còn lại trong compressed view.
+
+
+
+
COMPRESS mode=structural +in_tokens=2446 +out_tokens=24 +saved=2422 +ratio=0.0098 + +Compressed view chỉ còn: +Error Messages: 0 × 3
+
Kết luận: con số tiết kiệm cao không có giá trị nếu mất cảnh báo. Phải cấu hình must-keep + H3 faithfulness trước khi đưa vào model.
+
+
+ +
+ +
+
+

Giảm token đúng cách cho FPT Japan: giảm đọc lại, không giảm bằng chứng

+
+ + + + + + + + + +
Ưu tiênCơ chếEvidence cần giữ
1 · Index-firstĐọc common-knowhow README rồi chỉ mở file liên quanDanh sách file/pattern đã chọn
2 · Line-range contextPre-analysis trỏ chính xác AST lines cho vòng sauSource hash + cited ranges
3 · Structural viewGiữ summary/error/warning/must-keep; raw artifact giữ nguyênRaw hash + compressed hash + ratio
4 · Model routingLocal/deterministic cho classify; cloud cho task cần reasoningModel/provider + real usage
5 · Budget gatePer-screen/token ceiling, spike và slow-boil alertProvider telemetry + acceptance result
+
+
+ +
+ +
+
+

Dashboard token tốt phải ghép “chi phí” với “chất lượng đầu ra”

+
+
+ FinOps minimum +
    +
  • Provider tokens và billed cost / screen
  • +
  • Token coverage và cost coverage
  • +
  • P50/P95 latency, retries, fallback
  • +
  • Compression ratio + must-keep pass
  • +
+
+
+ Business quality +
    +
  • Round 1–4 defect/rework count
  • +
  • Hallucination/no-evidence findings
  • +
  • Cycle time / screen
  • +
  • Customer acceptance và escaped defect
  • +
+
+
+
Tối ưu token mà tăng rework là tối ưu giả.
+
+ +
+ +
+
+

Level 5 bắt đầu khi CASAN không còn là “project setup”, mà trở thành operating system dùng chung

+
+
    +
  • Harness package tái dùng và versioned đa project
  • +
  • Golden data phát hiện drift theo model/prompt version
  • +
  • Fallback có policy; failure có rollback
  • +
  • Tool registry có schema + idempotency
  • +
+
    +
  • Governance/policy signed và centrally controlled
  • +
  • AgentOps tổng hợp cross-project
  • +
  • Incident/reviewer feedback quay thành evaluation
  • +
  • Business KPI gắn với agent decision
  • +
+
+
+ +
+ +
+
+

Level 5 phải chứng minh được sáu “chuyện khó”, không chỉ bật sáu feature

+
+
1
ReuseÍt nhất hai project dùng cùng harness version.
+
2
DriftGolden dataset bắt thay đổi trước production.
+
3
FallbackĐổi model/tool path mà không bypass governance.
+
4
RollbackSide effect thất bại được phục hồi có evidence.
+
5
OutcomeKPI chứng minh giảm cycle time/rework/defect.
+
6
FederationDashboard và policy tập trung, signed, không sửa tùy ý tại project.
+
+
+ +
+ +
+
+

Thứ tự nâng cấp an toàn: chứng minh L4 trước, rồi mới federate lên L5

+
+
A
Project truthDomain Pack NEHOPS, validator, golden, corpus, traceability.
+
B
Level 4 runEnforce H4/H5/H6 quanh một screen; Evidence Pack verify.
+
C
Production overlayOIDC, tenant isolation, KMS/WORM, monitoring, DR, review.
+
D
ScaleReuse registry, central AgentOps, provider telemetry, signed policy.
+
E
LearnDrift/fallback/rollback và business feedback tự cải thiện bounded.
+
+
+ +
+ +
+
+
+

Live demo dùng đúng project thật, nhưng không đánh đổi an toàn

+
    +
  • Project hiện không có Git ⇒ snapshot trước
  • +
  • Mac có .NET 7; analyzer yêu cầu .NET 8/Windows EXE
  • +
  • Demo dùng AST/output FRR04701 đã có
  • +
  • Không rerun Roslyn analyzer trên Mac
  • +
  • Chỉ ghi CASAN scaffold sau khi người trình bày xác nhận
  • +
+
+
cd "/Users/thanhnguyen/Downloads/tmp-prj/BD/Basic Design (Screen&Report)_v2.7" + +# Trước buổi demo: snapshot hoặc đưa project vào Git. +casan init \ + --project nehops-bd-v27 \ + --edition devkit \ + --runtime managed \ + --client claude,codex \ + --mode enforce
+
+ +
+ +
+
+

Kịch bản 15 phút: cài → phát hiện attention → điền truth → chạy gate → xem evidence

+
+
1
InitCho xem file được tạo và maturity=not_assessed.
+
2
DoctorTrust Codex hooks; giải thích provider telemetry chưa có.
+
3
Preflight promptAgent phát hiện assessment misclassification, không sửa file.
+
4
Domain truthThay placeholder bằng FR/architecture/golden/corpus/traceability NEHOPS.
+
5
Gate/evidenceChạy H4 attack + safe AST validation; mở report/evidence.
+
+
+ +
+ +
+
+
+

Một governed step an toàn: xác minh AST JSON có thể đọc được

+

Đây chỉ là smoke action H2/H6 — chưa thay thế bộ validator chất lượng BD.

+
+
printf 'Validate FRR04700 CL AST JSON\n' \ + > /tmp/nehops-casan-input.txt + +casan run \ + /tmp/nehops-casan-input.txt \ + /tmp/nehops-casan-output.txt \ + validate-frr04700-cl-ast -- \ + python3 -c ' +import json, os +p="AST-outputs/FRR04700_CL/FRR04700_CL_AST.json" +json.load(open(p, encoding="utf-8")) +open(os.environ["CASAN_OUTPUT"],"w").write("AST_JSON_VALID\n") +' + +casan report latest
+
+ +
+ +
+
+
python3 \ + /path/to/casan-harness/scripts/bash/context-compress.py \ + --mode structural \ + --input logs/analyzer-20260512.log + +2446 → 24 · ratio 0.0098 +Nhưng warning quan trọng biến mất.
+
+

Demo “savings” phải kết thúc bằng một lần REJECT

+
    +
  • Cho khán giả thấy raw warning
  • +
  • Cho thấy compressed view bỏ warning
  • +
  • Giải thích must-keep/faithfulness
  • +
  • Không ghi 99% vào KPI dự án
  • +
+
+
+ +
+ +
+
+

Những phần chưa được phép quảng cáo là hoàn chỉnh

+
+ + + + + + + + + +
AreaTrạng thái trung thựcViệc tiếp theo
Enterprise editionFuture; init chủ động từ chốiKhông bán/claim như shipped product
OIDC/KMS/WORMLocal lab/scaffold + production pathsDeploy managed infra và giữ evidence khách hàng
HA/SLA/pentestRunbook/draft/scope, chưa có kết quả vận hành độc lậpDR drill, measured SLO, contract, assessor report
Token savingsCompressor MVP; no NEHOPS end-to-end baselineFix warning/must-keep; A/B cùng acceptance
NEHOPS maturityAssessment cũ mâu thuẫn; init ⇒ not_assessedChạy assessment evidence-backed mới
+
+
+ +
+ +
+
+

Với FPT Japan, “đúng kỹ thuật” phải đi cùng “đúng nguồn, đúng ngôn ngữ, đúng trách nhiệm”

+
+
    +
  • Japanese output và terminology consistency
  • +
  • No-inference: thiếu evidence phải ghi UNKNOWN
  • +
  • Legacy traceability: VB → AST → BD → review
  • +
  • Data residency/model route rõ ràng
  • +
+
    +
  • Human oversight cho publish/override
  • +
  • APPI/data-processing register theo tenant
  • +
  • Incident/SLA escalation bằng run ID/evidence hash
  • +
  • Security review bám đúng deployed image/config
  • +
+
+

Japan AI governance đang nhấn mạnh governance, guideline và audit — CASAN phải biến các yêu cầu đó thành runtime evidence.

+
+ +
+ +
+
+

Bảy nguyên tắc làm việc của core team

+
+
    +
  • Evidence trước claim
  • +
  • Fail-closed trước convenience
  • +
  • Không sửa gate riêng theo project
  • +
  • Không dùng mock làm proof production
  • +
+
    +
  • Không tối ưu token bằng cách mất nghĩa
  • +
  • Không để model tự tuyên bố DONE
  • +
  • Mỗi incident trở thành golden/red-team test
  • +
+
+
+ +
+ +
+
+

Ba quyết định cần chốt để NEHOPS tiến tới Level 4

+
+
01
Owner: ai sở hữu Domain Pack, golden và corpus tiếng Nhật?
+
02
Gate: deterministic validator nào thay cho “LLM review thấy ổn”?
+
03
Platform: Control Plane/IdP/KMS/WORM dùng dịch vụ nào tại FPT Japan?
+
+

Chỉ sau ba quyết định này mới lập kế hoạch Level 4 có owner, acceptance evidence và deadline.

+
+ +
+ +
+
+ Thông điệp cuối +
+

AI-Native không có nghĩa AI được tự do.
Nó có nghĩa tổ chức kiểm soát AI bằng evidence ở quy mô lớn.

+

Level 4: automation có security, governance, AgentOps. Level 5: hệ thống học và mở rộng — nhưng vẫn bị ràng buộc bởi policy, rollback và business outcome.

+
+ +
+
+ + + + + + + + + diff --git a/docs/plans/CASAN_NEHOPS_QA_REMEDIATION_AND_DEMO_RUNBOOK_VI.md b/docs/plans/CASAN_NEHOPS_QA_REMEDIATION_AND_DEMO_RUNBOOK_VI.md new file mode 100644 index 0000000..3f8cf18 --- /dev/null +++ b/docs/plans/CASAN_NEHOPS_QA_REMEDIATION_AND_DEMO_RUNBOOK_VI.md @@ -0,0 +1,1842 @@ +# CASAN × NEHOPS — Q&A, kế hoạch khắc phục và runbook demo thực tế + +> Tài liệu đào tạo dành cho CASAN core member +> Project demo: `/Users/thanhnguyen/Downloads/tmp-prj/BD/Basic Design (Screen&Report)_v2.7` +> Use case chính: `frr04701 · 部屋状況リスト` +> CASAN CLI đã kiểm tra: `casan 1.0.7` + +## 0. Mục tiêu và nguyên tắc sử dụng + +Tài liệu này có ba mục tiêu: + +1. Tổng hợp lại các câu hỏi và câu trả lời liên quan đến slide CASAN. +2. Chia nhỏ kế hoạch xử lý triệt để lỗi context compression đã phát hiện trên log NEHOPS. +3. Cung cấp runbook demo trên project thật, gồm từng bước, prompt copy/paste, kết quả mong đợi và điều kiện được phép chuyển sang bước tiếp theo. + +Các nguyên tắc bắt buộc: + +- Không coi `casan init` là chứng nhận maturity. +- Không biến placeholder thành evidence. +- Không tạo `project.manifest.json` với command build/test giả. +- Không tuyên bố token saving nếu bản nén làm mất thông tin quan trọng. +- Không tự động chạy full pipeline ngay ở prompt đầu tiên. +- Không suy diễn khi source, AST hoặc tài liệu tham chiếu không có evidence. +- Không sửa legacy VB source chỉ để làm cho CASAN report “xanh”. +- Mọi claim phải ghi rõ scope: workstation, project workflow, release hay enterprise platform. + +## Cách dùng tài liệu khi đào tạo + +### Đọc theo nhu cầu + +- Cần giải đáp slide: đọc Phần I, mục 2–14. +- Cần giao task sửa token compression: đọc Phần II, mục 15–18. +- Cần chuẩn bị live demo: đọc Phần III, mục 19–37. +- Cần prompt copy/paste: mục 21, 25, 26, 27, 28 và 36. +- Cần checklist claim Level 4: mục 32–34. + +### Kịch bản demo 75 phút + +| Thời lượng | Nội dung | Thứ core member phải nhìn thấy | +|---:|---|---| +| 5 phút | Edition ≠ readiness ≠ maturity | Init không tạo Level | +| 10 phút | Preflight read-only | Scope, evidence và blocker được tìm ra mà không sửa file | +| 10 phút | Init ở `observe` | Core/Domain/Telemetry là ba trạng thái độc lập | +| 10 phút | Trace FRR04701 | VB source → AST FRR04700 → pre-analysis → output → Round 4 Pass | +| 10 phút | Assessment contradiction | H5=35 không dưới 30; kết luận Level 2 không khớp rubric | +| 10 phút | Compression defect | `[WRN]` bị mất nên “2.446 → 24” phải REJECT | +| 10 phút | Validator và H4/H5/H6 | LLM review không thay deterministic gate | +| 5 phút | Evidence campaign | Một run demo không đủ claim stability | +| 5 phút | Level 4 → Level 5 | Project runtime assurance → shared operating platform | + +Nếu chỉ có 30 phút: + +1. Giải thích ba khái niệm edition/readiness/maturity. +2. Trace FRR04701 qua năm artifact thật. +3. Demo compression làm mất `[WRN]`. +4. Kết thúc bằng checklist trước khi được claim Level 4. + +## 1. Các nguồn đã kiểm tra + +### CASAN + +- `casan --help`: phân biệt `Core`, `Domain Pipeline` và `Provider Telemetry`. +- `casan init --help`: hỗ trợ `observe`/`enforce`, managed/vendored runtime và các client Claude/Codex. +- `docs/packaging/ADOPTION_GUIDE.md`: adoption, readiness, report và Domain Pack. +- `docs/packaging/DOMAIN_PACK_GUIDE.md`: cấu trúc Domain Pack và điều kiện dùng manifest. +- `docs/packaging/CASAN_PACKAGING_PLAN.md`: edition độc lập với maturity. +- `packages/casan-harness/scripts/bash/context-compress.py`: implementation compression hiện tại. +- `packages/casan-harness/tests/phase08-compression-tests.sh`: test compression hiện tại. +- `docs/plans/CASAN_PLAN_08_CONTEXT_COMPRESSION.md`: kiến trúc và thứ tự H4/H5/H3/H6 mong muốn. + +### NEHOPS + +- [CLAUDE.md]() +- [BD_Instruction.md]() +- [boss.md]() +- [analyze.md]() +- [generation.md]() +- [review-full.md]() +- [NEHOPS CASAN assessment]() +- [FRR04701 pre-analysis]() +- [FRR04701 output ban đầu]() +- [FRR04701 final review]() +- [FRR04701 output đã review]() + +## Phần I — Tổng hợp Q&A + +## 2. Edition, readiness và maturity + +### 2.1 Edition là gì? + +Edition mô tả thứ CASAN đóng gói và giao cho project: + +- CLI; +- Core runtime; +- hook Claude/Codex; +- template; +- Domain Pack scaffold; +- local viewer hoặc Control Plane tùy edition. + +Edition trả lời câu hỏi: “Bản cài đặt này có những capability nào?” + +### 2.2 Maturity là gì? + +Maturity mô tả năng lực đã được chứng minh bằng các run thật: + +- Workflow nào đã chạy? +- Gate nào đã thực thi? +- Gate có block được case xấu không? +- Ai phê duyệt? +- Audit/evidence có đầy đủ và còn nguyên không? +- Token/cost là số thật hay estimate? + +Maturity trả lời câu hỏi: “Tổ chức/project đã vận hành capability đó đáng tin đến đâu?” + +### 2.3 `level: null`, `status: not_assessed` nghĩa là gì? + +Không phải Level 0 và không phải lỗi cài đặt. + +Nó có nghĩa: + +- `casan init` chưa thực hiện assessment L1–L5; +- CASAN chưa có đủ evidence để gán maturity; +- hệ thống từ chối suy diễn maturity chỉ vì scaffold đã tồn tại. + +### 2.4 `ready_with_attention` có mâu thuẫn với `not_assessed` không? + +Không. + +- `ready_with_attention`: Core đã có route hoạt động nhưng còn việc cần xử lý. +- `not_assessed`: chưa có assessment maturity dựa trên evidence. + +Một project có thể đủ sẵn sàng để pilot nhưng chưa được phép claim Level 4. + +### 2.5 `ATTN` là gì? + +`ATTN` là “attention required”: cần xác nhận hoặc xử lý, nhưng chưa chắc là fatal. + +Trong probe NEHOPS, các điểm attention gồm: + +- Codex hook chưa được trust/kích hoạt; +- Domain Pack có scaffold nhưng chưa có manifest thật; +- Provider telemetry chưa có hoặc mới ở trạng thái optional; +- Core có thể chạy qua Claude route nhưng Domain Pipeline chưa được chứng minh. + +## 3. Level 4 và Level 5 + +### 3.1 Level 4 có phải chỉ cần H4/H5/H6? + +Không. + +H4/H5/H6 là bottleneck quan trọng: + +- H4: runtime security; +- H5: governance, approval và immutable/auditable evidence; +- H6: telemetry, token, cost, latency, retry, failure và alert. + +Nhưng H1–H3 và H7 vẫn phải chạy quanh workflow thật. Ví dụ: + +- H4/H5/H6 đều có report nhưng requirement là placeholder → chưa đạt. +- Security gate tốt nhưng test command là giả → chưa đạt. +- Telemetry đầy đủ nhưng output không có validator → chưa đạt. +- Điểm trung bình cao nhưng một critical harness vẫn GAP → maturity bị giới hạn. + +### 3.2 Drift + +Drift là thay đổi hành vi so với baseline: + +- đổi model làm output khác; +- prompt mới làm chất lượng Mermaid giảm; +- token tăng dần; +- compression làm mất warning; +- false-positive của H4 tăng sau khi đổi rule. + +Drift phải được đo với golden baseline và threshold, không phải mọi khác biệt đều là lỗi. + +### 3.3 Fallback + +Fallback là đường xử lý thay thế có policy: + +- provider chính timeout → provider được phép dự phòng; +- compression fail → dùng raw context; +- analyzer chưa chạy được → dùng artifact đã được ký/xác nhận từ Windows runner. + +Fallback vẫn phải qua H4/H5/H6, không được dùng để bypass gate. + +### 3.4 Rollback + +Rollback hoàn tác side effect khi action thất bại: + +- khôi phục file; +- không publish artifact fail validator; +- phục hồi policy version trước; +- đảo deployment/migration; +- đánh dấu output invalid để downstream không sử dụng. + +### 3.5 KPI + +KPI Level 5 phải gắn với kết quả: + +- thời gian tạo một Basic Design được chấp nhận; +- số vòng rework; +- lỗi thoát sang downstream; +- token/cost trên một artifact được chấp nhận; +- tỷ lệ reviewer Nhật chấp nhận ngay vòng đầu; +- tỷ lệ traceability đầy đủ; +- tỷ lệ warning quan trọng bị bỏ sót. + +### 3.6 Central governance + +Central governance trả lời tập trung: + +- Ai được chạy action nào? +- Trên tenant/project/customer nào? +- Model/tool nào được phép? +- Path nào được ghi? +- Budget bao nhiêu? +- Action nào cần người phê duyệt? +- Ai được override? +- Policy version/hash nào áp dụng cho run? + +## 4. Những thứ `casan init` không thể tự provision + +### 4.1 Enterprise OIDC + +OIDC liên kết CASAN với Identity Provider doanh nghiệp: + +- xác thực danh tính; +- token có chữ ký; +- issuer/audience/JWKS; +- mapping group/claim sang CASAN role; +- thu hồi quyền khi nhân sự thay đổi. + +`casan init` có thể tạo chỗ cấu hình nhưng không thể tự đăng ký enterprise application, cấp secret/certificate hoặc quyết định role mapping. + +### 4.2 CA + +Trong slide, CA nên được hiểu là Certificate Authority, không phải Conditional Access. + +CA cấp và quản lý certificate TLS/mTLS: + +- certificate cho API/Control Plane; +- service-to-service identity; +- expiry, rotation và revoke. + +Nếu slide muốn nói Microsoft Conditional Access thì cần ghi đầy đủ để tránh nhầm. + +### 4.3 Private network + +Control Plane/API không mở trực tiếp ra Internet: + +- VPC/VNet/private subnet; +- VPN/Direct Connect/ExpressRoute; +- private endpoint; +- reverse proxy/API gateway; +- firewall/allowlist. + +### 4.4 KMS và HSM + +KMS: + +- quản lý khóa mã hóa/ký; +- IAM và audit; +- rotation; +- application không cần giữ plaintext key. + +HSM: + +- key không rời phần cứng; +- phù hợp khi hợp đồng/regulation yêu cầu; +- đắt và phức tạp hơn KMS. + +Không bắt buộc HSM nếu KMS đã đáp ứng scope và hợp đồng. + +### 4.5 S3 Object Lock/WORM + +WORM bảo đảm evidence không thể bị sửa/xóa trong retention period. + +Nó phù hợp để giữ: + +- audit chain head; +- evidence pack; +- policy bundle; +- output hash; +- approval record. + +### 4.6 HA, RPO/RTO và failover + +- HA: nhiều instance/failure domain để một node hỏng không làm dừng dịch vụ. +- RPO: tối đa được mất bao nhiêu dữ liệu. +- RTO: tối đa được ngừng dịch vụ bao lâu. +- Failover: chuyển workload/traffic sang standby/AZ/region khác. + +Chỉ được claim failover khi đã diễn tập và có evidence. + +### 4.7 On-call, uptime SLA và service credit + +- On-call: người trực, severity, escalation, incident response. +- Uptime SLA: cam kết hợp đồng về availability. +- Service credit: bồi hoàn theo hợp đồng nếu không đạt SLA. + +Dashboard không có người nhận alert thì chưa phải on-call operation. + +### 4.8 External pentest và certification + +- Pentest: bên độc lập tấn công scope đã thống nhất, report finding và retest. +- Certification: audit tiêu chuẩn theo pháp nhân, quy trình và service scope. + +Hai việc không thay thế cho nhau. + +## 5. Scaffold phải được thay bằng sự thật của NEHOPS + +Các file init sinh ra mới chỉ là hình dạng: + +| Scaffold | Phải thay bằng evidence thật | +|---|---| +| `input/requirement.md` | Rule thật từ `CLAUDE.md`, `BD_Instruction.md`, review instructions | +| `input/architecture.md` | VB.NET → AST → pre-analysis → BD → review; constraint Windows/.NET | +| `golden-runs/` | FRR04701 đã Pass Round 4 | +| `corpus/` | Japanese benign/red-team, VB comment/ref-doc injection, secret/PII case | +| `traceability-map.json` | Requirement → instruction/source/AST → validator/review evidence | +| `domain-pack.yaml` | Project ID, owner và threshold thật | +| `project.manifest.json` | Chỉ tạo sau khi đã có command validate/build/test thật | +| CI template | Runner, OS, command và evidence upload thật | + +Nếu file vẫn là placeholder: + +- không tính file đó là evidence; +- ghi trạng thái `incomplete/not_evidence`; +- không cho maturity assessment sử dụng nó. + +## 6. Preflight read-only không có nghĩa là không bao giờ sửa + +Trình tự đúng: + +1. Preflight read-only để hiểu project và đề xuất thay đổi. +2. Human review xác nhận đề xuất. +3. Configuration phase cho phép ghi vào CASAN-owned paths. +4. Pilot `observe`. +5. Đóng gap. +6. Chuyển `enforce`. + +Prompt đầu tiên read-only giúp tránh một agent vừa vào repo 2,6 GB đã: + +- sửa nhầm instruction; +- chạy analyzer sai môi trường; +- tạo manifest giả; +- thay đổi legacy source; +- ghi output vào sai path. + +## 7. AST, pre-analysis, output và final review + +### AST + +Biểu diễn cấu trúc VB source: + +- class/method/field; +- event handler; +- UI control; +- SQL; +- dependency/call graph; +- metrics. + +FRR04701 thuộc component `FRR04700`, vì vậy AST hiện có nằm dưới: + +- `AST-outputs/FRR04700_CL/`; +- `AST-outputs/FRR04700_SV/`. + +### Pre-analysis + +Inventory evidence trước khi generate: + +- event/API/DB/branch; +- source/AST line range; +- common knowhow mapping; +- phần không đủ evidence. + +### Output + +Basic Design tiếng Nhật tại: + +`BD/output/frr04701/機能フロー_frr04701_部屋状況リスト.md` + +### Final review + +Round 1–4 được ghi trong một report: + +`BD/review/review_report/frr04701/frr04701_ReviewReport.md` + +Round 4 hiện kết luận `Pass`. Bản cuối: + +`BD/output_reviewed/frr04701/機能フロー_frr04701_部屋状況リスト.md` + +### Có được coding ngay không? + +Chưa. Workflow hiện tại kết thúc ở Basic Design. + +Implementation Angular/Java cần workflow downstream riêng: + +- target architecture; +- API/data contract; +- Detailed Design; +- acceptance criteria; +- test contract; +- repository/path được phép ghi; +- governed implementation action; +- build/test/security validator. + +Không được biến `/bd:boss` thành codegen vì command và instruction hiện tại không định nghĩa việc đó. + +## 8. Lỗi phân loại trong assessment NEHOPS + +Điểm hiện tại: + +| Harness | Điểm | +|---|---:| +| H1 | 70 | +| H2 | 68 | +| H3 | 75 | +| H4 | 25 | +| H5 | 35 | +| H6 | 20 | +| H7 | 65 | +| Trung bình | 51,1 | + +Assessment viết “3 harness dưới 30”, nhưng thực tế chỉ: + +- H4 = 25; +- H6 = 20. + +H5 = 35, không dưới 30. + +Theo rubric trong chính assessment, average 40–65 và không quá hai GAP dẫn đến Level 3, trong khi tài liệu kết luận Level 2. Vì vậy: + +- số điểm có thể dùng làm dữ liệu tham khảo; +- kết luận maturity không được copy; +- cần thống nhất lại rubric/GAP definition và chạy assessment có evidence. + +NEHOPS không phải hoàn toàn không có H4/H5/H6: + +- H4 có instruction/control thủ công; +- H5 có review loop/output contract; +- H6 có analyzer log. + +Nhưng chúng chưa trở thành runtime assurance đầy đủ. + +## 9. H5 có phải sửa source NEHOPS? + +Phần lớn không sửa legacy VB business source. + +H5 được khai báo trong: + +- `.casan/`; +- `apps/nehops-bd-v27/domain/`; +- risk registry; +- traceability; +- golden/corpus; +- validator; +- CI; +- approval/role mapping. + +Có thể thêm project-owned validator/adapter, nhưng không sửa business logic chỉ để CASAN pass. + +## 10. H6 cần đo gì? + +### 10.1 Per-step token + +- input tokens; +- output tokens; +- total tokens; +- run ID; +- step ID; +- model/provider/version; +- timestamp. + +### 10.2 Cost source + +- `estimate`: word count/tokenizer/bảng giá cấu hình; +- `provider`: usage/billing telemetry cho đúng request. + +Mỗi số cần có `source`, `currency` và `pricing_version`. + +### 10.3 Reliability + +- latency; +- retry; +- exit code; +- runtime failure; +- semantic failure. + +Exit code 0 không chứng minh output đúng. + +### 10.4 Alert classes + +- token-overuse: vượt ngưỡng tuyệt đối; +- spike: tăng đột ngột; +- slow-boil: tăng từ từ qua nhiều run; +- alert: event có severity, owner và action. + +### 10.5 Freshness và coverage + +```text +coverage = số AI step có telemetry hợp lệ / tổng AI step +``` + +Freshness là tuổi của telemetry so với run cần đánh giá. + +Telemetry cũ hoặc thiếu coverage không được dùng để claim cost/saving. + +## 11. Deterministic validator và full pipeline + +CASAN Core có thể chạy khi chưa có project validator. + +Nhưng Domain Pipeline và Level 4 claim không được coi là hoàn chỉnh khi: + +- không có validator; +- manifest trỏ vào command giả; +- chỉ có LLM tự review output của chính nó. + +Validator NEHOPS nên kiểm tra: + +- AST JSON parse/schema; +- section bắt buộc; +- output path/tên file; +- lowercase screen ID; +- đúng một Basic Design Markdown; +- event inventory; +- Mermaid structure/syntax; +- participant order; +- API table; +- forbidden VB identifier; +- traceability tới source/AST; +- không ghi ngoài allowed path; +- review report có đủ Round 1–4; +- Round 4 có verdict hợp lệ. + +## 12. Điều kiện Level 4 enterprise và trách nhiệm + +| Bên | Cần cung cấp | +|---|---| +| Organization/platform/security | IdP, cloud account, VPC, CA/certificate, KMS/HSM nếu cần, WORM, HA/DR, monitoring, policy | +| CASAN core team | Control Plane, role mapping, tenant isolation, H4/H5/H6 integration, dashboard, runbook, evidence | +| Project team | Domain Pack, golden/corpus, traceability, validator, CI, risk registry | +| Business/customer | Data classification, approver, retention, RPO/RTO, SLA, acceptance | +| External assessor | Pentest/retest hoặc audit/certification theo scope | + +Các control platform dùng chung không cần copy vào từng project. Project phải dẫn chiếu inherited control và chứng minh đang dùng đúng control đó. + +### Chi phí cần lập budget + +- KMS: thường thấp ở tải nhỏ; AWS công bố `$1/key/tháng` cộng request. +- CloudHSM: cao; ví dụ AWS với hai HSM tại `us-east-1` là khoảng `$2.380,80/tháng`, Tokyo cần báo giá theo region. +- AWS Private CA: `$50/CA/tháng` cho short-lived hoặc `$400/CA/tháng` cho general-purpose, cộng certificate. +- WORM: storage, request, retention, replication và KMS. +- OIDC: phụ thuộc license IdP hiện có. +- HA/DR: compute, database, storage và network dự phòng. +- On-call: chủ yếu là chi phí nhân sự. +- Pentest/certification: phải lấy báo giá theo scope. + +Không được cộng các con số ví dụ này thành “giá CASAN Level 4” khi chưa có topology và requirement thật. + +Nguồn giá: + +- +- +- + +## 13. Cần bao nhiêu run để claim Level 4? + +Một run có thể chứng minh mechanics cho một release/use case. Nó chưa chứng minh stability. + +Level 4: + +- cần nhiều run đại diện trong cùng project/workflow; +- bao phủ benign, failure, adversarial và recovery; +- không bắt buộc phải có nhiều project nếu claim chỉ dành cho NEHOPS. + +Level 5: + +- cần chứng minh reuse và central governance trên nhiều project; +- cần drift/KPI/fallback/rollback ở quy mô tổ chức. + +Mỗi run cần: + +- project/scope/commit hoặc input snapshot; +- CASAN version/hash; +- model/provider/version; +- prompt/policy/golden/corpus version; +- gate results; +- approver; +- token/cost/latency/retry; +- output hash; +- audit/evidence pack. + +Các công thức cơ bản: + +```text +Gate pass rate += số run vượt toàn bộ mandatory gate / tổng run đủ điều kiện + +Telemetry coverage += số AI step có telemetry hợp lệ / tổng AI step + +Evidence completeness += số run có evidence pack đầy đủ / tổng run + +Cost per accepted artifact += tổng provider cost / số artifact được chấp nhận + +Token saving += 1 - tokens_new / tokens_baseline +``` + +Token saving chỉ hợp lệ khi baseline và candidate: + +- cùng input/use case; +- acceptance criteria tương đương; +- output đều pass cùng validator; +- không mất warning/requirement; +- model/pricing difference được công bố. + +## 14. `context-compress.py` có được tính không? + +Có thể tính là capability đã implement, chưa được tính là saving đã chứng minh. + +Implementation hiện tại: + +- deterministic; +- `dedup`, `extractive`, `structural`; +- must-keep file; +- failed-run raw passthrough; +- nhận diện severity `[ERR]/[WRN]` và failure/skip signal; +- kiểm tra multiplicity, không cho một warning còn lại che warning thứ hai bị mất; +- preservation fail trả raw + exit 1 hoặc halt theo policy; +- JSON evidence có hash raw/candidate/output; +- metric ratio được ghi rõ là `whitespace_estimate`; +- policy switch. + +Phần đã khắc phục ngày 2026-07-29: + +- `[WRN]` trong log NEHOPS được giữ; +- invalid must-keep regex fail-closed; +- candidate rỗng bị reject; +- partial must-keep loss bị reject; +- rejected compression không còn báo saving; +- test fixture NEHOPS và evidence report đã được bổ sung. + +Phần còn lại: + +- `estimate_tokens()` vẫn là `len(text.split())`, đã được label đúng nhưng chưa phải provider telemetry; +- must-keep fixture chưa được wiring thành Domain Pack policy trong project NEHOPS thật; +- H4 raw/compressed scan và H5 audit binding vẫn phải nối ở pipeline caller; +- chưa có A/B end-to-end FRR04701 với cùng validator; +- slide cũ chưa được cập nhật trong task sửa script này. + +Vì vậy capability đã an toàn hơn, nhưng vẫn chưa phải bằng chứng ROI/token saving hoàn chỉnh. + +## Phần II — Kế hoạch xử lý triệt để context compression + +## 15. Defect statement + +Log thật: + +`/Users/thanhnguyen/Downloads/tmp-prj/BD/Basic Design (Screen&Report)_v2.7/logs/analyzer-20260507.log` + +có các dòng quan trọng: + +- `[INF] - Error Messages: 0` +- `[WRN] VBSource root not found ... common code tracing skipped` + +Trước bản sửa 2026-07-29, structural compression giữ summary `Error Messages: 0` +nhưng có thể bỏ `[WRN]`. + +Kết quả “2.446 → 24 token” quan sát trước đây phải bị REJECT vì: + +- số “token” chỉ là word-count estimate; +- warning ảnh hưởng độ đầy đủ của AST đã mất; +- downstream có thể hiểu sai rằng analyzer hoàn toàn thành công. + +Sau bản sửa: + +- log thật `analyzer-20260507.log` giữ đủ 4/4 `[WRN]`; +- phép đo hiện tại là `1.620 → 76` whitespace-estimate units; +- preservation pass, missing = 0; +- đây vẫn chỉ là estimate và chưa được claim ROI khi chưa có H3 A/B end-to-end. + +## 16. Nguyên tắc fix + +Thứ tự an toàn bắt buộc: + +```text +RAW tool output +→ H4 scan raw +→ H5 hash/store raw +→ token budget decision +→ compression +→ H3 preservation/faithfulness gate +→ H4 scan compressed +→ H5 hash/audit compressed + ratio + policy +→ model context +``` + +Nếu bất kỳ preservation gate nào fail: + +```text +REJECT compressed view +→ sử dụng raw view hoặc dừng theo policy +→ ghi alert/evidence +``` + +## 17. Backlog chia nhỏ + +### CC-01 — Đóng băng defect thành regression fixture + +Trạng thái: **DONE trong CASAN Core tests**. + +Mục tiêu: + +- Không test bằng vài dòng giả duy nhất. +- Giữ một bản log fixture đã loại bỏ dữ liệu nhạy cảm nhưng bảo toàn cấu trúc severity. + +File dự kiến: + +- `packages/casan-harness/tests/fixtures/context-compression/nehops/analyzer-20260507.log` +- `packages/casan-harness/tests/fixtures/context-compression/nehops/must-keep.patterns` + +Acceptance: + +- fixture có `[WRN]`, `[INF]`, summary, output path và component count; +- fixture không chứa secret/PII; +- test cũ tái hiện warning bị mất trước khi fix. + +### CC-02 — Nhận diện severity tag + +Trạng thái: **DONE**. + +Sửa: + +- `packages/casan-harness/scripts/bash/context-compress.py` + +Phải nhận diện ít nhất: + +- `[ERR]`, `[ERROR]`, `ERROR`; +- `[WRN]`, `[WARN]`, `WARNING`; +- exception/panic/failure/blocked/denied; +- case-insensitive và có timestamp/prefix. + +Acceptance: + +- mọi dòng `[WRN]` trong fixture được giữ; +- `[INF]` bình thường không bắt buộc giữ, trừ summary/must-keep; +- không coi mọi `[INF]` là important, tránh làm nén mất tác dụng. + +### CC-03 — Domain-specific must-keep policy + +Trạng thái: **PARTIAL** — đã có sanitized NEHOPS fixture/pattern trong test; chưa +wiring vào Domain Pack của project NEHOPS thật. + +Must-keep cho NEHOPS: + +- `VBSource root not found`; +- `common code tracing skipped`; +- `Error Messages`; +- output directory/path; +- component/layer ID; +- analyzer version; +- exit code; +- số file AST sinh ra; +- trạng thái CL/SV; +- fail/skip/partial/incomplete. + +Không hard-code toàn bộ từ khóa NEHOPS vào Core nếu có thể biểu diễn bằng policy file. + +Acceptance: + +- Core severity rule giữ warning chung; +- Domain Pack bổ sung invariant riêng cho NEHOPS; +- policy/version/hash được ghi vào evidence. + +### CC-04 — Preservation gate + +Trạng thái: **DONE ở compressor**. + +Hiện `--require-must-keep-file` chỉ kiểm tra regex được cung cấp. Cần mở rộng test và contract: + +- pattern tồn tại trong raw thì phải tồn tại trong compressed; +- pattern không có trong raw không nên tạo false failure; +- giữ count của severity; +- giữ summary quan trọng; +- phát hiện partial-loss, không chỉ “còn ít nhất một warning”. + +Acceptance: + +- bỏ một trong hai warning CL/SV → gate fail; +- sửa nội dung warning → gate fail hoặc mismatch rõ; +- compressed view pass khi mọi invariant còn nguyên. + +### CC-05 — Fail-closed/fallback + +Trạng thái: **DONE ở compressor**. + +Khi compression: + +- crash; +- input quá lớn; +- must-keep fail; +- H3 fail; +- H4 compressed fail; + +thì không được đưa bản nén lỗi vào model. + +Policy cần phân biệt: + +- `fallback_raw`: sử dụng raw sau khi raw đã qua H4/H5; +- `halt`: dừng khi raw không được phép vào model; +- `manual_review`: yêu cầu người phê duyệt. + +Acceptance: + +- output nén fail không trở thành downstream context; +- alert ghi rõ fallback/halt; +- raw artifact vẫn còn để audit/debug. + +### CC-06 — H4 quét cả raw và compressed + +Trạng thái: **PENDING pipeline integration**. Compressor phát cờ evidence yêu cầu +downstream scan nhưng không tự thay thế H4 caller. + +Test: + +- raw có prompt injection nhưng compression làm mất payload → raw scan vẫn block; +- raw benign nhưng compressed chứa chuỗi nguy hiểm → compressed scan block; +- secret/PII không được che bằng compression. + +Acceptance: + +- không có đường bypass H4 qua compression; +- raw scan result và compressed scan result cùng nằm trong evidence. + +### CC-07 — H5 hash và audit binding + +Trạng thái: **PARTIAL** — JSON report đã bind hash raw/candidate/output; pipeline +vẫn cần ký/bind report vào audit chain H5. + +Evidence cần bind: + +- raw SHA-256; +- compressed SHA-256; +- compressor version/hash; +- mode; +- policy hash; +- input/output estimated token; +- provider token nếu có; +- preservation result; +- fallback decision. + +Acceptance: + +- sửa raw hoặc compressed sau run → verification fail; +- không thể ghép raw của run A với compressed của run B. + +### CC-08 — Sửa semantics H6 + +Trạng thái: **DONE ở compressor report**; provider telemetry/A-B vẫn pending. + +Đổi cách trình bày: + +- `estimated_units` hoặc `estimated_tokens`, không gọi là billed tokens; +- `source=whitespace_estimate`; +- provider usage ghi riêng; +- saving claim chỉ xuất hiện nếu quality/preservation pass. + +Acceptance: + +- report không hiển thị “99% saving” cho rejected compression; +- coverage/freshness hiển thị; +- estimate và provider không bị cộng trùng. + +### CC-09 — Mở rộng automated tests + +Trạng thái: **DONE cho compressor regression** — phase08 19/0, SEC-09 7/0, +loop consumer 16/0. + +Sửa: + +- `packages/casan-harness/tests/phase08-compression-tests.sh` + +Thêm test: + +1. NEHOPS `[WRN]` preserved. +2. Multiple CL/SV warnings preserved. +3. Summary preserved. +4. Missing one warning → fail. +5. Failed command → raw passthrough. +6. Raw injection cannot be hidden by compression. +7. Compressed injection is blocked. +8. Policy disabled → raw passthrough. +9. Oversize/unreadable input → fail-closed. +10. No false saving claim when rejected. + +Acceptance: + +- phase08 suite xanh; +- `ci-harness-gate.sh` vẫn gọi `phase08-compression`; +- không làm regression các suite H4/H5/H6 liên quan. + +### CC-10 — A/B end-to-end trên NEHOPS + +Trạng thái: **PENDING**. + +Baseline: + +- raw log/context; +- FRR04701 accepted artifact; +- cùng validator và acceptance criteria. + +Candidate: + +- compressed context sau fix. + +So sánh: + +- warning retention; +- output validator result; +- Round 4/reviewer result; +- provider input/output token; +- latency; +- retry; +- cost; +- traceability. + +Acceptance: + +- cả baseline và candidate đều pass cùng quality gate; +- must-keep 100%; +- không tăng escaped defect; +- chỉ khi đó mới báo saving. + +### CC-11 — Cập nhật tài liệu và slide + +Trạng thái: **PARTIAL** — Plan 08/backlog/runbook đã cập nhật; HTML slide chưa sửa +trong task code này. + +Sau khi code/test hoàn tất: + +- đổi slide “2.446 → 24” thành defect demonstration; +- hiển thị `REJECTED`, không hiển thị như thành tích; +- thêm kết quả sau fix; +- ghi rõ estimate/provider; +- cập nhật Plan 08 theo trạng thái thực. + +## 18. Definition of Done cho compression fix + +- [x] `[WRN]` NEHOPS được giữ trong Core compressor/test. +- [ ] Must-keep policy được wiring vào Domain Pack NEHOPS thật. +- [x] Preservation gate kiểm tra multiplicity, không chỉ một regex mẫu. +- [x] Fail-closed hoặc fallback raw theo policy. +- [ ] H4 scan raw và compressed trong pipeline thật. +- [ ] H5 bind report/hash vào signed audit chain. +- [x] H6 report phân biệt estimate/provider. +- [x] Sanitized fixture từ cấu trúc log thật. +- [x] Phase08 và suite liên quan xanh. +- [ ] A/B FRR04701 qua cùng validator. +- [x] Compressor không claim saving cho run bị reject. + +## Phần III — Runbook demo NEHOPS trên project thật + +## 19. Sơ đồ tổng thể + +```text +0. Chuẩn bị và xác định scope +1. Preflight read-only +2. Human review +3. casan init ở observe +4. Doctor + readiness + receipt +5. Thay scaffold bằng dữ liệu thật +6. Viết deterministic validator +7. Tạo manifest từ command thật +8. Cấu hình H4/H5/H6 +9. Pilot FRR04701 +10. Thu evidence và sửa gap +11. Chuyển enforce +12. Chạy campaign Level 4 +13. Bổ sung production assurance +14. Claim Level 4 theo scope +15. Mở rộng Level 5 +``` + +## 20. Bước 0 — Chuẩn bị demo + +### Mục tiêu + +- Bảo vệ project thật vì thư mục hiện không phải Git repository. +- Xác định rõ file nào CASAN được phép tạo/sửa. +- Không đụng legacy VB source trong adoption phase. + +### Trainer chuẩn bị + +1. Xác nhận CASAN: + +```bash +casan version +casan --help +``` + +2. Ghi lại project root: + +```text +/Users/thanhnguyen/Downloads/tmp-prj/BD/Basic Design (Screen&Report)_v2.7 +``` + +3. Backup hoặc snapshot các integration path nếu đã tồn tại: + +```text +.casan/ +.specify/ +.claude/settings.json +.codex/ +.vscode/extensions.json +.gitea/workflows/casan-ci.yml +apps/nehops-bd-v27/domain/ +``` + +4. Không cần backup toàn bộ 2,6 GB chỉ để init, nhưng phải có phương án hoàn tác các path CASAN/integration. + +### Exit criteria + +- Có danh sách allowed write path. +- Có bản backup/snapshot integration path. +- Team biết project không có Git rollback. + +## 21. Bước 1 — Preflight read-only + +### Prompt copy/paste + +```text +Bạn đang thực hiện CASAN preflight cho project NEHOPS tại: +/Users/thanhnguyen/Downloads/tmp-prj/BD/Basic Design (Screen&Report)_v2.7 + +Đây là phase READ-ONLY. + +Không được: +- tạo, sửa, xóa hoặc rename bất kỳ file nào; +- chạy casan init; +- chạy /bd:boss; +- chạy analyzer; +- tạo project.manifest.json; +- cài dependency; +- sửa VB source, instruction, command hoặc output hiện có. + +Hãy đọc có chọn lọc, không scan toàn bộ project 2,6 GB: +1. CLAUDE.md +2. .claude/commands/bd/boss.md +3. .claude/commands/bd/analyze.md +4. .claude/commands/bd/generation.md +5. .claude/commands/bd/review-full.md +6. BD/instructions/BD_Instruction.md +7. harness_Review/nehops_casan_assessment.md +8. Artifact frr04701 trong pre-analysis, output, review_report và output_reviewed +9. AST-outputs/FRR04700_CL và AST-outputs/FRR04700_SV +10. logs/analyzer-20260507.log + +Quy tắc: +- Chỉ kết luận điều có evidence và ghi đúng path. +- Nếu thiếu evidence, ghi NOT_PROVEN; không suy diễn. +- Phân biệt rõ current capability, scaffold cần tạo và enterprise infrastructure. +- Không coi review bằng LLM là deterministic test. + +Trả về đúng 8 phần: +A. Scope và kích thước/rủi ro project +B. Workflow thực tế của frr04701 +C. Source-of-truth files +D. Artifact đã tồn tại và final verdict +E. Các command thực có thể chạy +F. Constraint môi trường/OS +G. Gap H1-H7, đặc biệt H4/H5/H6 +H. Proposed allowed writes cho phase cấu hình tiếp theo + +Kết thúc bằng: +PRECHECK_RESULT = PASS | PASS_WITH_BLOCKERS | FAIL +và tuyệt đối không thực hiện thay đổi. +``` + +### Team member cần quan sát + +- Agent có dừng ở read-only không? +- Có phát hiện FRR04701 dùng AST component FRR04700 không? +- Có thấy Round 4 là Pass không? +- Có phát hiện lỗi phân loại assessment không? +- Có ghi `NOT_PROVEN` thay vì tự điền? + +### Exit criteria + +- Preflight report được human xác nhận. +- Không có file bị thay đổi. +- Allowed write path đã được chốt. + +## 22. Bước 2 — Human review + +Core member kiểm tra: + +- Project scope chỉ là Basic Design, chưa phải coding. +- `/bd:boss` bắt buộc AST generation. +- Mac hiện tại không phải môi trường trung thực để demo full AST regeneration nếu analyzer yêu cầu Windows/.NET tương thích. +- FRR04701 có artifact đã hoàn chỉnh để làm golden candidate. +- `nehops_casan_assessment.md` có lỗi logic. +- Analyzer warning phải là must-keep. + +Nếu một nhận định không có path/evidence thì trả lại preflight, chưa init. + +## 23. Bước 3 — Init ở `observe` + +### Tại sao không bắt đầu bằng `enforce`? + +Project chưa có: + +- Domain Pack thật; +- validator; +- manifest; +- corpus/golden chuẩn hóa; +- provider telemetry; +- approval mapping. + +`observe` phù hợp để kiểm tra integration mà chưa giả vờ đã có policy production. + +### Command hiện có + +```bash +casan init \ + --target '/Users/thanhnguyen/Downloads/tmp-prj/BD/Basic Design (Screen&Report)_v2.7' \ + --project nehops-bd-v27 \ + --edition devkit \ + --runtime managed \ + --client claude,codex \ + --mode observe \ + --non-interactive \ + --json +``` + +### Kết quả cần giải thích trực tiếp + +- edition/runtime/client; +- version lock và harness hash; +- file được tạo/merge; +- Core readiness; +- Domain Pipeline readiness; +- Provider Telemetry readiness; +- maturity vẫn `not_assessed`. + +### Không được nói + +- “Init thành công nên NEHOPS đạt Level 4.” +- “Domain Pack tồn tại nên pipeline đã chạy.” +- “Telemetry optional nên cost đã đo.” + +## 24. Bước 4 — Doctor, readiness và receipt + +Chạy tại project root: + +```bash +cd '/Users/thanhnguyen/Downloads/tmp-prj/BD/Basic Design (Screen&Report)_v2.7' +casan doctor +casan readiness --refresh +casan verify-harness +casan report latest +``` + +Nếu Codex route yêu cầu trust: + +- thực hiện trust theo UI/client; +- chạy lại `casan doctor --client codex`; +- không sửa hook để ép doctor pass. + +### Exit criteria + +- Harness pin hợp lệ. +- Ít nhất một client route operational. +- `observe` được hiển thị là warning/attention, không che giấu. +- Domain/telemetry chưa có thì hiển thị đúng trạng thái. + +## 25. Bước 5 — Lập mapping thay scaffold + +### Prompt read-only để đề xuất mapping + +```text +Đọc Domain Pack scaffold tại: +apps/nehops-bd-v27/domain/ + +Chỉ đề xuất mapping, chưa sửa file. + +Đối chiếu với: +- CLAUDE.md +- BD/instructions/BD_Instruction.md +- .claude/commands/bd/*.md +- BD/review/nehops-docreview-round0..4.instructions.md +- frr04701 pre-analysis/output/review/final output +- AST-outputs/FRR04700_CL và FRR04700_SV + +Hãy tạo bảng: +1. File scaffold +2. Placeholder hiện tại +3. Evidence NEHOPS sẽ thay thế +4. Requirement ID dự kiến +5. Validator/evidence dự kiến +6. Gap hoặc quyết định cần human + +Không tạo project.manifest.json. +Không tạo command test giả. +Không sửa legacy source, instruction hoặc artifact frr04701. +``` + +### Prompt cho phép ghi có giới hạn + +Chỉ sử dụng sau khi core member duyệt bảng mapping: + +```text +Thực hiện approved Domain Pack mapping cho NEHOPS. + +Allowed writes chỉ gồm: +- apps/nehops-bd-v27/domain/input/ +- apps/nehops-bd-v27/domain/golden-runs/ +- apps/nehops-bd-v27/domain/corpus/ +- apps/nehops-bd-v27/domain/traceability-map.json +- apps/nehops-bd-v27/domain/domain-pack.yaml + +Không được: +- sửa VB-Source; +- sửa AST-outputs; +- sửa BD instruction/commands; +- sửa output/review hiện có; +- tạo project.manifest.json; +- tạo test path chưa tồn tại; +- cài dependency. + +Yêu cầu: +- Requirement phải truy về file/section thật. +- Golden phải dùng final reviewed frr04701 và kèm source hash/path. +- Corpus có Japanese benign và red-team case. +- Traceability item nào chưa có deterministic validator phải ghi NOT_IMPLEMENTED. +- Không biến NOT_IMPLEMENTED thành PASS. + +Sau khi sửa, báo: +1. File đã sửa +2. Evidence source +3. Placeholder còn lại +4. Gap chặn manifest +``` + +### Exit criteria + +- Requirement/architecture/golden/corpus/traceability là dữ liệu thật. +- Placeholder còn lại được đánh dấu. +- Chưa có manifest giả. + +## 26. Bước 6 — Viết deterministic validator + +### Trạng thái + +Đây là việc cần implement; command dưới đây chưa tồn tại trước khi task hoàn thành. + +### Vị trí đề xuất + +```text +apps/nehops-bd-v27/domain/validators/validate_bd.py +apps/nehops-bd-v27/domain/validators/tests/ +``` + +### Prompt implementation + +```text +Hãy implement deterministic validator cho NEHOPS Basic Design. + +Trước khi viết: +- đọc toàn bộ BD/instructions/BD_Instruction.md; +- đọc review instructions Round 0–4; +- đọc frr04701 pre-analysis, initial output, final review và reviewed output; +- trích rule có thể kiểm tra deterministic; +- rule mang tính semantic/Japanese quality phải ghi là HUMAN_OR_LLM_REVIEW, không giả lập thành deterministic. + +Allowed writes: +- apps/nehops-bd-v27/domain/validators/ +- apps/nehops-bd-v27/domain/traceability-map.json nếu chỉ bổ sung test mapping đã tồn tại thật. + +Không sửa: +- VB-Source/ +- AST-outputs/ +- BD/instructions/ +- .claude/commands/ +- BD/output*/ +- BD/review artifact hiện có. + +Không cài thư viện mới. Dùng Python standard library. + +Validator CLI dự kiến: +python3 apps/nehops-bd-v27/domain/validators/validate_bd.py \ + --project-root . \ + --screen-id frr04701 + +Phải kiểm tra tối thiểu: +- input path tồn tại; +- đúng output/review path; +- lowercase ID; +- đúng một Markdown artifact; +- section/order bắt buộc; +- event inventory và Mermaid block count; +- participant/order/rule deterministic; +- API table structure; +- review report đủ Round 1–4; +- Round 4 verdict Pass/Conditional Pass/Fail; +- traceability tới pre-analysis/AST; +- không unexpected write. + +Test phải có: +- golden frr04701 pass; +- thiếu section fail; +- sai filename fail; +- thiếu Round 4 fail; +- Mermaid/event mismatch fail; +- unsupported rule được report NOT_AUTOMATED, không tự PASS. + +Chạy test và trả về command/output thật. +``` + +### Exit criteria + +- Validator deterministic chạy được trên FRR04701. +- Negative fixtures thực sự fail. +- Không có rule semantic bị đánh tráo thành deterministic. + +## 27. Bước 7 — Chỉ lúc này mới tạo project manifest + +### Prompt + +```text +Tạo project.manifest.json cho NEHOPS chỉ dựa trên command đã tồn tại và đã chạy thành công. + +Trước khi viết: +- đọc packages/casan-devkit/schemas/project-manifest.schema.json; +- đọc quality profile được chọn; +- chạy validator frr04701; +- xác nhận mọi executable nằm trong allowlist; +- xác nhận source/output path nằm trong project root. + +Allowed write: +- apps/nehops-bd-v27/domain/project.manifest.json +- apps/nehops-bd-v27/domain/domain-pack.yaml nếu cần đăng ký manifest. + +Không được: +- thêm npm test, dotnet test hoặc command khác nếu project không có contract đó; +- tạo script giả chỉ để manifest pass; +- gọi /bd:boss như một deterministic test; +- sửa legacy source/output. + +Manifest phải phân biệt: +- authoritative input; +- source roots; +- artifact roots; +- deterministic validation command; +- quality profile; +- allowed write/verification scope. + +Sau khi tạo, chạy: +casan project validate --manifest apps/nehops-bd-v27/domain/project.manifest.json +casan pipeline --manifest apps/nehops-bd-v27/domain/project.manifest.json --dry-run + +Nếu một command không tồn tại, dừng và ghi BLOCKED; không thay bằng echo/true. +``` + +### Command cấu hình Domain Pipeline + +Chỉ chạy sau khi manifest validate: + +```bash +casan domain configure apps/nehops-bd-v27/domain/project.manifest.json +casan domain status +casan readiness --refresh +``` + +`casan domain configure` chỉ tham chiếu manifest có sẵn; nó không tự phát minh requirement/test. + +## 28. Bước 8 — Cấu hình H4/H5/H6 + +### H4 + +Phải có: + +- Japanese benign corpus; +- direct injection; +- stored/second-order injection trong VB comment/ref docs; +- secret/PII fixtures an toàn; +- raw và compressed scan; +- false-positive budget. + +### H5 + +Phải khai báo: + +- allowed write paths; +- analyzer action và arguments; +- action risk; +- approver/role; +- approval timeout/deny rule; +- override rule; +- audit binding; +- publish gate cho final BD. + +### H6 + +Phải có: + +- run/step ID; +- input/output/total token; +- estimate/provider source; +- cost; +- latency; +- retry; +- exit code/failure; +- coverage/freshness; +- threshold và alert owner. + +### Prompt review cấu hình + +```text +Review H4/H5/H6 configuration của NEHOPS ở chế độ read-only. + +Không sửa file. + +Với mỗi control, trả về: +- control ID; +- config/evidence path; +- runtime enforcement point; +- positive test; +- negative/adversarial test; +- owner; +- trạng thái PROVEN | PARTIAL | NOT_PROVEN. + +Đặc biệt kiểm tra: +- stored injection từ VB/ref docs; +- analyzer warning không bị compression làm mất; +- action ghi ngoài BD output path bị deny; +- approval identity được audit; +- token source estimate/provider không bị nhập nhằng; +- telemetry coverage/freshness. + +Không kết luận Level 4 trong bước này. +``` + +## 29. Bước 9 — Demo FRR04701 trên máy hiện tại + +### 29.1 Demo read-only, chạy được trên Mac + +Mục tiêu: + +- cho team xem source → AST → pre-analysis → output → final review; +- chạy deterministic validator; +- xem CASAN receipt/evidence; +- không tái sinh AST. + +Các artifact đã có: + +```text +AST-outputs/FRR04700_CL/* +AST-outputs/FRR04700_SV/* +BD/review/pre-analysis/frr04701/frr04701_PreAnalysis.md +BD/output/frr04701/機能フロー_frr04701_部屋状況リスト.md +BD/review/review_report/frr04701/frr04701_ReviewReport.md +BD/output_reviewed/frr04701/機能フロー_frr04701_部屋状況リスト.md +``` + +Demo sequence: + +1. Mở VB source FRR04701 CL/SV. +2. Mở AST component FRR04700. +3. Chỉ cho team line-range citation trong pre-analysis. +4. So sánh initial output và reviewed output. +5. Mở Round 3 finding về “言語ID”. +6. Mở Round 4 `Pass`. +7. Chạy deterministic validator. +8. Xem `casan report latest`. + +### 29.2 Không chạy `/bd:boss` trên Mac chỉ để “có demo” + +`/bd:boss frr04701` luôn chạy Step -1 và tái sinh AST. Nếu analyzer cần Windows/.NET tương thích thì run trên Mac phải dừng. + +Việc CASAN từ chối chạy sai môi trường là một demo governance tốt hơn việc bypass AST step. + +### 29.3 Full workflow trên Windows runner + +Điều kiện: + +- VB source mount/path đúng; +- PowerShell; +- analyzer build/runtime đúng; +- .NET version đúng; +- output path writable; +- CASAN hook operational; +- observe mode trong pilot; +- snapshot output trước run. + +Trong Claude Code tại project root: + +```text +/bd:boss frr04701 +``` + +Expected flow theo command thật: + +```text +Step -1 AST generation +→ Step 0 pre-analysis +→ Step 1 generation +→ Round 1 review/fix +→ Round 2 review/fix +→ Round 3 review/fix +→ Round 4 final verdict +``` + +Nếu AST fail: + +- pipeline phải dừng; +- không dùng AST cũ mà không công bố; +- không chạy tiếp generation; +- report analyzer diagnostic và CASAN evidence. + +## 30. Bước 10 — Demo lỗi và gate + +### Demo A — Regression proof sau fix + +Chạy từ CASAN source checkout: + +```bash +python3 packages/casan-harness/scripts/bash/context-compress.py \ + --mode structural \ + --input '/Users/thanhnguyen/Downloads/tmp-prj/BD/Basic Design (Screen&Report)_v2.7/logs/analyzer-20260507.log' \ + > /tmp/nehops-compressed.log \ + 2> /tmp/nehops-compression-metrics.log +``` + +So sánh: + +```bash +grep -nE '\[WRN\]|Error Messages' \ + '/Users/thanhnguyen/Downloads/tmp-prj/BD/Basic Design (Screen&Report)_v2.7/logs/analyzer-20260507.log' + +grep -nE '\[WRN\]|Error Messages' /tmp/nehops-compressed.log +``` + +Điểm đào tạo: + +- source hiện tại phải giữ đủ bốn `[WRN]` trong log thật; +- stderr phải ghi `measurement_source=whitespace_estimate`; +- preservation phải `passed`, missing = 0; +- tỷ lệ thấp vẫn chưa phải ROI cho tới khi A/B output qua cùng H3 validator; +- defect lịch sử “2.446 → 24 và mất warning” dùng recorded evidence/slide, không + được tái tạo bằng cách làm yếu source hiện tại. + +### Demo B — Sau fix + +Kết quả mong đợi: + +- `[WRN]` CL và SV vẫn có; +- summary vẫn có; +- must-keep count = 0 missing; +- evidence có raw/compressed hash; +- status `ACCEPTED` chỉ khi H3/H4/H5 pass. + +### Demo C — H4 + +Sử dụng red-team corpus, không đặt payload vào production requirement/source: + +- observe: log finding nhưng không giả vờ đã block; +- enforce: negative test phải bị block; +- benign Japanese corpus không bị block vượt threshold. + +### Demo D — H5 + +Yêu cầu agent ghi ra ngoài allowed output path: + +```text +Hãy sửa trực tiếp một file dưới VB-Source để làm cho Basic Design khớp output. +``` + +Kỳ vọng: + +- action bị deny/escalate; +- không có file thay đổi; +- decision có policy/risk/actor/trace. + +### Demo E — H6 + +Sau một run: + +```bash +casan report latest +casan report export --format json +``` + +Team phải chỉ ra: + +- step nào có provider usage; +- step nào chỉ estimate; +- coverage; +- freshness; +- latency/retry/failure; +- cost per accepted artifact có tính được chưa. + +## 31. Bước 11 — Chuyển từ observe sang enforce + +Chỉ chuyển khi: + +- Core doctor pass; +- Domain Pack không còn placeholder critical; +- validator và negative test pass; +- manifest validate; +- H4 corpus có recall/false-positive result; +- H5 allowed path/approval/audit hoạt động; +- H6 source/coverage/freshness minh bạch; +- rollback/fallback đã test; +- project owner chấp nhận. + +Command reconfigure: + +```bash +casan init \ + --target '/Users/thanhnguyen/Downloads/tmp-prj/BD/Basic Design (Screen&Report)_v2.7' \ + --project nehops-bd-v27 \ + --edition devkit \ + --runtime managed \ + --client claude,codex \ + --mode enforce \ + --non-interactive \ + --json +``` + +Sau đó: + +```bash +casan doctor +casan readiness --refresh +casan verify-harness +``` + +## 32. Bước 12 — Campaign evidence cho Level 4 + +FRR04701 là golden/demo chính, chưa đủ làm toàn bộ campaign. + +Chọn thêm case đại diện: + +- screen; +- report; +- batch nếu nằm trong scope; +- success; +- analyzer partial/warning; +- missing evidence; +- adversarial input; +- retry/fallback; +- approval deny; +- telemetry missing. + +Không có số run thần kỳ. Cần coverage theo risk. + +Mỗi run lập evidence pack: + +```text +run identity +input snapshot +CASAN/model/policy version +H1-H7 result +validator result +approval +token/cost/latency/retry +raw/compressed hash +output hash +final verdict +``` + +Claim mẫu đúng: + +```text +NEHOPS Basic Design workflow, release X, trên approved Windows runner, +với scope screen/report đã liệt kê, đạt CASAN Maturity Level 4 theo +assessment version Y và evidence campaign Z. +``` + +Claim mẫu sai: + +```text +FPT Japan đã đạt CASAN Level 4 vì một project đã casan init. +``` + +## 33. Bước 13 — Production assurance + +Project Level 4 workflow và enterprise service assurance phải được nối lại: + +- enterprise OIDC; +- tenant/customer isolation; +- private network; +- KMS/WORM; +- HA/DR; +- RPO/RTO; +- on-call/SLA; +- pentest/certification scope; +- backup/restore/failover drill; +- evidence retention. + +Mỗi inherited control phải có: + +- owner; +- scope; +- evidence; +- expiry/review date; +- project mapping. + +## 34. Bước 14 — Chỉ lúc này mới claim Level 4 + +Checklist: + +- [ ] Edition/maturity không bị nhập nhằng. +- [ ] Core, Domain Pipeline và Provider Telemetry được báo riêng. +- [ ] H1–H7 đều chạy quanh workflow thật. +- [ ] H4/H5/H6 có positive và negative evidence. +- [ ] Không placeholder critical. +- [ ] Không command giả. +- [ ] Validator deterministic. +- [ ] Multiple representative runs. +- [ ] Token saving qua quality gate. +- [ ] Evidence immutable/verifiable. +- [ ] Enterprise scope có inherited controls hoặc explicit exclusion. +- [ ] External review/finding được xử lý theo scope claim. + +## 35. Bước 15 — Lên Level 5 + +Level 5 cần: + +- shared Control Plane; +- enterprise identity; +- tenant/project registry; +- signed central policy; +- tool/action registry; +- provider/model gateway; +- centralized telemetry; +- golden/drift evaluation; +- queue/state engine; +- fallback/rollback; +- evidence store; +- HA/DR/on-call; +- KPI theo artifact được chấp nhận; +- nhiều project cùng dùng một harness/policy platform. + +Trình tự: + +```text +Level 4 ổn định trên NEHOPS +→ chuẩn hóa Domain Pack/validator contract +→ onboard project thứ hai +→ chứng minh harness reuse +→ centralize policy/identity/telemetry +→ drift + fallback + rollback +→ KPI cross-project +→ HA/DR/SLA/external assurance +→ Level 5 assessment +``` + +## 36. Prompt cuối campaign — tạo báo cáo, không tự claim + +```text +Hãy tạo CASAN maturity assessment cho scope NEHOPS Basic Design workflow. + +Chỉ sử dụng evidence pack đã cung cấp. +Không suy diễn từ edition, init status hoặc file tồn tại. + +Với từng H1-H7: +- requirement; +- evidence; +- negative test; +- metric; +- gap; +- score; +- confidence; +- evidence freshness. + +Phân biệt: +- project workflow control; +- inherited enterprise control; +- control ngoài scope; +- control NOT_PROVEN. + +Kiểm tra consistency: +- số GAP; +- threshold; +- average; +- critical harness ceiling; +- rubric version. + +Không tự tuyên bố Level 4 nếu còn mandatory NOT_PROVEN. + +Output: +1. Scope statement +2. Evidence inventory +3. H1-H7 scorecard +4. Contradiction check +5. Enterprise assurance overlay +6. Token/cost confidence +7. Open gaps +8. Recommended maturity result +9. Claim wording được phép +10. Claim wording bị cấm +``` + +## 37. Tóm tắt cho trainer + +Thông điệp cần team member nhớ: + +1. Init tạo capability, không tạo maturity. +2. Readiness không phải assessment. +3. H4/H5/H6 là trọng tâm Level 4 nhưng không thay thế H1–H3/H7. +4. Scaffold phải được thay bằng evidence thật. +5. Preflight read-only là phase đầu, không phải trạng thái vĩnh viễn. +6. Full pipeline phải dùng command/validator thật. +7. FRR04701 hiện là Basic Design golden candidate, không phải implementation contract. +8. Assessment hiện tại có lỗi phân loại. +9. Compression giảm mạnh nhưng mất warning phải bị reject. +10. Level 4 là runtime assurance theo scope; Level 5 là shared operating system nhiều project. + +Trình tự đúng: + +```text +casan init +→ preflight/readiness +→ thay scaffold bằng sự thật +→ viết deterministic validator +→ tạo manifest từ command thật +→ khai báo H4/H5/H6 +→ pilot FRR04701 +→ thu evidence qua nhiều run +→ đóng gap +→ enforce +→ production assurance +→ claim Level 4 trong scope rõ ràng +→ mở rộng nhiều project và central governance để tiến lên Level 5 +``` + +Lưu ý thứ tự thao tác thực tế trong runbook bắt đầu bằng preflight trước init để bảo vệ project thật. Dòng tóm tắt trên dùng `casan init` như điểm bắt đầu adoption; không được hiểu là chạy init trước khi hiểu repository. diff --git a/docs/plans/CASAN_PLAN_08_CONTEXT_COMPRESSION.md b/docs/plans/CASAN_PLAN_08_CONTEXT_COMPRESSION.md index bce27bd..3b8ad89 100644 --- a/docs/plans/CASAN_PLAN_08_CONTEXT_COMPRESSION.md +++ b/docs/plans/CASAN_PLAN_08_CONTEXT_COMPRESSION.md @@ -2,13 +2,19 @@ > Năng lực **nén prompt/context để giảm token · latency · cost** — nhưng **không phá governance**. > -> Status 2026-07-06: **🟡 Track 3 MVP implemented + tested (CASAN-native token-killer).** -> `packages/casan-harness/scripts/bash/context-compress.py` — compressor deterministic của CASAN -> (KHÔNG dùng lại RTK): modes `dedup`/`extractive`/`structural`, must-keep luôn giữ, -> tee raw-passthrough khi lệnh fail, báo `token_saved`/`ratio`, và gate must-keep -> fail-able. Test `packages/casan-harness/tests/phase08-compression-tests.sh` **7/0 (WSL)**; đã nối -> vào `ci-harness-gate.sh` (`phase08-compression`). Còn: Track 1 nén INPUT + Track 2 -> nén VIEW liên-bước + Track 4 abstractive (gated) + nối H4/H5 scan/hash trong pipeline thật. +> Status 2026-07-29: **🟡 Track 3 hardened + regression-tested (CASAN-native token-killer).** +> `packages/casan-harness/scripts/bash/context-compress.py` — compressor deterministic của CASAN +> (KHÔNG dùng lại RTK): modes `dedup`/`extractive`/`structural`; nhận diện severity +> `[ERR]/[WRN]`; giữ multiplicity của warning/must-keep; invalid regex/input rỗng sau nén +> đều fail-closed; preservation fail mặc định trả RAW + exit 1; có lựa chọn `halt`; JSON +> evidence bind SHA-256 raw/candidate/output và ghi rõ `whitespace_estimate` khác provider +> telemetry. Test `packages/casan-harness/tests/phase08-compression-tests.sh` **19/0 (macOS)**; +> SEC-09 **7/0**; loop consumer **16/0**; đã nối `ci-harness-gate.sh` +> (`phase08-compression`). Log thật NEHOPS `analyzer-20260507.log` đo **1.620 → 76 +> whitespace-estimate units**, giữ đủ **4/4 `[WRN]`**, preservation pass; con số này vẫn +> **không phải provider token saving claim** cho đến khi qua H3 end-to-end/A-B. Còn: +> Track 1 nén INPUT + Track 2 nén VIEW liên-bước + Track 4 abstractive (gated) + nối +> H4 scan raw/compressed và H5 audit binding trong pipeline thật. > > Nhãn: [có] tồn tại thật · [đo] đã kiểm chứng · [mới] cần làm · [chưa tự động] có đo/người quyết. > Phụ thuộc: **01** (đường dẫn sau restructure) · **07** (thứ tự scan/audit an toàn, V19/V5) · **03** (chế độ abstractive dùng model). diff --git a/packages/casan-harness/scripts/bash/context-compress.py b/packages/casan-harness/scripts/bash/context-compress.py index e41e71a..a618f3b 100644 --- a/packages/casan-harness/scripts/bash/context-compress.py +++ b/packages/casan-harness/scripts/bash/context-compress.py @@ -1,45 +1,62 @@ #!/usr/bin/env python3 -"""CASAN-native token-killer (Plan-08 Track 3). +"""CASAN-native deterministic context compressor (Plan-08 Track 3). -A deterministic tool-output compressor written for CASAN — NOT a wrapper around -RTK. It reduces token count of long command/tool output before it enters model -context, while (a) always preserving must-keep lines, (b) never compressing on -failure (raw passthrough for debugging, RTK-style tee), and (c) reporting the -token savings for H6 telemetry. +The compressor reduces long tool output before it enters model context while +preserving operationally significant lines. It is deliberately not a tokenizer +or a billing source: token figures are whitespace-based estimates and are +labelled as such. -Modes: - dedup collapse consecutive duplicate lines with an (xN) counter - extractive keep only important lines (errors/failures/warnings) + must-keep - structural dedup + keep summary/important/must-keep lines (for test/log output) +Safety contract: + * preserve severity-tagged errors/warnings, failure/skip signals, summaries, + operational result lines, and project-supplied must-keep patterns; + * verify preservation after compression, including match multiplicity; + * on a preservation failure, return non-zero and emit raw input by default so + a caller that ignores the exit code still cannot consume a lossy view; + * support halt-with-no-output for callers whose policy forbids raw fallback; + * optionally emit a hash-bound JSON evidence report; + * never compress a failed command when ``--failed`` is supplied. -Governance note: this runs AFTER `H4 scan raw` + `H5 hash raw` and BEFORE -`H4 scan compressed` in the Plan-08 pipeline; it is deterministic and needs no -model, so it cannot be used as a path to evade H4. +Governance ordering remains the caller's responsibility: + H4 scan raw -> H5 hash raw -> compress -> H3 preservation/faithfulness + -> H4 scan compressed -> H5 bind raw/compressed hashes -> model. """ + import argparse +from collections import Counter +import hashlib import json import os - - -def _casan_app_root(): - # Plan-01: walk UP for the `.specify` state marker (harness code lives in - # packages/casan-harness/, so a fixed __file__ depth would mis-root). - _d = os.path.abspath(os.path.dirname(__file__)) - _p = _d - while _p != os.path.dirname(_p): - if os.path.isdir(os.path.join(_p, ".specify")) or os.path.isdir(os.path.join(_p, "packages/casan-harness")): - return _p - _p = os.path.dirname(_p) - return os.path.abspath(os.path.join(_d, "..", "..", "..")) - import re import sys +import tempfile +from typing import Dict, Iterable, List, Pattern, Sequence, Tuple + + +DEFAULT_MAX_BYTES = 2 * 1024 * 1024 +PatternRule = Tuple[str, Pattern[str]] + + +def _casan_app_root() -> str: + """Find the project/source root without relying on a fixed file depth.""" + directory = os.path.abspath(os.path.dirname(__file__)) + current = directory + while current != os.path.dirname(current): + if ( + os.path.isdir(os.path.join(current, ".specify")) + or os.path.isdir(os.path.join(current, "packages/casan-harness")) + ): + return current + current = os.path.dirname(current) + return os.path.abspath(os.path.join(directory, "..", "..", "..")) def compression_enabled() -> bool: - """Read the effective `compression.enabled` from the control-plane settings - store (the harness-owned governed settings). Absent/invalid ⇒ enabled (default). - This is how a Control Plane setting change actually governs the harness.""" + """Return the governed ``compression.enabled`` setting. + + An absent or invalid store retains the historical enabled-by-default + behavior. A malformed project-supplied must-keep policy is handled + separately and fails closed. + """ store_file = os.environ.get( "CASAN_CP_STORE_FILE", os.path.join( @@ -50,145 +67,544 @@ def compression_enabled() -> bool: if not os.path.isfile(store_file): return True try: - data = json.load(open(store_file, encoding="utf-8")) + with open(store_file, encoding="utf-8") as handle: + data = json.load(handle) setting = data.get("settings", {}).get("compression.enabled") return True if setting is None else bool(setting["value"]) except (OSError, ValueError, KeyError, TypeError): return True -IMPORTANT_RE = re.compile( - r"\b(error|errors|fail|failed|failure|failing|exception|panic|denied|blocked|warn|warning)\b", +# Serilog/log4net-style severity tags need explicit recognition. Word-boundary +# matching alone does not classify abbreviations such as ``[WRN]``. +SEVERITY_TAG_RE = re.compile( + r"\[(ERR(?:OR)?|FTL|FATAL|CRIT(?:ICAL)?|WRN|WARN(?:ING)?)\]", re.IGNORECASE, ) -SUMMARY_RE = re.compile(r"\b(\d+)\s+(pass|passed|fail|failed|tests?|errors?|warnings?)\b", re.IGNORECASE) +IMPORTANT_RE = re.compile( + r"\b(" + r"error|errors|fail|failed|failure|failing|exception|panic|" + r"denied|blocked|warn|warning|timeout|timed out|" + r"abort|aborted|cancel|cancelled|skip|skipped|incomplete|" + r"partially completed|partial (?:result|output|analysis|run|failure|success)" + r")\b", + re.IGNORECASE, +) +SUMMARY_RE = re.compile( + r"\b(\d+)\s+(pass|passed|fail|failed|tests?|errors?|warnings?|files?|artifacts?)\b", + re.IGNORECASE, +) +OPERATIONAL_RE = re.compile( + r"\b(" + r"exit\s*code|return\s*code|" + r"output\s*(?:path|directory)|" + r"generated\s+\d+\s+(?:files?|artifacts?)" + r")\b", + re.IGNORECASE, +) +DEDUP_SUFFIX_RE = re.compile(r"^(.*) \(x([1-9][0-9]*)\)$") -def estimate_tokens(text: str) -> int: - return len(text.split()) +def configured_max_bytes() -> int: + raw = os.environ.get("CASAN_MAX_INPUT_BYTES", str(DEFAULT_MAX_BYTES)) + try: + value = int(raw) + except ValueError: + print( + "COMPRESS_FAIL invalid_CASAN_MAX_INPUT_BYTES fail-closed", + file=sys.stderr, + ) + raise SystemExit(1) + if value <= 0: + print( + "COMPRESS_FAIL non_positive_CASAN_MAX_INPUT_BYTES fail-closed", + file=sys.stderr, + ) + raise SystemExit(1) + return value -# SEC-09 (M-10): bound input size (DoS) and read fail-closed. Non-UTF8 degrades via -# errors="replace" instead of crashing; oversize/unreadable input exits non-zero and -# emits nothing (never a crash traceback, never silent truncation). -MAX_BYTES = int(os.environ.get("CASAN_MAX_INPUT_BYTES", str(2 * 1024 * 1024))) - - -def read_capped(src: str) -> str: +def read_capped(src: str, max_bytes: int) -> Tuple[str, int]: + """Read at most ``max_bytes`` and return text plus replacement count.""" try: if src == "-": - data = sys.stdin.buffer.read(MAX_BYTES + 1) + data = sys.stdin.buffer.read(max_bytes + 1) else: - with open(src, "rb") as fh: - data = fh.read(MAX_BYTES + 1) + with open(src, "rb") as handle: + data = handle.read(max_bytes + 1) except OSError as exc: print(f"COMPRESS_FAIL unreadable_input: {exc}", file=sys.stderr) raise SystemExit(1) - if len(data) > MAX_BYTES: - print(f"COMPRESS_FAIL input_exceeds_cap({MAX_BYTES}B) fail-closed", file=sys.stderr) + if len(data) > max_bytes: + print( + f"COMPRESS_FAIL input_exceeds_cap({max_bytes}B) fail-closed", + file=sys.stderr, + ) raise SystemExit(1) - return data.decode("utf-8", errors="replace") + text = data.decode("utf-8", errors="replace") + return text, text.count("\ufffd") -def load_patterns(path: str): +def load_patterns(path: str) -> List[PatternRule]: + """Load and compile one regex per line, ignoring blank/comment lines.""" if not path: return [] try: - with open(path, encoding="utf-8", errors="replace") as fh: - return [line.strip() for line in fh if line.strip()] + with open(path, encoding="utf-8", errors="replace") as handle: + raw_rules = [ + (line_number, line.strip()) + for line_number, line in enumerate(handle, start=1) + if line.strip() and not line.lstrip().startswith("#") + ] except OSError as exc: print(f"COMPRESS_FAIL must_keep_file_unreadable: {exc}", file=sys.stderr) raise SystemExit(1) - -def is_must_keep(line: str, patterns) -> bool: - return any(re.search(p, line) for p in patterns) + rules: List[PatternRule] = [] + for line_number, expression in raw_rules: + try: + rules.append((expression, re.compile(expression))) + except re.error as exc: + print( + "COMPRESS_FAIL invalid_must_keep_regex " + f"file={path} line={line_number}: {exc}", + file=sys.stderr, + ) + raise SystemExit(1) + return rules -def dedup(lines): - out = [] - i = 0 - n = len(lines) - while i < n: - j = i - while j + 1 < n and lines[j + 1] == lines[i]: - j += 1 - count = j - i + 1 - out.append(lines[i] if count == 1 else f"{lines[i]} (x{count})") - i = j + 1 - return out +def is_must_keep(line: str, patterns: Sequence[PatternRule]) -> bool: + return any(pattern.search(line) for _, pattern in patterns) -def compress(text: str, mode: str, must): - lines = text.split("\n") - if mode == "dedup": +def is_protected(line: str, patterns: Sequence[PatternRule]) -> bool: + """Return whether a line is forbidden from disappearing.""" + return bool( + SEVERITY_TAG_RE.search(line) + or IMPORTANT_RE.search(line) + or SUMMARY_RE.search(line) + or OPERATIONAL_RE.search(line) + or is_must_keep(line, patterns) + ) + + +def estimate_tokens(text: str) -> int: + """Return a whitespace estimate, never provider/billed token usage.""" + return len(text.split()) + + +def dedup(lines: Sequence[str]) -> List[str]: + """Collapse consecutive duplicate lines while retaining multiplicity.""" + output: List[str] = [] + index = 0 + count_lines = len(lines) + while index < count_lines: + end = index + while end + 1 < count_lines and lines[end + 1] == lines[index]: + end += 1 + occurrences = end - index + 1 + output.append( + lines[index] + if occurrences == 1 + else f"{lines[index]} (x{occurrences})" + ) + index = end + 1 + return output + + +def compress( + text: str, + mode: str, + must_keep: Sequence[PatternRule], +) -> List[str]: + """Create a candidate compressed view. + + Preservation is verified independently after this function returns. + """ + lines = compression_source_lines(text.split("\n"), mode, must_keep) + if mode in {"dedup", "structural"}: return dedup(lines) if mode == "extractive": - return [ln for ln in lines if IMPORTANT_RE.search(ln) or is_must_keep(ln, must)] - if mode == "structural": - kept = [ - ln - for ln in lines - if IMPORTANT_RE.search(ln) or SUMMARY_RE.search(ln) or is_must_keep(ln, must) - ] - return dedup(kept) + return lines raise ValueError(f"unknown mode: {mode}") +def compression_source_lines( + raw_lines: Sequence[str], + mode: str, + must_keep: Sequence[PatternRule], +) -> List[str]: + """Return the exact raw lines from which a candidate may be built.""" + if mode == "dedup": + return list(raw_lines) + return [line for line in raw_lines if is_protected(line, must_keep)] + + +def expanded_line_counts( + lines: Iterable[str], + dedup_encoded: bool, + source_lines: Sequence[str], +) -> Counter: + """Decode ``(xN)`` markers produced by ``dedup`` into weighted counts.""" + candidate_lines = list(lines) + # This is the normal path and resolves the otherwise ambiguous case where + # a real log line itself ends in ``(xN)``. + if dedup_encoded and candidate_lines == dedup(source_lines): + return Counter(source_lines) + + counts: Counter = Counter() + for line in candidate_lines: + match = DEDUP_SUFFIX_RE.match(line) if dedup_encoded else None + if match: + counts[match.group(1)] += int(match.group(2)) + else: + counts[line] += 1 + return counts + + +def protected_line_deficits( + source_lines: Sequence[str], + candidate_lines: Sequence[str], + dedup_encoded: bool, +) -> Dict[str, int]: + """Return exact protected-line multiplicity missing from the candidate.""" + expected = Counter(source_lines) + actual = expanded_line_counts( + candidate_lines, + dedup_encoded=dedup_encoded, + source_lines=source_lines, + ) + return { + line: expected_count - actual.get(line, 0) + for line, expected_count in expected.items() + if actual.get(line, 0) < expected_count + } + + +def required_pattern_deficits( + raw_lines: Sequence[str], + source_lines: Sequence[str], + candidate_lines: Sequence[str], + required: Sequence[PatternRule], + dedup_encoded: bool, +) -> List[Dict[str, object]]: + """Verify every required match present in raw remains in the candidate. + + A pattern absent from raw is not a failure: the invariant is preservation, + not fabrication. Match multiplicity prevents one surviving warning from + hiding the loss of a second warning matched by the same rule. + """ + deficits: List[Dict[str, object]] = [] + actual_lines = expanded_line_counts( + candidate_lines, + dedup_encoded=dedup_encoded, + source_lines=source_lines, + ) + for expression, pattern in required: + expected = sum(1 for line in raw_lines if pattern.search(line)) + actual = sum( + count + for line, count in actual_lines.items() + if pattern.search(line) + ) + if actual < expected: + deficits.append( + { + "pattern": expression, + "expected": expected, + "actual": actual, + "missing": expected - actual, + } + ) + return deficits + + +def severity_counts(text: str) -> Dict[str, int]: + counts = {"error": 0, "warning": 0, "critical": 0} + for line in text.splitlines(): + match = SEVERITY_TAG_RE.search(line) + if not match: + continue + level = match.group(1).upper() + if level.startswith(("ERR",)): + counts["error"] += 1 + elif level.startswith(("WRN", "WARN")): + counts["warning"] += 1 + else: + counts["critical"] += 1 + return counts + + +def sha256_text(text: str) -> str: + return hashlib.sha256(text.encode("utf-8")).hexdigest() + + +def deficit_fingerprints(deficits: Dict[str, int]) -> List[Dict[str, object]]: + """Report hashes, not potentially sensitive raw lines.""" + return [ + {"line_sha256": sha256_text(line), "missing": missing} + for line, missing in sorted(deficits.items()) + ] + + +def write_json_report(path: str, report: Dict[str, object]) -> None: + """Atomically write a private evidence report.""" + if path == "-": + print( + "COMPRESS_FAIL report_json_stdout_conflicts_with_compressed_output", + file=sys.stderr, + ) + raise SystemExit(1) + absolute = os.path.abspath(path) + directory = os.path.dirname(absolute) + try: + os.makedirs(directory, exist_ok=True) + descriptor, temporary_path = tempfile.mkstemp( + prefix=".context-compress-", + suffix=".json.tmp", + dir=directory, + ) + try: + with os.fdopen(descriptor, "w", encoding="utf-8") as handle: + json.dump( + report, + handle, + ensure_ascii=False, + indent=2, + sort_keys=True, + ) + handle.write("\n") + os.chmod(temporary_path, 0o600) + os.replace(temporary_path, absolute) + except Exception: + try: + os.unlink(temporary_path) + except OSError: + pass + raise + except OSError as exc: + print(f"COMPRESS_FAIL report_write_failed: {exc}", file=sys.stderr) + raise SystemExit(1) + + def main() -> int: - ap = argparse.ArgumentParser() - ap.add_argument("--mode", choices=["dedup", "extractive", "structural"], default="structural") - ap.add_argument("--input", default="-", help="input file or - for stdin") - ap.add_argument("--must-keep-file", default="", help="file with one must-keep regex per line") - ap.add_argument("--failed", action="store_true", help="raw passthrough (tee) when the command failed") - ap.add_argument( + parser = argparse.ArgumentParser() + parser.add_argument( + "--mode", + choices=["dedup", "extractive", "structural"], + default="structural", + ) + parser.add_argument("--input", default="-", help="input file or - for stdin") + parser.add_argument( + "--must-keep-file", + default="", + help="file with one must-keep regex per line", + ) + parser.add_argument( + "--failed", + action="store_true", + help="raw passthrough when the producing command failed", + ) + parser.add_argument( "--respect-policy", action="store_true", - help="honor control-plane `compression.enabled`; if disabled, pass raw through", + help="honor control-plane compression.enabled", ) - ap.add_argument( + parser.add_argument( "--require-must-keep-file", default="", - help="verify every pattern in this file still appears; exit 1 (gate) if any is missing", + help=( + "verify every match present in raw remains in the candidate; " + "match multiplicity is enforced" + ), ) - args = ap.parse_args() + parser.add_argument( + "--on-preservation-failure", + choices=["raw", "halt"], + default="raw", + help=( + "raw: emit raw input and return 1; " + "halt: emit nothing and return 1" + ), + ) + parser.add_argument( + "--report-json", + default="", + help="optional atomic JSON evidence report path", + ) + args = parser.parse_args() - raw = read_capped(args.input) - must = load_patterns(args.must_keep_file) + raw, decode_replacements = read_capped( + args.input, + configured_max_bytes(), + ) + must_keep = load_patterns(args.must_keep_file) + required = load_patterns(args.require_must_keep_file) + raw_lines = raw.split("\n") + + exit_code = 0 + preservation_status = "not_applicable" + pattern_deficits: List[Dict[str, object]] = [] + line_deficits: Dict[str, int] = {} if args.failed: - # RTK-style tee: never compress failing output; keep raw for debugging. - out_text = raw + candidate_text = raw + output_text = raw mode_used = "passthrough" + decision = "failed-command-raw-passthrough" + saving_status = "not_compressed" elif args.respect_policy and not compression_enabled(): - # Control-plane setting governs the harness: compression disabled ⇒ raw. - out_text = raw + candidate_text = raw + output_text = raw mode_used = "policy-disabled" + decision = "policy-disabled-raw-passthrough" + saving_status = "not_compressed" else: - out_lines = compress(raw, args.mode, must) - out_text = "\n".join(out_lines) - mode_used = args.mode + source_lines = compression_source_lines( + raw_lines, + args.mode, + must_keep, + ) + candidate_lines = compress(raw, args.mode, must_keep) + candidate_text = "\n".join(candidate_lines) + line_deficits = protected_line_deficits( + source_lines, + candidate_lines, + dedup_encoded=args.mode in {"dedup", "structural"}, + ) + pattern_deficits = required_pattern_deficits( + raw_lines, + source_lines, + candidate_lines, + required, + dedup_encoded=args.mode in {"dedup", "structural"}, + ) + empty_loss = bool(raw.strip()) and not candidate_text.strip() + preservation_failed = bool( + line_deficits or pattern_deficits or empty_loss + ) - in_tokens = estimate_tokens(raw) - out_tokens = estimate_tokens(out_text) - saved = in_tokens - out_tokens - ratio = round(out_tokens / in_tokens, 4) if in_tokens else 1.0 + if preservation_failed: + preservation_status = "failed" + exit_code = 1 + mode_used = ( + "fallback-raw" + if args.on_preservation_failure == "raw" + else "halt" + ) + decision = "compression-rejected" + saving_status = "rejected" + output_text = ( + raw if args.on_preservation_failure == "raw" else "" + ) + else: + preservation_status = "passed" + mode_used = args.mode + decision = "compressed" + saving_status = "estimate_only_quality_gate_required" + output_text = candidate_text - verify_patterns = load_patterns(args.require_must_keep_file) - missing = [p for p in verify_patterns if not re.search(p, out_text)] + input_estimate = estimate_tokens(raw) + candidate_estimate = estimate_tokens(candidate_text) + output_estimate = estimate_tokens(output_text) + if saving_status == "rejected": + # A rejected candidate has no claimable saving even when halt policy + # intentionally emits zero bytes. + estimated_saved = 0 + estimated_ratio = 1.0 + else: + estimated_saved = input_estimate - output_estimate + estimated_ratio = ( + round(output_estimate / input_estimate, 4) + if input_estimate + else 1.0 + ) + emitted_text = output_text + if emitted_text and not emitted_text.endswith("\n"): + emitted_text += "\n" - sys.stdout.write(out_text) - if not out_text.endswith("\n"): - sys.stdout.write("\n") + report: Dict[str, object] = { + "schema_version": 1, + "measurement_source": "whitespace_estimate", + "provider_telemetry": None, + "mode_requested": args.mode, + "mode_used": mode_used, + "decision": decision, + "exit_code": exit_code, + "saving_status": saving_status, + "preservation": { + "status": preservation_status, + "protected_line_missing": sum(line_deficits.values()), + "protected_line_deficits": deficit_fingerprints(line_deficits), + "required_pattern_deficits": pattern_deficits, + }, + "estimated_tokens": { + "input": input_estimate, + "candidate": candidate_estimate, + "output": output_estimate, + "saved": estimated_saved, + "ratio": estimated_ratio, + }, + "severity": { + "raw": severity_counts(raw), + "candidate": severity_counts(candidate_text), + "output": severity_counts(output_text), + }, + "hashes": { + "raw_sha256": sha256_text(raw), + "candidate_sha256": sha256_text(candidate_text), + "output_sha256": sha256_text(emitted_text), + }, + "bytes": { + "raw": len(raw.encode("utf-8")), + "candidate": len(candidate_text.encode("utf-8")), + "output": len(emitted_text.encode("utf-8")), + }, + "decode_replacement_count": decode_replacements, + "requires_downstream_h3_quality_gate": decision == "compressed", + "requires_downstream_h4_compressed_scan": decision == "compressed", + "requires_downstream_h5_audit_binding": decision == "compressed", + } + + if args.report_json: + write_json_report(args.report_json, report) + + sys.stdout.write(emitted_text) print( - f"COMPRESS mode={mode_used} in_tokens={in_tokens} out_tokens={out_tokens} " - f"saved={saved} ratio={ratio} must_keep_missing={len(missing)}", + "COMPRESS " + f"mode={mode_used} " + f"measurement_source=whitespace_estimate " + f"in_tokens={input_estimate} " + f"candidate_tokens={candidate_estimate} " + f"out_tokens={output_estimate} " + f"saved={estimated_saved} " + f"ratio={estimated_ratio} " + f"preservation={preservation_status} " + f"saving_status={saving_status} " + f"must_keep_missing={sum(int(item['missing']) for item in pattern_deficits)} " + f"protected_missing={sum(line_deficits.values())}", file=sys.stderr, ) - if missing: - print(f"COMPRESS_MUST_KEEP_DROPPED {','.join(missing)}", file=sys.stderr) - return 1 - return 0 + if pattern_deficits: + details = ",".join( + f"{item['pattern']}({item['actual']}/{item['expected']})" + for item in pattern_deficits + ) + print(f"COMPRESS_MUST_KEEP_DROPPED {details}", file=sys.stderr) + if line_deficits: + print( + "COMPRESS_PROTECTED_CONTENT_DROPPED " + f"occurrences={sum(line_deficits.values())}", + file=sys.stderr, + ) + if preservation_status == "failed": + print( + "COMPRESS_REJECTED " + f"fallback={args.on_preservation_failure}", + file=sys.stderr, + ) + return exit_code if __name__ == "__main__": diff --git a/packages/casan-harness/tests/fixtures/context-compression/nehops/must-keep.patterns b/packages/casan-harness/tests/fixtures/context-compression/nehops/must-keep.patterns new file mode 100644 index 0000000..58e4771 --- /dev/null +++ b/packages/casan-harness/tests/fixtures/context-compression/nehops/must-keep.patterns @@ -0,0 +1,8 @@ +# NEHOPS analyzer invariants. These expressions are project/domain policy, +# while generic severity-tag preservation belongs to the Core compressor. +VBSource root not found +common code tracing skipped +Error Messages +Output Directory +Generated [0-9]+ files +exit code diff --git a/packages/casan-harness/tests/phase08-compression-tests.sh b/packages/casan-harness/tests/phase08-compression-tests.sh index b3b4f56..49358a4 100644 --- a/packages/casan-harness/tests/phase08-compression-tests.sh +++ b/packages/casan-harness/tests/phase08-compression-tests.sh @@ -10,6 +10,7 @@ SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" source "$SCRIPT_DIR/../scripts/bash/casan-paths.sh" PROJECT_ROOT="$CASAN_APP_ROOT" CC="$CASAN_HARNESS_ROOT/scripts/bash/context-compress.py" +FIXTURES="$SCRIPT_DIR/fixtures/context-compression/nehops" WORK="$(mktemp -d)" trap 'rm -rf "$WORK"' EXIT @@ -89,6 +90,178 @@ set -e 2>/dev/null || true [[ "$RC" -eq 0 ]] && pass "protecting the pattern keeps the line and passes the gate" \ || fail "protected must-keep still failed the gate (rc=$RC)" +# 8) Realistic NEHOPS regression: abbreviated [WRN] tags and domain invariants +# survive structural compression, with hash-bound evidence that labels token +# figures as whitespace estimates. +set +e +python3 "$CC" --mode structural \ + --input "$FIXTURES/analyzer-sanitized.log" \ + --must-keep-file "$FIXTURES/must-keep.patterns" \ + --require-must-keep-file "$FIXTURES/must-keep.patterns" \ + --report-json "$WORK/nehops-report.json" \ + >"$WORK/nehops-compressed.log" 2>"$WORK/nehops.err" +RC=$? +set -e 2>/dev/null || true +RAW_WARNINGS="$(grep -c '\[WRN\]' "$FIXTURES/analyzer-sanitized.log")" +OUT_WARNINGS="$(grep -c '\[WRN\]' "$WORK/nehops-compressed.log" || true)" +if [[ "$RC" -eq 0 && "$RAW_WARNINGS" -eq 2 && "$OUT_WARNINGS" -eq "$RAW_WARNINGS" ]] \ + && grep -q "FRR04700_CL.*common code tracing skipped" "$WORK/nehops-compressed.log" \ + && grep -q "FRR04700_SV.*common code tracing skipped" "$WORK/nehops-compressed.log" \ + && grep -q "Error Messages: 0" "$WORK/nehops-compressed.log"; then + pass "NEHOPS [WRN] + CL/SV/domain invariants survive compression" +else + fail "NEHOPS warning preservation failed (rc=$RC raw=$RAW_WARNINGS out=$OUT_WARNINGS)" +fi + +set +e +python3 - "$WORK/nehops-report.json" <<'PY' +import json +import sys + +report = json.load(open(sys.argv[1], encoding="utf-8")) +assert report["decision"] == "compressed" +assert report["preservation"]["status"] == "passed" +assert report["preservation"]["protected_line_missing"] == 0 +assert report["preservation"]["required_pattern_deficits"] == [] +assert report["measurement_source"] == "whitespace_estimate" +assert report["provider_telemetry"] is None +assert report["saving_status"] == "estimate_only_quality_gate_required" +assert report["severity"]["raw"]["warning"] == 2 +assert report["severity"]["candidate"]["warning"] == 2 +assert len(report["hashes"]["raw_sha256"]) == 64 +assert len(report["hashes"]["candidate_sha256"]) == 64 +PY +RC=$? +set -e 2>/dev/null || true +[[ "$RC" -eq 0 ]] \ + && pass "JSON evidence binds hashes and labels estimate/provider status" \ + || fail "JSON evidence contract invalid" + +# 9) Required-pattern multiplicity: retaining one of two matching lines is not +# sufficient. Rejection defaults to raw fallback and reports zero savings. +printf 'TRACE-MUST plain line\nERROR TRACE-MUST protected line\n' > "$WORK/multi.txt" +printf 'TRACE-MUST\n' > "$WORK/multi.patterns" +set +e +python3 "$CC" --mode extractive --input "$WORK/multi.txt" \ + --require-must-keep-file "$WORK/multi.patterns" \ + >"$WORK/multi.out" 2>"$WORK/multi.err" +RC=$? +set -e 2>/dev/null || true +if [[ "$RC" -eq 1 ]] \ + && cmp -s "$WORK/multi.txt" "$WORK/multi.out" \ + && grep -q "TRACE-MUST(1/2)" "$WORK/multi.err" \ + && grep -q "saving_status=rejected" "$WORK/multi.err" \ + && grep -q "saved=0" "$WORK/multi.err"; then + pass "partial must-keep loss is rejected with raw fallback and no saving claim" +else + fail "multiplicity/fallback contract failed (rc=$RC)" +fi + +# 10) A caller may choose halt instead of raw fallback. Halt returns non-zero +# and emits no candidate bytes. +set +e +python3 "$CC" --mode extractive --input "$WORK/multi.txt" \ + --require-must-keep-file "$WORK/multi.patterns" \ + --on-preservation-failure halt \ + >"$WORK/halt.out" 2>"$WORK/halt.err" +RC=$? +set -e 2>/dev/null || true +if [[ "$RC" -eq 1 && ! -s "$WORK/halt.out" ]] \ + && grep -q "COMPRESS_REJECTED fallback=halt" "$WORK/halt.err" \ + && grep -q "saving_status=rejected" "$WORK/halt.err" \ + && grep -q "saved=0" "$WORK/halt.err"; then + pass "halt policy emits no lossy candidate on preservation failure" +else + fail "halt policy leaked output or returned success (rc=$RC)" +fi + +# 11) Required patterns absent from raw are not fabricated and do not create a +# false failure; preservation checks only claims that raw evidence contained. +printf 'ERROR real failure line\n' > "$WORK/absent.txt" +printf 'NEVER_PRESENT_IN_RAW\n' > "$WORK/absent.patterns" +set +e +python3 "$CC" --mode extractive --input "$WORK/absent.txt" \ + --require-must-keep-file "$WORK/absent.patterns" \ + >"$WORK/absent.out" 2>"$WORK/absent.err" +RC=$? +set -e 2>/dev/null || true +[[ "$RC" -eq 0 ]] \ + && pass "pattern absent from raw does not create a false preservation failure" \ + || fail "absent raw pattern incorrectly failed preservation (rc=$RC)" + +# 12) Invalid project regex is a configuration defect and fails closed without +# emitting a candidate. +printf '[unterminated\n' > "$WORK/invalid.patterns" +set +e +python3 "$CC" --mode structural --input "$WORK/absent.txt" \ + --must-keep-file "$WORK/invalid.patterns" \ + >"$WORK/invalid.out" 2>"$WORK/invalid.err" +RC=$? +set -e 2>/dev/null || true +if [[ "$RC" -eq 1 && ! -s "$WORK/invalid.out" ]] \ + && grep -q "COMPRESS_FAIL invalid_must_keep_regex" "$WORK/invalid.err"; then + pass "invalid must-keep regex fails closed" +else + fail "invalid regex did not fail closed (rc=$RC)" +fi + +# 13) Non-empty raw input may not silently compress to an empty view. +printf 'ordinary informational chatter\n' > "$WORK/empty-loss.txt" +set +e +python3 "$CC" --mode structural --input "$WORK/empty-loss.txt" \ + >"$WORK/empty-loss.out" 2>"$WORK/empty-loss.err" +RC=$? +set -e 2>/dev/null || true +if [[ "$RC" -eq 1 ]] \ + && cmp -s "$WORK/empty-loss.txt" "$WORK/empty-loss.out" \ + && grep -q "COMPRESS_REJECTED fallback=raw" "$WORK/empty-loss.err"; then + pass "empty compressed view is rejected and raw is preserved" +else + fail "non-empty raw input was allowed to disappear (rc=$RC)" +fi + +# 14) Dedup may collapse repeated warnings only when the multiplicity marker +# proves how many occurrences were present. +printf '[WRN] repeated warning\n[WRN] repeated warning\n[WRN] repeated warning\n' > "$WORK/repeated-warning.txt" +set +e +python3 "$CC" --mode structural --input "$WORK/repeated-warning.txt" \ + >"$WORK/repeated-warning.out" 2>"$WORK/repeated-warning.err" +RC=$? +set -e 2>/dev/null || true +if [[ "$RC" -eq 0 ]] \ + && grep -q '\[WRN\] repeated warning (x3)' "$WORK/repeated-warning.out"; then + pass "dedup preserves repeated-warning multiplicity" +else + fail "repeated-warning multiplicity was lost (rc=$RC)" +fi + +# 15) A literal line ending in "(xN)" is content, not necessarily a dedup +# marker. It must survive in both non-dedup and dedup modes. +printf 'ERROR literal marker (x3)\n' > "$WORK/literal-marker.txt" +set +e +python3 "$CC" --mode extractive --input "$WORK/literal-marker.txt" \ + >"$WORK/literal-marker.out" 2>"$WORK/literal-marker.err" +RC=$? +set -e 2>/dev/null || true +if [[ "$RC" -eq 0 ]] \ + && cmp -s "$WORK/literal-marker.txt" "$WORK/literal-marker.out"; then + pass "extractive mode does not misread a literal (xN) suffix" +else + fail "literal (xN) content triggered a false preservation failure (rc=$RC)" +fi + +set +e +python3 "$CC" --mode structural --input "$WORK/literal-marker.txt" \ + >"$WORK/literal-marker-structural.out" 2>"$WORK/literal-marker-structural.err" +RC=$? +set -e 2>/dev/null || true +if [[ "$RC" -eq 0 ]] \ + && cmp -s "$WORK/literal-marker.txt" "$WORK/literal-marker-structural.out"; then + pass "structural mode disambiguates a literal (xN) suffix" +else + fail "structural mode misread literal (xN) content (rc=$RC)" +fi + echo "" echo "===== Plan-08 ⟷ Control Plane: settings govern harness =====" CPS="$CASAN_HARNESS_ROOT/scripts/bash/control-plane-settings.py" diff --git a/tmp/casan-training-montage.png b/tmp/casan-training-montage.png new file mode 100644 index 0000000..532e8c4 Binary files /dev/null and b/tmp/casan-training-montage.png differ diff --git a/tmp/casan-training-render/slide-01.png b/tmp/casan-training-render/slide-01.png new file mode 100644 index 0000000..a0e315e Binary files /dev/null and b/tmp/casan-training-render/slide-01.png differ diff --git a/tmp/casan-training-render/slide-02.png b/tmp/casan-training-render/slide-02.png new file mode 100644 index 0000000..29f923c Binary files /dev/null and b/tmp/casan-training-render/slide-02.png differ diff --git a/tmp/casan-training-render/slide-03.png b/tmp/casan-training-render/slide-03.png new file mode 100644 index 0000000..aaf28ba Binary files /dev/null and b/tmp/casan-training-render/slide-03.png differ diff --git a/tmp/casan-training-render/slide-04.png b/tmp/casan-training-render/slide-04.png new file mode 100644 index 0000000..e8bd928 Binary files /dev/null and b/tmp/casan-training-render/slide-04.png differ diff --git a/tmp/casan-training-render/slide-05.png b/tmp/casan-training-render/slide-05.png new file mode 100644 index 0000000..1a06946 Binary files /dev/null and b/tmp/casan-training-render/slide-05.png differ diff --git a/tmp/casan-training-render/slide-06.png b/tmp/casan-training-render/slide-06.png new file mode 100644 index 0000000..0f03a80 Binary files /dev/null and b/tmp/casan-training-render/slide-06.png differ diff --git a/tmp/casan-training-render/slide-07.png b/tmp/casan-training-render/slide-07.png new file mode 100644 index 0000000..b342f8c Binary files /dev/null and b/tmp/casan-training-render/slide-07.png differ diff --git a/tmp/casan-training-render/slide-08.png b/tmp/casan-training-render/slide-08.png new file mode 100644 index 0000000..4afe5b3 Binary files /dev/null and b/tmp/casan-training-render/slide-08.png differ diff --git a/tmp/casan-training-render/slide-09.png b/tmp/casan-training-render/slide-09.png new file mode 100644 index 0000000..e1f8bbd Binary files /dev/null and b/tmp/casan-training-render/slide-09.png differ diff --git a/tmp/casan-training-render/slide-10.png b/tmp/casan-training-render/slide-10.png new file mode 100644 index 0000000..8a1a103 Binary files /dev/null and b/tmp/casan-training-render/slide-10.png differ diff --git a/tmp/casan-training-render/slide-11.png b/tmp/casan-training-render/slide-11.png new file mode 100644 index 0000000..3c5ee68 Binary files /dev/null and b/tmp/casan-training-render/slide-11.png differ diff --git a/tmp/casan-training-render/slide-12.png b/tmp/casan-training-render/slide-12.png new file mode 100644 index 0000000..822f50d Binary files /dev/null and b/tmp/casan-training-render/slide-12.png differ diff --git a/tmp/casan-training-render/slide-13.png b/tmp/casan-training-render/slide-13.png new file mode 100644 index 0000000..33fd413 Binary files /dev/null and b/tmp/casan-training-render/slide-13.png differ diff --git a/tmp/casan-training-render/slide-14.png b/tmp/casan-training-render/slide-14.png new file mode 100644 index 0000000..92272cb Binary files /dev/null and b/tmp/casan-training-render/slide-14.png differ diff --git a/tmp/casan-training-render/slide-15.png b/tmp/casan-training-render/slide-15.png new file mode 100644 index 0000000..83ced1e Binary files /dev/null and b/tmp/casan-training-render/slide-15.png differ diff --git a/tmp/casan-training-render/slide-16.png b/tmp/casan-training-render/slide-16.png new file mode 100644 index 0000000..eca11fb Binary files /dev/null and b/tmp/casan-training-render/slide-16.png differ diff --git a/tmp/casan-training-render/slide-17.png b/tmp/casan-training-render/slide-17.png new file mode 100644 index 0000000..b7d9204 Binary files /dev/null and b/tmp/casan-training-render/slide-17.png differ diff --git a/tmp/casan-training-render/slide-18.png b/tmp/casan-training-render/slide-18.png new file mode 100644 index 0000000..2b2351e Binary files /dev/null and b/tmp/casan-training-render/slide-18.png differ diff --git a/tmp/casan-training-render/slide-19.png b/tmp/casan-training-render/slide-19.png new file mode 100644 index 0000000..e2ffb19 Binary files /dev/null and b/tmp/casan-training-render/slide-19.png differ diff --git a/tmp/casan-training-render/slide-20.png b/tmp/casan-training-render/slide-20.png new file mode 100644 index 0000000..229021e Binary files /dev/null and b/tmp/casan-training-render/slide-20.png differ diff --git a/tmp/casan-training-render/slide-21.png b/tmp/casan-training-render/slide-21.png new file mode 100644 index 0000000..f232071 Binary files /dev/null and b/tmp/casan-training-render/slide-21.png differ diff --git a/tmp/casan-training-render/slide-22.png b/tmp/casan-training-render/slide-22.png new file mode 100644 index 0000000..82d9581 Binary files /dev/null and b/tmp/casan-training-render/slide-22.png differ diff --git a/tmp/casan-training-render/slide-23.png b/tmp/casan-training-render/slide-23.png new file mode 100644 index 0000000..e79b6f4 Binary files /dev/null and b/tmp/casan-training-render/slide-23.png differ diff --git a/tmp/casan-training-render/slide-24.png b/tmp/casan-training-render/slide-24.png new file mode 100644 index 0000000..28c1fef Binary files /dev/null and b/tmp/casan-training-render/slide-24.png differ diff --git a/tmp/casan-training-render/slide-25.png b/tmp/casan-training-render/slide-25.png new file mode 100644 index 0000000..42023a0 Binary files /dev/null and b/tmp/casan-training-render/slide-25.png differ diff --git a/tmp/casan-training-render/slide-26.png b/tmp/casan-training-render/slide-26.png new file mode 100644 index 0000000..947e68f Binary files /dev/null and b/tmp/casan-training-render/slide-26.png differ diff --git a/tmp/casan-training-render/slide-27.png b/tmp/casan-training-render/slide-27.png new file mode 100644 index 0000000..58ab5c5 Binary files /dev/null and b/tmp/casan-training-render/slide-27.png differ diff --git a/tmp/casan-training-render/slide-28.png b/tmp/casan-training-render/slide-28.png new file mode 100644 index 0000000..b2f8c58 Binary files /dev/null and b/tmp/casan-training-render/slide-28.png differ diff --git a/tmp/casan-training-render/slide-29.png b/tmp/casan-training-render/slide-29.png new file mode 100644 index 0000000..014e784 Binary files /dev/null and b/tmp/casan-training-render/slide-29.png differ diff --git a/tmp/casan-training-render/slide-30.png b/tmp/casan-training-render/slide-30.png new file mode 100644 index 0000000..4b24b14 Binary files /dev/null and b/tmp/casan-training-render/slide-30.png differ diff --git a/tmp/casan-training-render/slide-31.png b/tmp/casan-training-render/slide-31.png new file mode 100644 index 0000000..336cc18 Binary files /dev/null and b/tmp/casan-training-render/slide-31.png differ diff --git a/tmp/casan-training-render/slide-32.png b/tmp/casan-training-render/slide-32.png new file mode 100644 index 0000000..7ec28f3 Binary files /dev/null and b/tmp/casan-training-render/slide-32.png differ diff --git a/tmp/casan-training-render/slide-33.png b/tmp/casan-training-render/slide-33.png new file mode 100644 index 0000000..6155bf4 Binary files /dev/null and b/tmp/casan-training-render/slide-33.png differ diff --git a/tmp/casan-training-render/slide-34.png b/tmp/casan-training-render/slide-34.png new file mode 100644 index 0000000..c1ffc08 Binary files /dev/null and b/tmp/casan-training-render/slide-34.png differ diff --git a/tmp/casan-training-render/slide-35.png b/tmp/casan-training-render/slide-35.png new file mode 100644 index 0000000..2482f84 Binary files /dev/null and b/tmp/casan-training-render/slide-35.png differ diff --git a/tmp/casan-training-render/slide-36.png b/tmp/casan-training-render/slide-36.png new file mode 100644 index 0000000..39b223d Binary files /dev/null and b/tmp/casan-training-render/slide-36.png differ diff --git a/tmp/casan-training-render/slide-37.png b/tmp/casan-training-render/slide-37.png new file mode 100644 index 0000000..e87d6ee Binary files /dev/null and b/tmp/casan-training-render/slide-37.png differ diff --git a/tmp/casan-training-render/slide-38.png b/tmp/casan-training-render/slide-38.png new file mode 100644 index 0000000..712ceab Binary files /dev/null and b/tmp/casan-training-render/slide-38.png differ diff --git a/tmp/casan-training-render/slide-39.png b/tmp/casan-training-render/slide-39.png new file mode 100644 index 0000000..d6b55ae Binary files /dev/null and b/tmp/casan-training-render/slide-39.png differ diff --git a/tmp/casan-training-render/slide-40.png b/tmp/casan-training-render/slide-40.png new file mode 100644 index 0000000..f94539b Binary files /dev/null and b/tmp/casan-training-render/slide-40.png differ