scorecard: chấm live 2 mục hardcode; pipeline: fallback real + honest scoring doc
Hướng A — scorecard.sh (video demo): - h5_1 approval workflow: hardcode 0 → governance-check deploy live (approval_required) - h6_2 hallucination rate: hardcode 0 → hallucination-scan phân biệt dirty>clean live - "N/5 mục" chuyển từ text cứng sang đếm động - H4/H5/H6 → 100/100 (5/5 gate live), Average 57.9 → 90.0 Hướng B — run-casan-pipeline.mjs: - fallback: stub 'exit 9' → 'cat /nonexistent' (real failure, nhất quán adversarial T3) - drift: giữ so fallback-output vs golden (clean run=1.0); năng lực phát hiện drift thật chứng minh ở adversarial suite - Full 12-step run verify: H1 CONTEXT_VALID=24, H2 tool-audit records=25 signed, H5 audit-chain records=22 signed, H6 provider_telemetry per-step thật, H7 rollback real phase3-real-run-scoring.md: giải thích vì sao scorecard cũ cho H5=60/H6=80 (hardcode), phân biệt scorecard-90 vs re-score-84 (2 mục đích khác nhau). Verify: adversarial 44/0, security-gate 11/0/0, pipeline 12 steps OK. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Sonnet 4.6
parent
fda0d6447d
commit
2f06662f5d
@@ -0,0 +1,86 @@
|
||||
# CASAN Phase 3 — Chấm điểm từ Pipeline Run Thật
|
||||
|
||||
**Ngày:** 2026-07-03
|
||||
**Môi trường:** macOS + Ollama local `ornith:9b` @ 127.0.0.1:11434
|
||||
**Trả lời câu hỏi:** "Vì sao H5/H6 thấp trong scorecard? Có cách nào chấm giống điểm thực tế project thật không?"
|
||||
|
||||
---
|
||||
|
||||
## Vì sao scorecard demo cho H5=60, H6=80 (KHÔNG phải project yếu)
|
||||
|
||||
`optimize-docs/video-steps/scorecard.sh` chấm mỗi harness = (số mục ✓ / 5) × 100.
|
||||
Bản cũ **hardcode 2 mục = 0** dù tính năng có thật và chạy được:
|
||||
|
||||
| Mục | Bản cũ | Sự thật |
|
||||
|---|---|---|
|
||||
| `h5_1` approval workflow | `=0` "chưa demo" | `governance-check.sh deploy` → **GOVERNANCE_DENIED approval_required** (chạy live) |
|
||||
| `h6_2` hallucination rate | `=0` "chưa sinh rate" | `hallucination-scan.py` phân biệt dirty=4 > clean=0 (chạy live) |
|
||||
|
||||
→ Trần cứng H5/H6 tối đa 80. Đây là hạn chế của **cách chấm demo**, không phải thiếu năng lực.
|
||||
|
||||
---
|
||||
|
||||
## Hướng A — Sửa scorecard chấm 2 mục đó LIVE THẬT
|
||||
|
||||
Thay hardcode `=0` bằng gate chạy thật, fail-able:
|
||||
|
||||
- **h5_1**: `governance-check.sh <file> deploy` → pass nếu output có `approval_required|GOVERNANCE_DENIED`
|
||||
- **h6_2**: scan 1 file có marker vs 1 file sạch → pass nếu `dirty > clean` (scanner phân biệt được)
|
||||
|
||||
Phần "N/5 mục ✓" cũng chuyển từ text cứng sang đếm động.
|
||||
|
||||
**Kết quả scorecard sau sửa:**
|
||||
|
||||
```
|
||||
H4 · Security → 100/100 (5/5 gate live)
|
||||
H5 · Governance → 100/100 (5/5 gate live)
|
||||
H6 · AgentOps → 100/100 (5/5 gate live)
|
||||
Average: 57.9 → 90.0/100 CASAN Level 4 — Automated
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Hướng B — Chạy full pipeline THẬT rồi verify từng harness từ artifact sinh ra
|
||||
|
||||
`node scripts/run-casan-pipeline.mjs` chạy 12 bước (SRS → BD → Spec → Review → Plan×2 →
|
||||
DD → Testkit → Tasks → ReviewCode) qua `casan-harness.sh`, mỗi bước đi qua chuỗi
|
||||
H4 security → H5 governance → H2 tool gate → H6 metrics. Verify **từ chính log/artifact vừa sinh**:
|
||||
|
||||
| Harness | Lệnh verify trên artifact pipeline | Kết quả thật |
|
||||
|---|---|---|
|
||||
| **H1** Context | `context-validate.sh pipeline-context.yaml` | `CONTEXT_VALID checked=24` |
|
||||
| **H2** Tool | `verify-tool-audit.sh` | `TOOL_AUDIT_VALID records=25 anchor=signed` |
|
||||
| **H5** Governance | `verify-audit-chain.sh` | `AUDIT_CHAIN_VALID records=22 anchor=signed` (tăng từ 9 — records mới từ run) |
|
||||
| **H6** AgentOps | `metrics.jsonl` per-step | step gọi model (`08-reviewplan`) = `provider_telemetry` 373 real Ollama tokens; step không gọi LLM = `word_count_estimate` — **honest, không đồng nhất giả tạo** |
|
||||
| **H7** Orchestration | rollback + fallback + drift | rollback before==after (restore thật); fallback `primary_exit=1` (real failure từ `cat /nonexistent`, KHÔNG phải stub exit 9); drift PASS vs golden |
|
||||
|
||||
### 2 điểm đã sửa trong pipeline runner để honest
|
||||
|
||||
1. **Fallback**: `bash -c "exit 9"` (stub) → `cat /nonexistent/casan/primary-model-endpoint` (real failure, nhất quán với adversarial suite T3).
|
||||
2. **Drift**: giữ so fallback-output vs golden-baseline (similarity=1.0 = "clean run, no drift"). Năng lực **phát hiện** drift thật (similarity<1.0 trên 2 tài liệu khác nhau) được chứng minh riêng ở `adversarial-harness-tests.sh` (H7 drift).
|
||||
|
||||
---
|
||||
|
||||
## Trạng thái verify cuối (tất cả chạy lại sau thay đổi)
|
||||
|
||||
```bash
|
||||
bash .specify/tests/adversarial-harness-tests.sh # PASS=44 FAIL=0
|
||||
bash .specify/scripts/bash/security-gate.sh # PASS=11 FAIL=0 SKIP=0
|
||||
node scripts/run-casan-pipeline.mjs # 12 steps OK, fallback real, drift PASS
|
||||
NO_COLOR=1 bash optimize-docs/video-steps/scorecard.sh # Average 90.0, H4/H5/H6=100
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Lưu ý quan trọng về con số 90.0 của scorecard
|
||||
|
||||
Average 90.0 trong `scorecard.sh` gồm **4 điểm baseline mang sang** (H1=90, H2=75, H3=85, H7=80
|
||||
từ assessment 2026-06-26) + **3 điểm đo mới** (H4/H5/H6=100). Đây là điểm của **battery gate cô lập**,
|
||||
KHÁC với bản re-score honest per-harness ([phase3-final-rescore.md](phase3-final-rescore.md), ~84 avg)
|
||||
vốn tính cả các gap còn lại (CI gate, cloud recall, KMS).
|
||||
|
||||
**Hai con số phục vụ 2 mục đích khác nhau:**
|
||||
- **Scorecard 90** = năng lực gate H4/H5/H6 khi chạy live (mỗi checklist item = 1 gate thật).
|
||||
- **Re-score ~84** = đánh giá thận trọng per-harness gồm cả residual gaps cần infra.
|
||||
|
||||
Cả hai đều honest, không hardcode, mọi test fail-able.
|
||||
Reference in New Issue
Block a user