Closes the last fully-[planned] Track-C dimension (was scored 1).
- incident.sh raise <event>: classify severity via incident-severity.map
(LOW/MED/HIGH/CRIT), record a structured entry (owner routing), and for
HIGH/CRIT auto-engage the scoped kill-switch + dispatch an alert (reuses H6
alert-dispatch.sh). Exit 2 on HIGH/CRIT so a pipeline gate goes red.
- kill-switch.sh engage/clear/check/status, scoped by project/model/provider
(+ global). `check` exits 2 when engaged so gates honor it.
- casan-harness.sh honors an engaged kill-switch before running (opt-in
CASAN_KILLSWITCH_ENFORCE=1, default OFF → baseline unchanged).
- incident-runbook.md: severity→owner→response + postmortem template + prod TODO.
- phase-c7-incident-tests.sh: 15 checks — severity grading, auto kill-switch on
HIGH/CRIT, MED-only records, lifecycle, global scope, structured record, and
the production wrapper refusing to run under an engaged switch.
Baselines: run-casan4 35/35, adversarial 44/44. New suite total: 175 → 190.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Self-contained, full-screen-projectable HTML visualizing the base → competition
→ now journey: OLED-glass design, 3-milestone timeline, 7 harness bento cards
(color-coded H1–H7 with each file's role), 5 upgrade-track tables (Track A /
C-MVP / Evidence / H5+ / H6+) with gap-closed columns, honest-maturity panel +
one-line flow diagram. UTF-8 standalone doc (opens directly in a browser),
IntersectionObserver scroll reveals, custom cubic-bezier motion, mobile fallback.
Verified rendering in preview. Cross-linked from both architecture md files.
Adds .claude/launch.json (static server for local preview of optimize-docs).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Two companion docs tracing the .specify/scripts evolution from the 5-file
spec-kit scaffold shown in the file tree:
- CASAN_ARCHITECTURE_BEFORE.md — state brought to the competition (freeze
fbcef96): base 5 → 37 scripts implementing all 7 harnesses, each file's
purpose grouped by H1–H7, + the 8 competition test suites. Honest maturity:
demo/PoC (~3.0/5).
- CASAN_ARCHITECTURE_AFTER.md — the feat/plan07-track-a-hardening upgrades:
37 → 60 scripts (+23) grouped by Track A / Track C-MVP / Evidence Pack /
H5+ / H6+, each new file's purpose + the gap it closes, notes on in-place
modifications (strict fail-closed, cost caps, KMS rotate, window breaker),
+ the 6 new test suites (+96 checks). Honest maturity: Level 4 proven by
attack (~4.0/5), not full production.
File counts verified against git (ls-tree fbcef96 vs HEAD).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Deep-dive defense ammo tied to the video scenes so a presenter can field
follow-up questions live. For each of the 26 attacks across the 3 videos:
plain-language what-it-is, why-dangerous, how-CASAN-blocks (flagging which
layer is deterministic vs the 3-4 AI scenes), likely challenge question +
answer, and the honest limit. Opens with 3 "mantra" lines that cover most
hard questions and closes with a 10-toughest-questions cheat sheet.
Link it from CASAN_SCRIPT_2VIDEO.md. Numbers kept consistent: V1 = 23 vectors,
V2 = 140 checks, V3 = 175 checks; ~90% of controls are deterministic.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The cross-layer showpiece stops ONE attack at 4 sequential checkpoints
(injection block → tool-input schema reject → runaway timeout → signed audit),
which span H4 + the Tool layer (H2) + H5 — not "four harness layers". That
contradicted the intro's "this video focuses on 3 layers (H4/H5/H6)". Reframe
the close as one attack caught at each successive stage, listing the 4 concrete
checkpoints, so it's accurate to the demo without mislabeling them as harnesses.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Add an explicit difficulty/depth progression the presenter can say out loud:
- Video 1 = basic & classic attacks (everyone must block these)
- Video 2 = advanced evasion + proof-you-can-trust (not just harder attacks —
half of it is production maturity: fail-closed, evidence pack, FP budget)
- Video 3 = deepest gaps — real governance & ops hardening
Update each video's subtitle and the quick-reference table with a Level column.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
- Rewrite Video 1 (attack battery) and Video 2 (Track A + C-MVP + Evidence Pack)
in plain, presentation-friendly language: each scene now has a 【what's on screen】
cue + 🎙️ spoken line, with everyday analogies (sealed ledger, fresh signature,
key-in-a-vault) so a non-technical viewer follows while the terminal video plays.
- Add Video 3 covering the feat/plan07-track-a-hardening work that was NOT in the
recorded videos: H5+ governance (approval-identity, KMS-in-vault, external WORM)
and H6+ AgentOps (live alerting + dead-letter, provider-API cost reconciliation,
stale-aware hosted dashboard, sliding-window breaker). Framed as the
"lift the weakest link" arc, ending on the honest 175/0 · Level 4 · not-yet-full-
production close.
- Add a quick-reference table + 60-second highlight cut for presenters.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Close the three gaps the scoring report itself flagged for H6 plus V15,
each as a real MVP + fail-able adversarial test (same pattern that lifted H5):
- D1 alert-dispatch.sh: alerts POST to a real HTTP webhook (severity routing,
dedup window, retry) + dead-letter queue with redelivery; fail-loud in strict.
Wired into agent-metrics.sh so a failing step pages live end-to-end.
- D2 provider-usage-fetch.sh + telemetry-reconcile.sh: pull usage from a provider
usage HTTP API (all-or-nothing schema gate, fail-loud) + reconcile local vs
provider ground truth — token under-reporting/hidden runs => TELEMETRY_DISCREPANCY.
- D3 dashboard-serve.sh + dashboard-server.py: serve the dashboard over HTTP with
a stale-aware /healthz probe (fresh=200 ok, telemetry silent-death=503 stale).
- D4 circuit-breaker-check.sh: sliding-window failure-rate breaker (V15) — interleaved
successes no longer evade the consecutive-failure breaker (CIRCUIT_OPEN_WINDOW).
New suite phase-h6-agentops-tests.sh: 20/20, all live against local HTTP endpoints
(webhook sink, mock provider API, dashboard server) — deterministic, no model needed.
Also fix sign-policy-bundle.sh key-sync invariant: the local-fallback branch only
exported policy-public.pem when generating a NEW key, so a Vault-DOWN run after a
Vault-signed run verified a local-key signature against the Vault pubkey (RSA padding
error, run-casan4 died mid-suite). Now always re-exports the pubkey before signing —
same fix class as tool-audit-lib.sh / governance-check.sh.
Full battery re-run sequentially: 175/175 PASS, 0 FAIL across 8 suites
(KMS SKIP this run — Vault down; validated live 2026-07-04). Docs synced:
scoring-run-report (H6 79→80, no harness below 80, 155→175), CASAN_HARDENING_STATUS
(Phase 5 D1–D4), Plan-07, submission README, and run-hardening.sh (H6+ scenes HO1–HO4).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
① Key management (B3/KMS): vault-kms.sh gains `rotate` (Transit key rotation)
and `assert-nonexportable` (proves private material never leaves the KMS).
Validated live against a Vault dev server: sign→verify (v1) → rotate →
sign→verify (v2) → export denied. sign-audit-head.sh already routes to Vault
when VAULT_ADDR/TOKEN are set, so this is the real production signing path.
② External WORM audit (C5/V21): worm-ledger.py + audit-ship.sh append the audit
head to a hash-linked append-only ledger (chattr +a best-effort on Linux;
S3 Object Lock/QLDB in production). verify-audit-gap.sh detects local audit
rollback (AUDIT_GAP_DETECTED — the durable ledger still holds the later head)
and ledger tampering (AUDIT_LEDGER_TAMPERED).
phase-h5-infra-tests.sh: 7 checks — KMS sign/rotate/non-exportable (skip-aware,
live when Vault reachable) + WORM in-sync/rollback/tamper (always local).
Baselines: run-casan4 35/35, adversarial 44/44, approval 8/8. Lifts H5
key-mgmt 2.5→~4 (KMS live path + rotation + non-exportable) and external-audit
1.5→~3.5 (WORM ledger + gap detection). Suites now 7 (+7 = 155 checks).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Under CASAN_APPROVAL_STRICT=1, a high-risk approval is trusted ONLY when a
REGISTERED reviewer cryptographically signs THIS exact request and their role is
authorized for the action — a plain env-var CASAN_APPROVER is no longer enough.
- approval-sign.sh: reviewer signs assertion
"casan-approval|v1|<action>|<actor>|<input_sha256>|<approver_id>" with their key.
- approval-verify.sh: gate looks up reviewer role+pubkey in reviewers.registry,
enforces role→action authorization, verifies the RSA signature (fail-closed).
- governance-check.sh: strict branch requires a valid signed approval; SoD still
enforced; default (non-strict) env-var path UNCHANGED (baseline preserved).
- reviewers.registry: role-scoped reviewer identity registry (pubkeys off-repo;
production replaces with OIDC/JWT from a real IdP).
- phase-h5-approval-tests.sh: 8 checks — valid/authorized approve; unsigned,
wrong-role, forged-key, unregistered, replay-to-other-request, self-approval
all denied; non-strict backward-compat.
Baselines: run-casan4 35/35, adversarial 44/44. Lifts H5 policy-approval (C4)
2.5 -> ~3.5-4 / 5. Total suites now 6 (+8 checks = 148).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
One plain-language closing paragraph per video for time-constrained delivery:
Video 1 (attacks blocked, proven by exit codes) and Video 2 (production
hardening + signed evidence pack). Jargon-light, keeps both closers.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
CASAN_SCRIPT_2VIDEO.md — tight summary voiceover for just the two demo videos,
one screen each: Video 1 (battery: H4/H5/H6 + chain) and Video 2 (hardening
Track A + C-MVP + Evidence Pack). Strips the per-scene cues/notes/soundbites of
the full narration; keeps the spoken beats and both closers.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
CASAN_VIDEO_NARRATION.md — spoken lines timed to each on-screen scene of
run-all.sh (Video 1) and run-hardening.sh (Video 2). Golden rule: terminal says
WHAT, narrator says WHY — never read the screen. Per-scene ▶cue / 🎙️line / ⏸pause
markers, hero-scene emphasis (A5, B1, B6, D1, chain, HA1, HA4, HE3 money-shot),
clustered narration for fast scenes, delivery notes (tone/pace/verdict timing),
a soundbite bank for judge Q&A, and a duration table (~13-16 min). Aligned with
existing claims: Level 4 proven, AI as optional escalation, sandbox scaffold
honesty, "casan-old wins a demo; casan5 survives production".
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- CASAN_ATTACK_CATALOG.md: one-page visual map of ALL ~38 attack scenes from
run-all.sh (A1-A9, B1-B8, D1-D6, chain) + run-hardening.sh (HA1-7, HC1-4,
HE1-4). Legend for 12 attack directions (input/output/artifact/tool-out/audit/
telemetry/cost/action/supply/exfil/runtime/evidence), a defense-in-depth
diagram, per-scene matrix with control + verdict + AI marker.
- CASAN_ATTACK_PLAYBOOK.md: internal deep-dive — per attack: scenario, why
dangerous, exact blocking mechanism, AI-or-deterministic, verify command +
expected result, threat-model ref. Prominent AI-usage answer: only A3/A8/D4
invoke the model; everything else is deterministic. Semantic AI is an optional
escalation that only ADDS a block, never removes one.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
- CASAN_HARDENING_STATUS.md: canonical implemented/scaffold/planned record with
per-control test mapping and test inventory (baseline 79 + 61 new = 140 checks).
- README claim boundary: separate implemented+tested controls from planned;
explicitly does NOT claim full production-readiness (Track B, C-Gov/Ops, true
sandbox isolation, IdP/WORM still planned).
- INDEX status table: Plan-07 Track A + C-MVP done, Plan-09 MVP done.
- video guide: core demo battery counts unchanged (hardening lives in separate
suites); added deep-dive commands + 3 Q&A rows + sandbox honesty note.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
evidence-pack.sh {pack|verify-pack} assembles a per-run proof pack from REAL
on-disk logs (summaries only — no raw secret/PII copied; decision-log passes the
data-exfil guard or the pack aborts). Produces the standard set: run-summary,
h1..h7 reports, redteam-result, benign-fp-report, artifact-manifest, decision-log,
plus a signed manifest head (evidence-pack.sig).
Tamper-evident: verify-pack recomputes every file hash vs artifact-manifest.json
(any change fails) and verifies the RSA signature over manifest-head.txt (a
manifest re-forge fails without the off-repo key).
Certified run: run-summary.certified is true ONLY when required gates pass
(H4 exercised, H5 audit chain valid, H5 telemetry verified, no unresolved cost
spike, benign-FP within budget) and none was silently skipped — missing evidence
records an honest reason and does NOT certify.
CLI mapping: `casan pack <id>` -> evidence-pack.sh pack; `casan verify-pack <id>`
-> evidence-pack.sh verify-pack. phase3-evidence-pack-tests.sh covers it.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
C3 (V19) data-exfil-guard.sh: destination-aware egress checkpoint built on the
H4 detectors. cloud/artifact boundaries fail closed on any secret; audit
boundary emits a PII-masked copy (fail closed on an unmaskable secret).
Covers secret-to-cloud, artifact-leaks-env, and PII-in-audit.
C6 (V22) sandbox-run.sh: static policy pre-check (BLOCK on reading ~/.ssh/creds,
network egress, fork bomb, writes outside workspace, huge-file/disk-fill) plus
ulimit file-size/CPU backstops and the wall-clock timeout. HONEST SCOPE: this
is not kernel isolation — the production target (docker --network=none
--read-only --pids-limit / nsjail) is documented as TODO(C6-prod). Process cap
is opt-in so it never breaks legitimate commands on a busy host.
phase2-track-c-tests.sh: 29 adversarial checks (C1 13, C2 6, C3 4, C6 6).
Baselines preserved: run-casan4 35/35, adversarial 44/44.
Running total: 35 + 44 + 25 + 29 = 133 checks.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
C1 (V17) action-gate.sh: gates the ACTION, not just the tool name. Outcome model
ALLOW/WARN/REQUIRE_APPROVAL/BLOCK. BLOCK on sensitive-file writes (.env, *.pem,
id_rsa, .github/workflows, .ssh, .aws/credentials, .npmrc) and destructive/
remote-exec commands (rm -rf /, curl|bash, chmod 777, git push --force);
REQUIRE_APPROVAL on dependency installs and non-local network egress (clears
only with an audited CASAN_ACTION_APPROVER). Decisions logged to action-gate.jsonl.
C2 (V18) supply-chain-gate.sh + supply-chain-scan.py: diffs package.json /
requirements.txt / pom.xml / build.gradle against a baseline (explicit or git
HEAD). BLOCK on denylisted/known-malicious packages, typosquats (edit-distance 1
to a known package), and dangerous lifecycle scripts (pre/post/install);
REQUIRE_APPROVAL on any new dependency. Emits a dep-diff report and records
which live scanners (npm audit / pip-audit / osv-scanner) are available.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
A4 (V9): telemetry-integrity.sh binds provider-usage.jsonl + cost/metrics.jsonl
to a signed manifest head. Tampering a token flips the head (MISMATCH); an
attacker who rewrites the head cannot re-sign it (SIGNATURE_INVALID) without
the off-repo key. sign-audit-head.sh now also signs telemetry (best-effort).
Supports CASAN_AUDIT_PRIV/PUB overrides for self-contained verification.
A5 (V12/V13/V14): cost-spike-detect.sh adds an absolute per-call cap
(CASAN_COST_ABSOLUTE_MAX_TOKENS, enforced from record #1 → catches slow-boil
and cold-start) and a cumulative budget (CASAN_COST_CUMULATIVE_BUDGET_TOKENS →
catches under-threshold spray), keeping the existing median×mult spike test.
Backward compatible: <3 records with no caps still exits 3.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
The cloud branch of model-call.py was a stub (cloud_backend_not_implemented,
failed even with a key set); casan-step.mjs gated the judge on a hard-coded
Ollama ping. Wire up the real cloud path so a model can run without Ollama.
- model-call.py: add call_openai() and call_anthropic() (raw urllib, no new
dependency — matches the existing call_ollama). Endpoints hard-pinned to the
SSRF allowlist; keys read from env, never logged. Anthropic sends no
temperature/thinking (rejected as 400 on Opus 4.8/4.7; omitting thinking
keeps the terse one-word classify/judge answer). main() routes by
ollama:/openai:/anthropic: prefix; key-unset still fails closed honestly.
provider-usage.jsonl cost_source is per-backend, keeping ollama's exact
"ollama_local_real_tokens" tag that evidence/tests key on.
- casan-step.mjs: ollamaAvailable() -> modelAvailable() — when
CASAN_MODEL_PRIMARY is a cloud spec with its key set, the judge runs through
the cloud path; otherwise it pings local Ollama as before. Default
(unset CASAN_MODEL_PRIMARY) is unchanged.
- CASAN_MASTER_RUNBOOK.md: update sections 0/1/4/7/8 — cloud is now
implemented (not a stub); keep the honest "untested with a real key" +
CA-cert caveats.
Not verified against a live API key (none available); confirmed key-set makes
a real HTTPS call and key-unset fails closed. Gates unchanged:
security-gate PASS=11 FAIL=0, adversarial PASS=44 FAIL=0.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add REAL=1 to the video-steps demo so attack vectors flow through the real
production entry-point instead of calling sub-scripts directly.
- run-all.sh: REAL=1 feeds each H4 vector (A1/A2/A4/A6/A7 + cross-layer step
1) as the INPUT of an agent step run through casan-harness.sh, so the
BLOCK/PASS verdict is produced by the wrapper itself (H4-in -> H5 -> H6 ->
exec -> H4-out) exactly as when the real pipeline meets malicious input.
After the battery it runs a real pipeline slice (STEP1 okr.srs via
casan-harness.sh -- node casan-step.mjs) and shows audit.jsonl growing by a
real record. An inline inventory documents which vectors intentionally keep
calling a single control directly (artifact-scan, audit tamper/re-forge,
detectors on synthetic telemetry) and why. Default mode (no REAL) unchanged.
- map-live.sh: show the PIPELINE (STEP1) row only under REAL=1, driven by a
mode sidecar file written by run-all.sh.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Add a shared log taxonomy (error<warn<info<debug<trace, default info) wired
through all three layers, plus a machine-readable per-step run log.
- scripts/casan-log.mjs + .specify/scripts/bash/casan-log.sh: shared logger
([LEVEL] ts [component] msg), stderr-only so stdout contracts and exit codes
are byte-identical. Node side includes trace-level redaction (secret/PII).
- run-casan-pipeline.mjs: per-STEP info line, loop-activation warnings
(BACK-TO-PLAN, FAIL->STEP6), debug harness rc + trace_id, trace payload
excerpt, end-of-run 13-STEP summary table, and one JSONL line/step in
.specify/logs/pipeline-run.jsonl. New --dry-run stubs all agents but keeps
the real wrapper in the loop (deterministic, offline, full 13 STEP + loops).
- casan-step.mjs: debug logs for judge verdict, checkpoint, rollback; logger
import degrades to noop when the file is copied standalone (T1 test).
- casan-harness.sh: debug-log each phase H4-in -> H5 -> [H2-gate] -> H6-exec
-> H4-out with its rc; optional CASAN_PHASE_REPORT JSON for the Boss.
Default level (info) keeps output close to before; behavior opt-in via env.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Hướng A — scorecard.sh (video demo):
- h5_1 approval workflow: hardcode 0 → governance-check deploy live (approval_required)
- h6_2 hallucination rate: hardcode 0 → hallucination-scan phân biệt dirty>clean live
- "N/5 mục" chuyển từ text cứng sang đếm động
- H4/H5/H6 → 100/100 (5/5 gate live), Average 57.9 → 90.0
Hướng B — run-casan-pipeline.mjs:
- fallback: stub 'exit 9' → 'cat /nonexistent' (real failure, nhất quán adversarial T3)
- drift: giữ so fallback-output vs golden (clean run=1.0); năng lực phát hiện
drift thật chứng minh ở adversarial suite
- Full 12-step run verify: H1 CONTEXT_VALID=24, H2 tool-audit records=25 signed,
H5 audit-chain records=22 signed, H6 provider_telemetry per-step thật, H7 rollback real
phase3-real-run-scoring.md: giải thích vì sao scorecard cũ cho H5=60/H6=80
(hardcode), phân biệt scorecard-90 vs re-score-84 (2 mục đích khác nhau).
Verify: adversarial 44/0, security-gate 11/0/0, pipeline 12 steps OK.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The runtime stage reinstalled deps with `npm ci --ignore-scripts`, which skips
bcrypt's install script and never produces bcrypt_lib.node. The app then
crash-looped at runtime ("Cannot find module .../bcrypt_lib.node") — the Prisma
seed and NestJS auth both require bcrypt — so nginx returned 502 on login.
Reuse the builder's node_modules (bcrypt built WITH scripts + generated Prisma
client) instead of reinstalling. builder and runtime share node:20-slim, so the
native binaries are ABI/platform-compatible.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
docker compose reads env_file client-side as the deploy user (ubuntu), so
the root-managed /opt/webapps/webapp-mysql.env (mode 600) caused
"open ...: permission denied" at `docker compose up`. Extend the deploy
preflight to grant docker-group read (chgrp docker + chmod 640) when the
deploy user cannot read it — idempotent, self-heals a rebuilt web VPS.
The live VPS file was already fixed out-of-band; this prevents recurrence.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Two comprehensive CI fixes:
1. Adversarial harness flake (PASS=42 FAIL=2, state-dependent):
governance-check.sh re-exported audit-public.pem only when GENERATING a
new signing key. On a CI runner where the off-repo private key PERSISTS
across runs, the checked-out (Vault-signed) audit-public.pem drifted out
of sync with the local re-signing key, so verify-audit-chain.sh rejected a
genuine head ("H5/H2 verifies the genuine signed chain"). Now always
re-export the public key matching the signing key — mirrors the same fix
already applied to tool-audit-lib.sh (c31987c). Reproduced the exact
FAIL=2 locally with a drifted persisted key; now deterministically 44/44.
2. Deploy docker.sock permission denied:
Added a self-healing preflight to deploy-okr that ensures the deploy user
is in the docker group on the web VPS (idempotent, passwordless sudo) and
proves a fresh SSH session can reach the daemon before streaming images.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Previously the public key was only written on first key generation.
If a CI step (sign-audit-head.sh via Vault KMS) overwrote audit-public.pem
after the key was generated, subsequent calls to append_tool_audit signed
with the local key while audit-public.pem held the Vault key — causing
verify-tool-audit.sh to fail with signature mismatch.
Now the public key is re-exported on every call so audit-public.pem always
matches the private key used to sign tool-calls-head.sig.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Root cause: tool-audit-lib.sh signs tool-calls-head.sig with a local RSA
key (~/.casan/audit-keys/), but sign-audit-head.sh (called as a CI step)
overwrites audit-public.pem with the Vault KMS public key. On the second
run inside security-gate.sh, the local key still exists so audit-public.pem
is NOT updated, leaving a Vault key vs local-key mismatch that causes
verify-tool-audit.sh to exit 1.
Fix 1 — sign-audit-head.sh: after signing the audit.jsonl chain via Vault
KMS, also re-sign the tool-calls chain head with the same casan-audit-key.
Both chains are now anchored to the same Vault public key in audit-public.pem.
Fix 2 — run-casan4-harness-tests.sh: call sign-audit-head.sh just before
the inline verify-tool-audit.sh check (line 217). This re-signs both chains
with Vault KMS so the inline check sees anchor=signed instead of mismatched
local key vs Vault pub.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
printf '%s' with inline secret expansion could leave out trailing newline
or preserve \r from browser-pasted keys, causing OpenSSH 'error in libcrypto'.
Use env: block + printf '%s\n' | tr -d '\r' to normalize the key file.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
Embedding $response as Python triple-quoted string caused json.loads()
to fail with 'Invalid control character' when Vault's JSON contained
\n sequences in the PEM public key (bash \n → Python newline → invalid
JSON control char).
Fix: write response to mktemp, pass path as argv, read with open().
Instead of skipping the frontend test when vitest is not installed,
security-gate.sh now runs `npm ci -w frontend` automatically.
Also add `cache: "npm"` to security-gate's actions/setup-node so the
npm cache from the frontend-tests job is reused — prevents OOM on
the 1GB VPS (cache restore is disk-only, not 300MB download).
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
security-gate.sh checked only `command -v node` before running
`npm test -w frontend`. In the CI security-gate job, node is in PATH
(from actions/setup-node) but root node_modules are NOT installed
(npm ci was removed to prevent OOM). This caused the frontend test to
fail with "Cannot find module vitest".
Fix: also require node_modules/.bin/vitest to exist. Without it the
step SKIPs gracefully — frontend tests are already covered by the
dedicated frontend-tests CI job which runs first.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
- Prisma schema: sqlite → mysql provider
- Migration SQL rewritten as MySQL DDL (utf8mb4, DATETIME(3), AUTO_INCREMENT)
- Add migration_lock.toml for mysql provider
- Dockerfile.backend: drop node:22/sqlite deps, use node:20-slim
- entrypoint.sh: replace SQLite first-run logic with prisma migrate deploy + db seed
- docker-compose.prod.yml: production compose for /opt/webapps/okr on web VPS
- reads DB creds from /opt/webapps/webapp-mysql.env
- reads app secrets from /opt/webapps/okr/.env.app (written by CI)
- port 80 (frontend), no conflict with Gitea 3000/Vault 8200
- ci.yml deploy-okr: moves from ubuntu-latest (web VPS) to ci-runner (161.33.149.243)
- builds images on CI runner VPS (no heavy build on web/Gitea VPS)
- transfers images via docker save | gzip | ssh | docker load
- deploys via SSH + docker compose up on web VPS
- scripts/setup-ci-runner.sh: one-time setup script for CI runner VPS
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
act-runner restart cancelled orphaned security-gate task from run #5.
Concurrency cancel-in-progress will clear run #5 and start fresh.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
The checkout action clones from http://gitea:3000/admin/casan5 — hostname 'gitea'
only resolves on the docker-compose network (gitea_default), not on the default
Docker bridge. Changing container.network: bridge → gitea_default fixes DNS.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>