Files
CASAN/AINative_OKR_CASAN5/packages/casan-harness/tests/phase-sec21-tests.sh
T
thanhnvandClaude Opus 4.8 664bd1f00c feat(plan-01): Phase 1 — relocate harness code to packages/casan-harness (symlink facade)
Physically move the pure-code subtrees out of .specify into the package, leaving
compat symlinks at the old .specify/<dir> paths so every existing reference (internal
CASAN_HARNESS_ROOT + external CI/docker/mjs) keeps resolving. Runtime state stays put.

Moved (git mv): scripts/ tests/ security/ templates/ config/ governance/ memory/
  .specify/<dir>  ->  packages/casan-harness/<dir>   (+ .specify/<dir> symlink)
Stays in .specify (state/governance/domain, handled later): logs/ agentops/ level5/
  init-options.json traceability-map.json

Python `.resolve()` self-location followed the compat symlink into packages and lost
the app root; generate-casan-demo-context.py, generate-agentops-dashboard.py and
dashboard-server.py now walk UP for the `.specify` state marker instead of a fixed
parent depth (fixes "missing trace files" in run-casan4).

Full gate: PASS=64 FAIL=0 SKIP=3 (CASAN_CI_STEP_TIMEOUT_SEC=1200 — track-a ~450s runs
close to the 600s default and can tip over under load; this is timing variance, not a
regression — it passed cleanly with headroom). Runtime log/audit artifacts kept unstaged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-08 00:06:00 +09:00

58 lines
2.5 KiB
Bash

#!/usr/bin/env bash
set -uo pipefail
# CASAN Plan-16 SEC-21 (ARCH-07) — bounded model timeout + per-run call budget.
#
# A 180s-per-call timeout across many steps let a hung model stall a run for tens
# of minutes. Now the per-call timeout is lower + configurable, and total model
# calls per run are capped so a wedged model cannot amplify into a DoS. Proves the
# budget refuses once exhausted, and no cap set means no limit (dev default).
#
# Deterministic (budget is checked BEFORE any backend call, so it holds whether or
# not a model is reachable). No network required.
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
source "$SCRIPT_DIR/../scripts/bash/casan-paths.sh"
PROJECT_ROOT="$(cd "$SCRIPT_DIR/../.." && pwd)"
MC="$CASAN_HARNESS_ROOT/scripts/bash/model-call.py"
WORK="$(mktemp -d)"
trap 'git -C "$PROJECT_ROOT" checkout -- .specify/logs/level5/provider-usage.jsonl 2>/dev/null; rm -rf "$WORK"' EXIT
PASS=0; FAIL=0
pass() { echo "PASS: $1"; PASS=$((PASS + 1)); }
fail() { echo "FAIL: $1"; FAIL=$((FAIL + 1)); }
echo "===== Plan-16 SEC-21: model timeout + per-run call budget ====="
# Per-call timeout is reduced from 180s and configurable.
grep -q 'timeout=180' "$MC" && fail "hardcoded 180s timeout still present" \
|| pass "no hardcoded 180s per-call timeout"
grep -q 'CASAN_MODEL_TIMEOUT_SEC' "$MC" \
&& pass "per-call timeout is configurable via CASAN_MODEL_TIMEOUT_SEC" \
|| fail "timeout not configurable"
# Budget: cap=1. Call 1 charges the budget (proceeds); call 2 is refused BEFORE any
# backend work, with a clear budget error, regardless of model availability.
echo "hello" > "$WORK/p.txt"
CNT="$WORK/counter"
export CASAN_MODEL_MAX_CALLS=1 CASAN_MODEL_CALL_COUNTER_FILE="$CNT"
python3 "$MC" "$WORK/p.txt" "$WORK/o1.json" --role classify >/dev/null 2>&1 # charges to 1
ERR2="$(python3 "$MC" "$WORK/p.txt" "$WORK/o2.json" --role classify 2>&1 >/dev/null)"; RC2=$?
if [[ "$RC2" -ne 0 ]] && echo "$ERR2" | grep -q "run_call_budget_exceeded"; then
pass "call over budget is REFUSED (rc=$RC2, budget error)"
else
fail "over-budget call not refused (rc=$RC2 err='$ERR2')"
fi
# No cap set → no budget error (dev default unchanged).
unset CASAN_MODEL_MAX_CALLS CASAN_MODEL_CALL_COUNTER_FILE
ERR3="$(python3 "$MC" "$WORK/p.txt" "$WORK/o3.json" --role classify 2>&1 >/dev/null || true)"
echo "$ERR3" | grep -q "run_call_budget_exceeded" \
&& fail "budget error fired with no cap set" \
|| pass "no cap set → no budget limit (dev default)"
echo ""
echo "===== SEC-21 SUMMARY: PASS=$PASS FAIL=$FAIL ====="
[[ "$FAIL" -eq 0 ]] || exit 1