Files
CASAN/AINative_OKR_CASAN5/.specify/scripts/bash/tool-output-scan.sh
T
thanhnvandClaude Opus 4.8 ac40c0b281 feat(h4-hardening): Plan-07 Track A — A1 strict semantic, A2 unicode/encoding, A3 tool-output scan
A1 (V1): CASAN_SECURITY_STRICT=1 makes semantic classification REQUIRED and
  fail-closed — model unavailable/no-verdict → BLOCK, never a silent SKIP.
  Non-strict CASAN_SEMANTIC_CLASSIFY=1 keeps regex verdict but logs
  SEMANTIC_SKIPPED loudly (sourced casan-log.sh). Default (no flags) unchanged.
A2 (V3/V4): unicode-normalize.py (NFKC + zero-width strip + Cyrillic/Greek
  homoglyph fold) and decode-suspicious.py (base64/hex decode + rescan, printable
  filter to avoid false positives) feed new match_either/secret_match haystacks.
  Blocks homoglyph, zero-width, fullwidth, base64/hex-smuggled injection & secrets.
A3 (V7): tool-output-scan.sh scans tool output for injection/secret before it
  re-enters model context; wrapper runs it after H6-exec (mode off|warn|block,
  strict→block). warn is default to preserve benign-draft behaviour.

Baseline preserved: run-casan4 35/35, adversarial 44/44.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-07-03 22:20:50 +09:00

53 lines
2.0 KiB
Bash
Executable File

#!/usr/bin/env bash
set -uo pipefail
# CASAN H4 — Tool-output indirect-injection scanner (Track A, V7).
#
# A tool (shell command, file read, web fetch, sub-agent) can return content
# that is then fed back into a downstream model's context. If that content
# carries a prompt injection, the model can be hijacked even though the ORIGINAL
# user input was clean. This scans a tool-output file the same way an untrusted
# artifact is scanned, BEFORE the output is allowed to re-enter model context.
#
# It reuses security-check.sh in `input` mode (block-pattern + unicode/encoding
# normalization + secret detection) but forces the semantic/strict model path
# OFF so the scan is deterministic and needs no model backend — this is a
# pattern scan of machine output, not a user-intent classification.
#
# Usage:
# tool-output-scan.sh <tool-output-file> [context-label]
# Exit:
# 0 — safe to reuse
# 2 — injection / secret pattern detected (caller should reject/quarantine)
# 64 — usage error (file missing)
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
OUTPUT_FILE="${1:-}"
LABEL="${2:-unknown-tool}"
if [[ -z "$OUTPUT_FILE" || ! -f "$OUTPUT_FILE" ]]; then
echo "Usage: tool-output-scan.sh <tool-output-file> [context-label]" >&2
exit 64
fi
WORK="$(mktemp -d)"; trap 'rm -rf "$WORK"' EXIT
SCAN_OUT="$WORK/tool-output-scan.txt"
TIMESTAMP="$(date -u +"%Y-%m-%dT%H:%M:%SZ")"
# Deterministic pattern scan: semantic + strict explicitly disabled here so a
# tool-output scan never depends on (or is blocked by) model availability.
CASAN_SECURITY_STRICT=0 CASAN_SEMANTIC_CLASSIFY=0 \
bash "$SCRIPT_DIR/security-check.sh" "$OUTPUT_FILE" "$SCAN_OUT" input >/dev/null 2>&1
SC_RC=$?
if [[ "$SC_RC" -eq 2 ]]; then
echo "TOOL_OUTPUT_SCAN_BLOCKED label=$LABEL file=$OUTPUT_FILE reason=injection_or_secret timestamp=$TIMESTAMP"
exit 2
elif [[ "$SC_RC" -ne 0 ]]; then
echo "TOOL_OUTPUT_SCAN_ERROR label=$LABEL rc=$SC_RC" >&2
exit 2 # fail closed on scan error
fi
echo "TOOL_OUTPUT_SCAN_CLEAN label=$LABEL file=$OUTPUT_FILE timestamp=$TIMESTAMP"
exit 0