EPIC R05 (Team Hoa) - one security/approval path for every tool call.
R05-T01 domain/tools/{tool_descriptor,tool_registry}.py
ToolCapability (READ/WRITE/EXECUTE/NETWORK, composable) + ToolDescriptor +
ToolRegistry, replacing three independently-maintained gating lists
(core/tools.py::WRITE_TOOLS, code_agent.py's WRITE_TOOLS|MS365_WRITE_TOOLS,
chat_agent.py's literal ("run_command","install_package") tuple) with one
capability lookup.
R05-T02 infrastructure/filesystem/{file_tools,command_tools,fetch_tools,tool_context}.py
core/tools.py's execute_tool if/elif chain split into per-concern modules.
core/tools.py is now a strangler-fig shim: re-exports ToolContext/ToolError,
dispatches through a {name: handler} dict built from the split modules.
core/tools.py: 566 -> 291 lines.
R05-T03 application/conversations/tool_policy_gateway.py
ToolPolicyGateway.allow(name, gate, payload) - capability-driven ALLOW vs
ask-the-gate decision. Wired into both chat_agent.py::run_cowork and
code_agent.py::run_code, replacing their separate hand-rolled checks.
Verified equivalent to the old hardcoded sets by test.
R05-T04 (behavior change, not just refactor)
MCP/connector tools (core/mcp_client.py, core/ext_connectors.py) reached
chat_agent.py via extra_executor(name, args) with NO permission check at
all. They are now tagged with a conservative default capability
(WRITE|EXECUTE|NETWORK - no MCP tool self-declares risk) and routed through
the SAME ToolPolicyGateway as built-ins. When "confirm before running
commands" is on, MCP/connector calls now prompt like run_command already
did - a real gap closed, and a user-visible change worth calling out.
R05-T05 infrastructure/mcp/mcp_source_manager.py
McpToolSourceManager extracts the connection cache/lock/start-or-skip
lifecycle out of state.py::AppContext (_mcp_connections/_conn_lock) into a
standalone, directly-testable class. AppContext.build_mcp_tools and
_ms365_builtin_connection now call ensure()/stop(); _ext_connections
(unified Connectors) is out of scope for this task and keeps its own lock.
New tests: tests/unit/test_tool_registry_and_policy.py,
test_code_agent_tool_policy.py, test_cowork_extra_tool_policy.py,
test_mcp_source_manager.py (26 new tests).
Suite: 254 passed, 4 pre-existing failures unrelated to R05 (2 EPIC R02
config-security, 2 environment-dependent routing tests - see checklist).
check_imports: PASS. All new files < 400 LOC.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
63 lines
2.2 KiB
Python
63 lines
2.2 KiB
Python
"""EPIC R05-T03/T04: ``core/code_agent.py::run_code`` used to gate tool calls
|
|
with ``if name in (WRITE_TOOLS | MS365_WRITE_TOOLS): gate.request(...)``. This
|
|
pins that the switch to ``ToolPolicyGateway`` still gates exactly the same
|
|
calls: ``write_file`` (a WRITE tool) consults the gate; ``list_dir``
|
|
(read-only) never does.
|
|
|
|
Runs the REAL engine (``run_code``) via :class:`FakeProvider`, same approach
|
|
``tests/characterization/test_run_cowork.py`` uses for the Cowork engine.
|
|
"""
|
|
from __future__ import annotations
|
|
|
|
from typing import Any, Dict, List
|
|
|
|
from cowork_local.core.code_agent import run_code
|
|
from cowork_local.core.tools import ToolContext
|
|
from tests.fakes import FakeProvider, ScriptedTurn
|
|
|
|
|
|
class _RecordingGate:
|
|
def __init__(self, approve: bool):
|
|
self.approve = approve
|
|
self.calls: List[Dict[str, Any]] = []
|
|
|
|
def request(self, payload: Dict[str, Any]) -> bool:
|
|
self.calls.append(payload)
|
|
return self.approve
|
|
|
|
|
|
def _run(tmp_path, provider, gate):
|
|
ctx = ToolContext(tmp_path)
|
|
events: List[Dict[str, Any]] = []
|
|
messages: List[Dict[str, Any]] = [{"role": "user", "content": "do it"}]
|
|
run_code(provider, messages, ctx, gate, events.append)
|
|
return events
|
|
|
|
|
|
def test_write_file_consults_the_gate_and_honors_rejection(tmp_path):
|
|
provider = FakeProvider([
|
|
ScriptedTurn(tool_calls=[("write_file", {"path": "a.txt", "content": "hi"})]),
|
|
ScriptedTurn(text="done"),
|
|
])
|
|
gate = _RecordingGate(approve=False)
|
|
events = _run(tmp_path, provider, gate)
|
|
|
|
assert len(gate.calls) == 1 and gate.calls[0]["name"] == "write_file"
|
|
results = [e for e in events if e.get("type") == "tool_result"]
|
|
assert results[0]["ok"] is False
|
|
assert not (tmp_path / "a.txt").exists() # rejected, never actually written
|
|
|
|
|
|
def test_read_only_tool_never_consults_the_gate(tmp_path):
|
|
(tmp_path / "existing.txt").write_text("x", encoding="utf-8")
|
|
provider = FakeProvider([
|
|
ScriptedTurn(tool_calls=[("list_dir", {})]),
|
|
ScriptedTurn(text="done"),
|
|
])
|
|
gate = _RecordingGate(approve=False) # would reject if ever asked
|
|
events = _run(tmp_path, provider, gate)
|
|
|
|
assert gate.calls == []
|
|
results = [e for e in events if e.get("type") == "tool_result"]
|
|
assert results[0]["ok"] is True
|