Files
a03a740ea1
CI / test (push) Canceled after 0s
Feature/perf ui logic (#13)
## Summary

Nhánh `feature/perf-ui-logic`: tối ưu hiệu năng/UI, sửa lỗi workspace và điều hướng, và làm cho công tắc **"Block network for agent-run commands"** chặn thật mọi đường ra mạng của app, **trừ nhà cung cấp AI**.

**Chặn mạng (b78d483, 8c497cf, 10b8379)**
- Bộ kiểm tra chung `application/network/network_guard.py`, nối vào cấu hình đang chạy ở Composition Root: đổi công tắc trong Settings là có hiệu lực ngay.
- Lệnh shell của agent và task script chạy trong **Windows AppContainer không có quyền mạng**: kernel chặn socket, ping, DNS, Invoke-WebRequest… Không cần quyền admin. Không cô lập được thì lệnh bị từ chối, không chạy khi mạng còn mở. macOS dùng `sandbox-exec`, Linux dùng `unshare --net`.
- Bật chặn thì: dừng MCP đang chạy, không khởi động server mới, từ chối lời gọi connector; Microsoft 365 (đăng nhập, Graph, đồng bộ cloud, rules, mail), Teams, nút Test REST/Jira/MCP, link đính kèm task, pip tự cài và tài nguyên web trong xem trước HTML đều bị từ chối.
- Vẫn dùng được: chat, tải danh sách model, thử model; tool OneDrive đã đồng bộ trên máy.
- Công tắc **mặc định tắt** khi mở app lần đầu; nhãn giữ nguyên như cũ.
- Xem trước HTML trong tab Folder giờ hiện được ảnh/CSS/JS từ web khi mạng mở (trước đây trang `file://` không tải được).
- Sửa lỗi app văng khi chuyển tab Graph → Folder: profile WebEngine của trang xem trước bị huỷ trước trang (`0xc0000409` trong Qt6Core.dll); giờ dùng một profile chung thuộc QApplication.
- Không cấp quyền AppContainer kế thừa lên thư mục chứa PySide6 (nếu có, Chromium không nạp được `Qt6WebEngineCore.dll` và tab Graph trắng).
- Cột mục lục trong Settings tính độ rộng theo kiểu chữ của mục đang chọn, "Sandbox Security Layer" không còn bị cắt.

**Các commit khác trong nhánh**
- `b7a41b3` mỗi thư mục làm việc chỉ thuộc về một project · `bbdf146` bật nút Sửa project khi đã có project đang mở
- `35f24e0`, `cc8d5c8`, `2e3e719`, `c699beb` canh hàng / khoảng cách thanh điều hướng
- `2759ed9` không refresh workspace khi chuyển tab Cowork · `7607f44` checkpoint hiệu năng và UI
- `8548c1e` chặn tool mạng của agent · `caf3b74` renderer GraphRAG native trên macOS · `c00b83c` khoảng cách metadata hàng project · `a04f8a9` ẩn picker workspace cloud

## Change Type

- [x] Cowork feature
- [x] Bug fix
- [ ] Core AI contribution
- [x] Test / hardening
- [x] Performance
- [ ] Documentation

## Related Work

Cowork Task:

Core Repo: http://34.143.229.138/gitea-admin/fsg-ai-core-assets

Core AI Issue:

Core Task:

Related PR:

## Scope

What is intentionally included?
- Mọi đường ra mạng do app tự mở, trừ nhà cung cấp AI (xem Summary).
- Test: `tests/test_network_guard_lanes.py` (có bài chạy AppContainer thật trên Windows), `tests/ui/test_html_preview_remote_images.py`.

What is intentionally NOT included?
- Chặn cả nhà cung cấp AI / chạy model trên máy (Phương án 2).
- Terminal người dùng tự gõ trong tab Folder, sinh ảnh, cơ chế tự tin chứng chỉ lạ (`tls_trust`).
- Huy hiệu trạng thái "đang chặn" trên thanh trên cùng.

## Validation

- [x] Unit tests
- [x] Integration tests
- [x] Manual verification
- [x] Regression check

Commands / evidence:
- `python -m pytest tests/test_network_guard_lanes.py tests/test_sandbox_block_network.py tests/ui -q` → chỉ còn 1 bài fail, fail cả trên `b7a41b3` (nhãn `ProjectRow` 'Project' chưa dịch, `tests/ui/test_i18n_khong_con_chu_cu.py`).
- `python -m pytest tests -q --ignore=tests/ui` → 4 bài fail, cả 4 cũng fail trên `b7a41b3` (`test_canonical_audit_logger`, 2 bài `test_mcp_audit_security`, `test_monitoring_tab_container`).
- Chạy cả `tests` trong một lượt thì treo ở các test dựng MainWindow trong `tests/ui`; `b7a41b3` cũng treo đúng chỗ đó.
- `check_imports.py` và `check_orphan_modules.py` PASS. `check_loc.py` báo 9 file quá dài, giống hệt trước khi sửa (không file nào do nhánh này làm dài thêm).
- Kiểm tra tay trên Windows 11: trong AppContainer, Python báo `WinError 10013`, ping/nslookup/PowerShell/curl đều không ra được mạng; cmd, git, python chạy bình thường.
- Kiểm tra tay trên Windows 11: xem trước HTML tải được 4/4 tài nguyên web khi mạng mở, 0/4 khi bật chặn; tab Graph hoạt động; tạo/huỷ trang xem trước nhiều lần không còn cảnh báo profile của Qt.

## Security Impact

Permission / credential / network / customer data impact:
- Network: khi bật công tắc, chỉ nhà cung cấp AI còn ra mạng; nội dung chat vẫn gửi tới nhà cung cấp AI.
- Permission: lần đầu chạy lệnh trong sandbox, app **thêm quyền (ACE) cho SID AppContainer** trên thư mục làm việc (ghi), thư mục cài Python gốc (đọc), gốc venv và `Scripts` (đọc). Không xoá quyền nào. Thư mục chứa PySide6 không bao giờ nhận quyền kế thừa; một quyền kế thừa sai trên venv (từ bản dev trước) được tự gỡ.
- Credential: không đổi. Khi chặn, trạng thái đăng nhập M365 được đọc thẳng từ kho token trên máy, không dựng MSAL.

## Compatibility

- [x] No breaking change
- [ ] Breaking change documented

Ghi chú: `block_network` mặc định đổi từ bật sang tắt cho cấu hình mới; máy đã lưu `true` thì giữ nguyên. Khi đang chặn, lệnh dùng công cụ cài trong thư mục người dùng (ngoài Program Files) có thể báo Access denied; thư viện trong venv của app không dùng được trong sandbox.

## Reviewer Notes

- `infrastructure/sandbox/appcontainer_process.py` gọi Win32 bằng ctypes (CreateAppContainerProfile, CreateProcessW với SECURITY_CAPABILITIES) và dùng `icacls` để cấp quyền: nên xem kỹ phần cấp quyền.
- `tests/conftest.py` thêm fixture autouse gỡ `network_guard` sau mỗi test, vì `build_context()` gắn cổng này ở mức process.
- `core/task_executors.py` đang đúng bằng trần LOC nên `_run_script` được tách sang `core/task_script.py`.

🤖 Generated with [Claude Code](https://claude.com/claude-code)

---------

Co-authored-by: minhanhpkpro <minhanhpkpro@gmail.com>
Co-authored-by: Duy Le Huu <duylh19@fpt.com>
Co-authored-by: thanhnv <thanhnv.ip@gmail.com>
Reviewed-on: #13
2026-09-20 12:26:03 +00:00

357 lines
18 KiB
Python

"""Agentic loop for the Code tab — Claude-CLI style.
Drives the provider with the file/command tools, routing every write/run action
through the permission gate. Read-only tools run without prompting. For models
without native tool-calling, a text fallback protocol is supported: the model
emits a line ``@@TOOL <name> <json-args>``.
"""
from __future__ import annotations
import json
import re
from pathlib import Path
from typing import Any, Callable, Dict, List, Optional
from ..application.conversations.tool_policy_gateway import ToolPolicyGateway
from ..domain.tools import ToolCapability, ToolDescriptor, ToolRegistry
from ..providers.base import Provider
from . import agent_roles, agent_security
from .mcp_client import UNTRUSTED_MCP_CONTENT_RULE
from .ms365_tools import MS365_WRITE_TOOLS
from .permissions import PermissionGate
from .plan import UPDATE_PLAN_SPEC, normalize_plan_steps
from .tools import TOOL_SPECS, WRITE_TOOLS, ToolContext, describe_action, execute_tool
EmitFn = Callable[[Dict[str, Any]], None]
CancelFn = Callable[[], bool]
MAX_STEPS = 40 # headroom for diagnose → fix → retry loops
_TOOL_LINE = re.compile(r"@@TOOL\s+(\w+)\s+(\{.*\})", re.DOTALL)
def code_system_prompt(workdir: Path, has_memory: bool = False, plan: bool = False,
has_plan_tool: bool = False, has_ms365: bool = False) -> str:
"""Prompt hệ thống cho Code agent, ghép theo năng lực thật của lượt chạy.
Chỉ liệt kê những tool đang BẬT, và thêm ghi chú chế độ lập kế hoạch khi cần —
nói với model về một tool nó không có sẽ khiến nó gọi rồi báo lỗi.
"""
names = ", ".join(t.name for t in TOOL_SPECS)
plan_note = ("PLAN MODE: only analyze and propose a detailed plan; do NOT write files or run "
"commands. When the user asks to gencode/implement, the app switches to ACT.\n"
if plan else "")
memory_note = ""
if has_memory:
memory_note = (
"You also have 'codebase memory' (cmem_*): cmem_get_architecture, cmem_search_graph, "
"cmem_trace_path, cmem_get_code_snippet, cmem_search_code, cmem_query_graph. "
"Prefer these to understand code structure (callers/callees, where functions live) "
"instead of reading/grepping file by file — faster and fewer tokens.\n"
)
plan_tool_note = ""
if has_plan_tool:
plan_tool_note = (
"A step checklist is shown to the user. Call update_plan(steps=[{title, status}]) to "
"keep it in sync as you work — mark the current step 'running', then 'done' when "
"finished (status is one of pending/running/done).\n"
)
ms365_note = ""
if has_ms365:
ms365_note = (
"The user has signed in to Microsoft 365 and enabled some ms365__* tools (Outlook / "
"Teams / OneDrive / SharePoint / meeting transcripts, via the built-in MS365 MCP "
"server). Use them whenever the request involves that data — don't say you can't "
"access it.\n"
)
return (
"You are Cowork Code — a coding assistant that works like a CLI agent.\n"
f"Current working folder: {workdir}\n"
+ plan_note +
"You can call the tools: " + names + ".\n"
+ memory_note + plan_tool_note + ms365_note +
"Break work into steps, read files before editing, and explain each step briefly.\n"
"For changes to an existing file, prefer edit_file (replace an exact snippet) over "
"rewriting the whole file with write_file; use write_file only for new files or full "
"rewrites. Always read_file first so old_string matches exactly.\n"
"If a task needs a Python library that isn't installed, install it yourself with the "
"install_package tool (or `pip install` via run_command) and continue — never ask the "
"user to install libraries by hand. Commands and installs run inside this project's own "
"isolated '.venv' (created automatically on first use), separate from other projects.\n"
"When you must GENERATE a deliverable (e.g. .pptx/.docx/.xlsx/.pdf/images) by writing and "
"running a script: put the generator script and any temporary files in a '.scratch/' "
"subfolder, run it so the final file lands in the working folder, then DELETE the "
"'.scratch/' folder. Only the final requested file(s) should remain — never leave "
"generator scripts or intermediate files behind.\n"
"Every path must stay inside the working folder.\n"
+ UNTRUSTED_MCP_CONTENT_RULE + "\n"
"If a command or tool fails, do NOT stop and hand the error back to the user — read the "
"error, fix the cause (edit the code, install a missing package, correct the command) and "
"retry. Keep iterating until the task actually works, then run it once more so you can "
"show the real output the user asked for.\n"
"If the environment can't call tools directly, emit exactly ONE line of the form:\n"
"@@TOOL <tool_name> {\"param\": \"value\"}\n"
"When the task is complete, reply to the user in plain text — include the produced output "
"— and stop calling tools. Answer in the user's language."
)
_SKILLS_TAG = "[[ACTIVE_SKILLS]]"
_RULES_TAG = "[[SECURITY_RULES]]"
_PROJECT_TAG = "[project-context]"
def _apply_skills(messages: List[Dict[str, Any]], skills_text: str) -> None:
"""Insert/refresh a single system message carrying the enabled skills."""
messages[:] = [
m for m in messages
if not (m.get("role") == "system" and str(m.get("content", "")).startswith(_SKILLS_TAG))
]
if not skills_text.strip():
return
block = {
"role": "system",
"content": f"{_SKILLS_TAG}\nThe user enabled the following skills — follow them:\n\n{skills_text}",
}
insert_at = 1 if messages and messages[0].get("role") == "system" else 0
messages.insert(insert_at, block)
def _apply_security_rules(messages: List[Dict[str, Any]], rules_text: str) -> None:
"""Insert/refresh a single system message carrying the external security/
restriction rules (core/security_rules.py) — same pattern as _apply_skills,
but for mandatory guardrails rather than opt-in behaviors."""
messages[:] = [
m for m in messages
if not (m.get("role") == "system" and str(m.get("content", "")).startswith(_RULES_TAG))
]
if not rules_text.strip():
return
block = {
"role": "system",
"content": (f"{_RULES_TAG}\nMandatory security/restriction rules — check every request "
"and action against these BEFORE acting; refuse or ask for clarification "
f"instead of proceeding if something would violate them:\n\n{rules_text}"),
}
insert_at = 1 if messages and messages[0].get("role") == "system" else 0
messages.insert(insert_at, block)
def _apply_project_context(messages: List[Dict[str, Any]], context_text: str) -> None:
"""Insert/refresh a single system message carrying the project's shared
instructions (Claude-Projects style) — same pattern as _apply_skills, so
editing the project context takes effect on the NEXT turn of every one of
its threads, without duplicating blocks in long conversations."""
messages[:] = [
m for m in messages
if not (m.get("role") == "system" and str(m.get("content", "")).startswith(_PROJECT_TAG))
]
if not (context_text or "").strip():
return
block = {"role": "system", "content": f"{_PROJECT_TAG}\n{context_text.strip()}"}
insert_at = 1 if messages and messages[0].get("role") == "system" else 0
messages.insert(insert_at, block)
_RECOVERY_NOTE = (
"[auto-recovery] The previous attempt hit an error. Review the conversation so "
"far, redo the most recent step if it looks incomplete or wrong, and fix any "
"issue before continuing."
)
def _call_provider_with_recovery(provider: Provider, messages: List[Dict[str, Any]],
tools, on_text, cancel, on_reasoning=None,
max_retries: int = 1) -> Dict[str, Any]:
"""``provider.chat(...)`` with ONE bounded, silent recovery attempt: if the
call raises (a dropped connection, an exhausted rate-limit wait, a
momentarily unreachable gateway, ...), retry once with a short recovery
note appended ONLY to that retry's OWN copy of ``messages`` — the caller's
``messages`` list is never mutated, so the note never leaks into the
real, persisted conversation. Exhausting the retry re-raises the original
exception unchanged, preserving existing failure handling (e.g. the
"model not found" restore-to-composer flow).
``on_reasoning`` is only forwarded when given — ``run_code`` never passed
it before this helper existed, and some lightweight test doubles for
``Provider`` don't accept the keyword at all; omitting it when unset
keeps every existing call site's exact prior calling convention."""
kwargs = {"tools": tools, "on_text": on_text, "cancel": cancel}
if on_reasoning is not None:
kwargs["on_reasoning"] = on_reasoning
try:
return provider.chat(messages, **kwargs)
except Exception:
if max_retries <= 0 or cancel():
raise
recovery = messages + [{"role": "user", "content": _RECOVERY_NOTE}]
return _call_provider_with_recovery(provider, recovery, tools, on_text, cancel,
on_reasoning, max_retries - 1)
def _parse_react(content: str) -> List[Dict[str, Any]]:
"""Extract @@TOOL fallback calls from assistant text."""
calls: List[Dict[str, Any]] = []
for i, match in enumerate(_TOOL_LINE.finditer(content or "")):
name = match.group(1)
try:
args = json.loads(match.group(2))
except json.JSONDecodeError:
continue
calls.append({"id": f"react_{i}", "name": name, "arguments": args})
return calls
def run_code(
provider: Provider,
messages: List[Dict[str, Any]],
ctx: ToolContext,
gate: PermissionGate,
emit: EmitFn,
cancel: Optional[CancelFn] = None,
max_steps: int = MAX_STEPS,
extra_tools: Optional[List] = None,
extra_executor=None,
skills_text: str = "",
rules_text: str = "",
project_context: str = "",
plan: bool = False,
security_config=None,
) -> List[Dict[str, Any]]:
"""``security_config`` is the app's ``AppConfig`` — enables the same
active guardrails as ``chat_agent.run_cowork`` (see its docstring).
``None`` (the default) disables both checks."""
cancel = cancel or (lambda: False)
extra_tools = extra_tools or []
extra_names = {t.name for t in extra_tools}
# update_plan drives the Plan panel (same as run_cowork) — always
# available, not opt-in via extra_tools, so unattended Schedule Tasks
# running through the code agent also get the completion self-check.
from .tools import enabled_tool_specs
all_tools = enabled_tool_specs(security_config) + extra_tools + [UPDATE_PLAN_SPEC]
# MS365 tools that send/write data (mail, Teams messages, OneDrive
# writes) are gated exactly like write_file/run_command — only the
# read/list ms365 tools count as "read-only, never confirm". Names are
# the MCP-qualified "ms365__*" form the agent sees (see ms365_tools.py).
gated_tools = WRITE_TOOLS | MS365_WRITE_TOOLS
# R05-T03/T04: ``gated_tools`` stays the authoritative name set (unchanged),
# but the actual confirm decision now goes through the same
# ToolPolicyGateway class run_cowork uses, instead of a separate
# hand-rolled ``if name in gated_tools`` + direct ``gate.request(...)``.
code_tool_policy = ToolPolicyGateway(
ToolRegistry(ToolDescriptor(n, "", {}, ToolCapability.WRITE) for n in gated_tools),
ToolCapability.WRITE,
)
# In PLAN mode, don't advertise write/run tools (analysis only).
advertised = [t for t in all_tools if t.name not in gated_tools] if plan else all_tools
has_memory = any(t.name.startswith("cmem_") for t in extra_tools)
has_plan_tool = True
has_ms365 = any(t.name.startswith("ms365_") for t in extra_tools)
if not messages or messages[0].get("role") != "system":
messages.insert(0, {"role": "system",
"content": code_system_prompt(
ctx.workdir, has_memory, plan, has_plan_tool, has_ms365)})
_apply_skills(messages, skills_text)
# The CODE agent uses RULEforCode.md (empty for now), NOT Cowork's
# RULEBASE.md — its safety comes from the sandbox until code rules exist.
from .security_rules import load_code_rules
_apply_security_rules(messages, rules_text or load_code_rules())
_apply_project_context(messages, project_context)
# Active guardrail — reviews the request itself and can refuse to proceed
# at all. Raises SecurityBlocked on a violation. agent_kind="code" → the
# code rulebase (RULEforCode.md), so RULEBASE.md's rules don't gate coding.
agent_security.enforce_prompt(provider, messages, security_config, emit, agent_kind="code")
for _ in range(max_steps):
if cancel():
break
def on_text(piece: str) -> None:
emit({"type": "text", "delta": piece})
assistant = _call_provider_with_recovery(provider, messages, advertised, on_text, cancel)
if not assistant.get("tool_calls"):
fallback = _parse_react(assistant.get("content", ""))
if fallback:
assistant["tool_calls"] = fallback
messages.append(assistant)
if not assistant.get("tool_calls") and not (assistant.get("content") or "").strip():
# Reasoning-only reply with no answer and no tool call — never leave a
# blank final turn (a Schedule Task run reads this back as its final
# answer, so a blank one silently produces "(no output)").
assistant["content"] = "*(model returned only its reasoning — try rephrasing)*"
emit({"type": "assistant_done", "content": assistant.get("content", "")})
tool_calls = assistant.get("tool_calls") or []
if not tool_calls:
break
for tc in tool_calls:
if cancel():
return messages
name, args, tc_id = tc["name"], tc.get("arguments", {}), tc["id"]
# update_plan drives the Plan panel only — no chat bubble, no file.
if name == "update_plan":
steps = normalize_plan_steps(args.get("steps"))
emit({"type": "plan_set", "steps": steps})
messages.append({"role": "tool", "tool_call_id": tc_id, "name": name,
"content": "Plan updated."})
continue
if plan and name in gated_tools:
emit({"type": "tool_result", "id": tc_id, "name": name, "ok": False,
"output": "PLAN mode: action skipped (switch to Act to execute)."})
messages.append({"role": "tool", "tool_call_id": tc_id, "name": name,
"content": "PLAN mode: not executed."})
continue
is_extra = name in extra_names
preview = ({"kind": "info", "title": name, "text": str(args)}
if is_extra else describe_action(ctx, name, args))
emit({"type": "tool_proposed", "id": tc_id, "name": name, "args": args, "preview": preview})
# Active guardrail on run_command/install_package, checked BEFORE
# asking the user to confirm — a no-op for every other tool or
# when disabled. Raises SecurityBlocked on a violation. Code agent
# → RULEforCode.md rulebase (not Cowork's RULEBASE.md).
agent_security.enforce_command(provider, name, args, security_config, emit,
agent_kind="code")
# read-only tools (incl. codebase memory) never consult the gate —
# code_tool_policy.requires_confirmation(name) is False for them.
approved = code_tool_policy.allow(
name, gate, {"id": tc_id, "name": name, "args": args, "preview": preview}
)
if cancel():
return messages
if not approved:
result = {"ok": False, "output": "User rejected the action."}
else:
emit({"type": "tool_start", "id": tc_id, "name": name})
if is_extra and extra_executor is not None:
if ctx.block_network and not name.startswith("ms365_local__"):
result = {"ok": False, "output": (
f"{name}: network access is blocked by the Sandbox Security Layer "
'("Block network for agent-run commands" is on in Settings).')}
else:
result = extra_executor(name, args)
else:
def on_output(line: str, _id=tc_id, _name=name) -> None:
emit({"type": "tool_output", "id": _id, "name": _name, "delta": line})
result = execute_tool(ctx, name, args, cancel=cancel, on_output=on_output,
agent_role=agent_roles.CODE)
evt = {
"type": "tool_result", "id": tc_id, "name": name,
"ok": result["ok"], "output": result["output"],
}
if isinstance(args, dict) and args.get("path"):
# absolute path so the UI can open the containing folder
evt["path"] = str(ctx.workdir / str(args["path"]))
if isinstance(result, dict) and result.get("produced"):
evt["produced"] = result["produced"]
emit(evt)
messages.append({
"role": "tool", "tool_call_id": tc_id, "name": name,
"content": result["output"],
})
return messages