Ba khối, mỗi người một khối: cột trái là thứ phải có trong tay mới làm được kèm
nguồn, cột phải là thứ bắt buộc giao ra kèm người nhận. Nhãn 'có rồi' đánh dấu
những gì mục chung đã giao xong hôm nay (SecretStore, ConfigRepository, fake,
script CASAN).
Kèm bảng output bắt buộc với cả ba mỗi PR, mỗi dòng có lệnh tự kiểm.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sáu việc trong "mục chung" của bản phân công, làm trước khi ba nhánh tính năng
tách ra.
1. Khung 5 tầng theo đúng đường dẫn plan.md: domain/ application/
infrastructure/ presentation/ platform/ + tests/fakes/ — 38 __init__.py.
Trước đó là 0 file, mà mọi task của cả ba người đều ghi vào đây.
Đã kiểm platform/ không che khuất module platform của stdlib.
2. Hợp đồng SecretStore và ConfigRepository (Protocol, chưa cài đặt) + fake
chạy trong bộ nhớ. Danh sách thuộc tính không bịa: đếm 156 lời gọi
ctx.config.* trong 29 file rồi lấy những cái dùng thật, xếp theo số lần.
Cố ý bỏ config.data (36 lời gọi, nhiều nhất) — bê dict thô sang kiến trúc
mới là bê nguyên vấn đề cũ.
3. tests/test_contracts.py — bài nghiệm thu, không phải test cho vui. Bài
chính chạy tiến trình riêng và khẳng định dùng fake KHÔNG kéo theo
cowork_local.config lẫn PySide6; đó là điều kiện để N2 và N3 code ngay hôm
nay thay vì đợi bản thật ngày 23 và 26/08.
4. scripts/audit_security.py — CASAN Check 1, Gamma chủ trì (hạn 30/08). Viết
sớm để kiểm liên tục trong lúc chuyển API key, không đợi tới ngày cổng.
Lần chạy đầu ra 3 báo động giả (secret_in_output là tên quy tắc, api_key="x"
là dữ liệu test) nên đã siết: ngưỡng độ dài, hằng liệt kê, hình dạng khoá
i18n, và dấu "# casan: allow" làm lối thoát chuẩn.
--self-test cắm 4 credential thật + 5 mẫu vô hại để chứng minh nó còn cắn
được — một máy quét không tìm thấy gì chỉ có giá trị nếu chứng minh được nó
biết tìm.
5. Ba check CASAN vào CI, chạy mọi PR thay vì dồn tới 30/08. Check 2 và 3
thuộc Team Hoa và Team Duy, chưa có script — bước CI bỏ qua nếu file chưa
tồn tại, để thêm cổng không làm đỏ CI của hai team kia.
6. docs/refactor/GammaTeam_decisions.md — hai quyết định chờ nhóm trưởng chốt:
provider_conf() còn trả api_key hay không (ảnh hưởng 5 nơi, 3 nằm ngoài
team), và số phận 24 checker UI sẽ vỡ khi file bị dời.
96 test xanh (90 cũ + 6 mới). CASAN Check 1: 0 credential lộ.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Trang HTML tự đứng một mình, mở bằng trình duyệt là xem được, không cần mạng.
Chia toàn bộ phần việc của team trong plan.md (R02, R07-T06, R08-T07…T10, R09,
CASAN Check 1) cho 3 người:
- Một mục chung nhóm trưởng làm trước, xong mới chia nhánh: dựng khung 5 thư
mục đích (hiện là 0 file), interface + fake cho Config/Secrets, chốt số phận
api_key, script CASAN Check 1, đưa 3 check vào CI, quyết số phận 24 checker
UI sẽ vỡ khi file bị dời.
- Ba nhánh tính năng ngang nhau, mỗi nhánh ~2.700 dòng: N1 cấu hình và vỏ ứng
dụng (nhóm trưởng giữ, vì chạm app.py / config.py / theme.py / i18n.py),
N2 giám sát, N3 Co4E.
- Bảy quy ước cho N2 và N3, ba trong đó là bắt buộc.
Số dòng code, 156 lời gọi ctx.config, 24 lời gọi audit_log.record và baseline
90 test đều đo trực tiếp trên main ngày 21/08, không lấy từ tài liệu.
Footer ghi rõ phần nào là đề xuất, phần nào lấy từ ba tài liệu gốc — mục chung,
cách chia nhánh, quy ước và nghiệm thu là đề xuất.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs/refactor/BaoCao_TeamDuy_R01_R03_R04.md records what was delivered against
each of the 16 tasks, the measured evidence (243 tests, 218 of them in 1.22s;
check_imports PASS; no production file over 400 LOC), the three real defects
found while working - the routing_application() deadlock, the swallowed
"notice" event, and the suite silently testing a different checkout - plus the
six open decisions and, explicitly, what was NOT tested (no manual app launch,
no real provider traffic, tools/check_*.py not run).
Refactoring_Checklist.md now links to it from the progress block.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
The previous commit recorded Team Duy as owning R01/R02/R04/R10. That is wrong.
Feature_Architecture_Proposal.md line 7 and DeltaTeam_prompt.md line 17 both
state R01, R03, R04, R08 (Chat UI) and R10; R02 belongs to Team Nam, which is
also who owns the two failing config-security tests.
The completed work itself (R01, R03, R04) was already correct and is unchanged.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Verification gap closed. The suite proved the new services correct in isolation,
but three paths I had modified had no test actually running them:
tests/integration/test_task_executor_flow.py (7 tests)
The Schedule Task path after R04-T05. Pins that History is still re-saved from
the LIVE message list mid-run (the reason begin_turn() exists - the pre-turn
copy would have frozen progress at the first user message), that update_plan
tracking still reports an unfinished checklist, and that a failed run still
raises so execute_task writes error.txt.
tests/integration/test_routing_surfaces.py (11 tests)
Real offscreen CoworkTab/Co4ETab/FolderTab calling the shared routing service:
correct surface key per screen, Auto switches, Off does not consult the engine,
Manual switches only on approval, a pinned Admin agent still wins, and AI-Edit
still pins TaskType.CODING. Also pins the field contract ui/routing_toggle.py
reads off RoutingDecision (from_model/to_model as provider/model keys) - a
rename there would only fail inside a modal dialog.
Also updates docs/refactor/Refactoring_Checklist.md: the 16 completed R01/R03/R04
tasks, the Team Duy daily rows, and a status block recording the measured
numbers, the scope correction (team owns R01/R02/R04/R10), and what is still
outstanding.
Suite: 243 passed, 2 pre-existing failures (EPIC R02). Fast suite (unit +
contracts + characterization + routing): 218 passed in 1.16s.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
EPIC R04 (Team Duy) - the turn lifecycle leaves the widget.
R04-T01 domain/agents/conversation_execution_request.py
Frozen snapshot of one turn, captured on the UI thread at submit time. The
job closure used to read widget/workspace state from inside the worker
thread, so a turn could run on a mix of submit-time and later state
depending on thread timing.
R04-T02 domain/agents/agent_event.py
13 frozen event types replacing untyped emit() dicts, with a two-way bridge
so existing widgets keep consuming the legacy shape until EPIC R08. Adds
TurnCompletedEvent - the end-of-turn signal the engine never had, which is
why a cancelled turn and a failed turn look identical to the UI today.
R04-T03 application/conversations/conversation_application_service.py
Runs a turn from a request and reports typed events. Never raises across the
worker boundary; TurnResult.raise_if_failed() preserves the existing
exception-based failure path. begin_turn()/execute_turn() expose the live
message list for callers that autosave history mid-run.
R04-T04 ui/cowork_tab.py::build_job -> snapshot + service.
R04-T05 core/task_executors.py::_run_agent -> same service (was a second,
slightly different assembly of the same call).
Caught while wiring the bridge: the first event vocabulary had no "notice"
event, so Agent Security warnings and auto-compaction notices would have been
silently swallowed. Added NoticeEvent plus a test that scans the engine sources
for emit() tags and fails when one has no typed counterpart.
New: tests/integration/ - real offscreen CoworkTab running a scripted turn end
to end (7 tests), including a characterisation of the extra provider call Agent
Security spends reviewing each request.
Suite: 225 passed, 2.74s. check_imports: PASS. All new files < 400 LOC.
2 pre-existing failures remain in test_config_security.py (EPIC R02/Team Nam).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
EPIC R03 (Team Duy) - one provider catalogue, one routing flow, one usage seam.
R03-T01 tests/contracts/test_providers.py
29 contract tests every provider must satisfy: canonical assistant message,
streamed text == returned content, reasoning never joins the answer, parsed
tool arguments, ProviderError for every failure. Real adapters exercised
offline by stubbing Provider._request.
R03-T02 domain/models/provider_descriptor.py
infrastructure/providers/provider_registry.py
Provider facts declared once (was split across providers/factory.py,
DEFAULT_CONFIG and PROVIDER_LABELS). ProviderRegistry.build() also stamps the
descriptor id onto the instance, so ollama/github_copilot/codex usage is no
longer all attributed to "openai_compat", and never mutates the caller config.
R03-T03 application/model_routing/routing_application_service.py
Pure-Python routing policy with four modes: Off, Auto, Manual and the new
Fallback (switch only AFTER the current model fails). Depends on a RoutingPort
protocol; production wires the existing core.routing engine underneath.
R03-T04/T05 ui/chat_panel.py, ui/co4e_tab.py, ui/folder_tab.py
Three near-identical routing copies (~40 lines each) replaced by a call to
ctx.routing_application() plus a confirm callback. Mode vocabulary now lives
in one place (normalize_mode/is_valid_mode) instead of four literal tuples.
R03-T06 infrastructure/telemetry/usage_sink.py
Token usage extracted from both providers into UsageEvent + UsageEventSink.
Estimation pinned against core.usage_tracker so no recorded number changes.
Also fixes a deadlock introduced while wiring AppContext: routing_application()
held _routing_lock and called routing(), which takes the same non-reentrant lock.
Suite: 186 passed, 1.22s. check_imports: PASS. All new files < 400 LOC.
2 pre-existing failures remain in test_config_security.py (EPIC R02/Team Nam).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
EPIC R01 (Team Duy) - safety net before the parallel refactor starts.
R01-T01 docs/architecture/ADR-001-layered-architecture.md
4-tier boundaries, allowed dependency directions, invariants I1-I6 and
the strangler-fig migration strategy.
R01-T02 tests/fakes/{fake_provider,fake_tool_executor}.py
Scripted, offline Provider and extra-tool executor doubles.
R01-T03 scripts/check_imports.py
AST-based Clean Architecture Guard (CASAN Check 3). Also covers relative
imports and function-local imports; ASCII-only output for cp932 consoles.
R01-T04 tests/characterization/test_run_cowork.py
13 snapshot tests pinning run_cowork's current observable contract before
EPIC R04 moves its orchestration into application/.
R01-T05 docs/architecture/dormant-code.md
Import-graph scan: 43 unimported modules verified down to 6 genuinely
dormant items (~1887 LOC); the rest run via subprocess/CLI entry points.
tests/conftest.py binds `cowork_local` to THIS checkout by absolute path -
previously sys.path discovery could import a sibling checkout and the suite
would silently test the wrong code.
Suite: 104 passed, 1.08s (2 pre-existing failures in test_config_security.py
remain - config.py still ships a hardcoded default password, EPIC R02/Team Nam).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copy of sections VI-IX from Feature_Architecture_Proposal.md
(roadmap, team assignment/KPI, anti-patterns, function migration map).
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>