Multi-agent workbench · desktop and web多 agent 工作台 · 桌面版與網頁版

Stage your coding agents. Gate every handoff.

讓 agent 接力,交接都要過關。

Claude Code, Codex, Grok, and Antigravity run through their own headless runtimes, and OpenRouter models run through Codex. An orchestrator agent gates each stage, your test command can check each patch, and the whole run stays replayable.

Claude Code、Codex、Grok、Antigravity 都透過各自的 headless runtime 執行,OpenRouter 的模型則借用 Codex 執行。每個階段由協調者 agent 把關,patch 可以交給你自己的測試指令檢查,整個執行過程都能回放。

MITTypeScript · Tauri 2Binds 127.0.0.1 by default預設只綁定 127.0.0.1
Run workspace執行工作區Running執行中
Implement → Test → Review
Pipeline preset · verify command: npm test內建 Pipeline 流程 · 驗證指令:npm test
✓Implement · verified實作 · 驗證通過claude · sonnet · worktree
✓Gate · advance關卡 · 繼續orchestrator · claude sonnet
…Test · running測試 · 執行中codex · gpt-5.6-sol · worktree
…Review · queued審查 · 排隊中claude · sonnet · safe
v0.2.10Latest release, July 2026最新版本,2026 年 7 月
5Providers, plus a test mock個 provider,另有測試用 mock
4Builtin workflows個內建工作流
12Agents per stage, at most每個階段最多的 agent 數
Why it exists為什麼要做

Side-by-side answers are not a workflow.

並排的答案,還不是工作流程。

Its predecessor, multi-ai-chat-desktop, orchestrated web chats. That is fine for comparing answers, but a chat window cannot edit your repository, run your tests, or record who decided what. Multi-AI Terminal (MAT) drives each agent's headless runtime inside your workspace, so every edit, check, and decision is kept as data.

前一代的 multi-ai-chat-desktop 串的是網頁聊天。拿來比較答案沒問題,但聊天視窗改不了你的 repo、跑不了你的測試,也記不下誰做了什麼決定。Multi-AI Terminal(MAT)直接驅動各家 agent 的 headless runtime,在你的工作區裡動手,每一次修改、檢查與決策都會存成資料。

A degraded or unverified advance is labeled, never hidden.
降級放行或未經驗證的結果,都會清楚標示,絕不藏起來。
How it works運作方式

Compose, run, gate, record.

組流程、跑 agent、過關卡、留紀錄。

A run applies a workflow to a workspace: ordered stages, each holding one or more agent slots, with an optional gate after each stage.

一次執行,就是把一個工作流套用到一個工作區:依序排好的階段,每個階段放一個以上的 agent 槽位,階段結束後可以設一道關卡。

01

Add a workspace

加入工作區

Point MAT at a project directory by absolute path, or pick it with Browse in the desktop app. Optionally set a verification command such as npm test.

用絕對路徑指定專案目錄,桌面版也可以用 Browse 選資料夾。可以順便設定驗證指令,例如 npm test。

02

Pick or shape a workflow

挑選或調整工作流

Start from Planning, Build, Review, or Pipeline. Under Customize, each slot sets provider, model, reasoning effort, permission tier, prompt template, and count.

從規劃、建置、審查或 Pipeline 四個內建流程開始。在「進階設定」裡,每個槽位可以指定 provider、模型、推理強度、權限層級、提示範本與數量。

03

Agents work, a gate decides

agent 動手,關卡決定

A stage's agents run in parallel, each optionally in its own git worktree. At a gated stage, an orchestrator agent reads a digest of the candidates and answers advance, retry, or abort in strict JSON.

同一階段的 agent 平行執行,也可以各自在獨立的 git worktree 裡動手。遇到有關卡的階段,協調者 agent 會讀取候選結果的摘要,用嚴格的 JSON 回覆繼續、重試或中止。

04

Read the record

回頭看紀錄

Conversation lists each answer, decision, check, and failure. Timeline replays the raw event log. Export a Markdown report or a debug bundle when you hand the work on.

「對話」列出每個回答、決策、檢查與失敗,「時間軸」回放原始事件紀錄。要交接時,可以匯出 Markdown 報告或除錯套件。

Then the gate decides

跑完再由關卡決定

The slots in one stage run together. One gate then answers for the whole stage.

同一個階段的槽位一起跑。再由一個關卡替整段回答。

  1. Stage

    A stage holds at most 12 slots, and the same provider waits 1.5 seconds between starts.

  2. Candidates

    The text stays, and a worktree run also keeps a binary patch.

  3. Gate

    The gate answers go on, try again, or stop.

  • slots started apart
  • answer kept
  • gate wants JSON
Before it answers, the gate reads a summary of that stage.
  1. 階段

    一個階段最多 12 個槽位,同一家 provider 要隔 1.5 秒才啟動下一個。

  2. 候選結果

    文字會留下,走 worktree 時再多一份二進位 patch。

  3. 關卡

    關卡的回答是繼續、再試一次,或停下來。

  • 啟動有間隔
  • 回答留下了
  • 關卡要 JSON
回答之前,關卡會先讀這一段的摘要。

The run lands in files

一次執行留下檔案

Environment values are hidden before anything is written down.

環境變數的值會先遮掉,再寫進檔案。

  1. Engine

    Events get appended from the stage engine while the run is still going.

  2. Data dir

    Both the web UI and the desktop app read and write here.

  3. Event log

    The timeline replays whatever landed in events.jsonl.

  4. Patches

    A binary patch from one attempt sits with that run.

  5. Report

    Generated, reviewed, advanced, and verified each get their own part of the Markdown.

  • events appended
  • patches stored
  • values hidden
The debug bundle is one zip of that same run, and the hidden values stay hidden inside it.
  1. 引擎

    執行還在跑的時候,階段引擎就把事件一筆筆寫進去。

  2. 資料目錄

    網頁和桌面版讀寫的是同一個地方。

  3. 事件紀錄

    events.jsonl 裡有什麼,時間軸就回放什麼。

  4. patch

    某次嘗試的二進位 patch,就放在這次執行旁邊。

  5. 報告

    產生、審查、放行、驗證,在 Markdown 裡各寫一段。

  • 事件已寫入
  • patch 留著
  • 變數值已遮
除錯套件是這次執行的一個 zip,遮掉的值在裡面也還是遮著的。
What you get你會得到什麼

Every step leaves evidence.

每一步都留下證據。

MAT treats agent output as evidence to check, not as an answer to trust.

MAT 把 agent 的輸出當成要查核的證據,而不是直接採信的答案。

01

A real runtime per provider

每家都用真正的 runtime

Claude runs through the Agent SDK and Codex through a persistent codex app-server. Grok and Antigravity run as headless CLIs behind FIFO managers, and OpenRouter models run through Codex with an isolated config.

Claude 透過 Agent SDK 執行,Codex 透過常駐的 codex app-server。Grok 與 Antigravity 以 headless CLI 執行,各自前面有一個 FIFO manager 負責排隊;OpenRouter 的模型則借用 Codex 執行,並使用獨立的設定。

server/src/providers/
02

An orchestrator with a budget

有額度上限的協調者

Any provider can orchestrate. A retry can target specific nodes and add a prompt addendum. If the answer cannot be parsed or the retry budget runs out (2 per stage by default), the stage advances and is marked degraded.

任何 provider 都能擔任協調者。重試可以只針對特定節點,並附上補充提示。回覆無法解析,或用完重試額度(預設每階段 2 次)時,階段會繼續往下,並標為降級。

server/src/orchestrator/
03

Worktrees and patches

Worktree 與 patch

Nodes can run in their own git worktree, and each attempt is captured as a binary patch. Apply one from the UI: MAT checks it with git apply first and lists conflicts instead of forcing it.

節點可以在獨立的 git worktree 裡執行,每次嘗試的修改都存成二進位 patch。從介面套用時,MAT 會先用 git apply 檢查,有衝突就列出來,不會硬套。

server/src/engine/worktree.ts
04

Verification you define

驗證指令由你決定

Each workspace can set a command such as npm test, with a 600-second default timeout. Worktree candidates with a non-empty patch run it, and a stage with requireVerified retries failed checks when nothing passed.

每個工作區可以設定一個驗證指令,例如 npm test,預設逾時 600 秒。若候選結果使用 worktree 且 patch 不是空的,就會執行這個指令;開啟 requireVerified 的階段,如果沒有任何候選通過,就會重試沒過的那些。

server/src/engine/verify.ts
05

Steering mid-run

執行中途追加指示

Interrupt stops the active candidates, keeps their partial logs and patches, and runs your instruction. Queue waits for the next stage boundary. Up to 8 per run, first in, first out, never typed into a running process.

「立即插入」會停下正在執行的候選,保留已產生的日誌與 patch,再執行你的新指示;「排隊」則等到下一個階段交界。每次執行最多 8 則,先進先出,絕不會寫進執行中行程的 stdin。

server/src/engine/steer.ts
06

Reports and debug bundles

報告與除錯套件

The Markdown report separates generated, reviewed, advanced, and verified work, with CLI versions, usage, patches, and checks. The debug bundle is one zip of the whole run, with environment variable values redacted.

Markdown 報告會分開標示已產生、已審查、已放行與已驗證的工作,並附上 CLI 版本、用量、patch 與驗證結果。除錯套件把整次執行打包成一個 zip,環境變數的值都會遮蔽。

GET /api/runs/:id/report

Five runtimes, one slot

一個槽位一種跑法

You fill the same kind of slot no matter which provider you pick.

不管選哪一家,槽位要填的東西都一樣。

  1. Workflow slot

    You pick a provider, a model, an effort, and one of safe, auto, or full.

  2. Runtime

    Claude runs in its agent library, Codex keeps a server process, and grok and agy are command line tools with no window.

  • provider picked
  • runtime follows
  • same slot for all
OpenRouter models run through Codex, with a config of their own.
  1. 流程槽位

    你選 provider、模型、推理強度,再從 safe、auto、full 裡挑一個。

  2. 執行期

    Claude 走程式庫,Codex 開一個伺服器行程,grok 和 agy 用不開視窗的指令列。

  • 選好一家
  • 跑法跟著走
  • 槽位都一樣
OpenRouter 的模型透過 Codex 跑,用自己的一份設定檔。

You check the patch

用你的指令看 patch

Only a worktree candidate with a real patch gets this command.

只有 worktree 候選,而且 patch 真的有內容,才會跑這道指令。

  1. Patch

    Nothing to check means the command never starts.

  2. Your command

    A command like npm test stops after 600 seconds, unless the workspace sets another limit.

  3. Result

    The log keeps whichever of passed, failed, error, or skipped came back.

  • empty patch skipped
  • stops at 600 seconds
  • result saved
When a stage requires a passing check and no candidate passes, the failed ones are tried again within that stage's retry budget.
  1. patch

    沒有內容可看,指令就不會啟動。

  2. 你的指令

    npm test 這類指令,工作區沒另設的話,600 秒就停。

  3. 結果

    日誌會記下回來的是通過、失敗、錯誤還是略過。

  • 空 patch 跳過
  • 600 秒會停
  • 結果存好了
階段要求驗證通過、卻沒有候選通過時,沒過的候選會用這個階段的重試額度再試。
Architecture系統架構

One local server, many runtimes.

一個本機伺服器,多種 runtime。

The browser UI and the desktop shell talk to the same Node server. It owns runs, gates, and storage, and starts and stops each provider's runtime as a child process or SDK session.

瀏覽器介面和桌面版都連到同一個 Node 伺服器。執行、關卡與儲存都由它負責,各家 provider 的 runtime 也由它以子行程或 SDK session 的形式啟動與停止。

REST + WebSocket → engine → agents → diskREST + WebSocket → 引擎 → agent → 磁碟
Web UI網頁介面React + ViteConversation對話Timeline時間軸Health健康狀態
Server伺服器Fastify + WSStage engine階段引擎Gates關卡Verification驗證
AgentsAgentclaudecodexgrok · agyopenrouter
Data資料run.jsonevents.jsonlpatches, logspatch、日誌
Data lives in ~/.multi-ai-terminal. The Tauri 2 desktop app runs the same server on a random loopback port.資料放在 ~/.multi-ai-terminal。Tauri 2 桌面版跑的是同一個伺服器,只是改用隨機的 loopback 連接埠。

Stays on this machine

預設只聽這台機器

Agent processes start from argument arrays, not from a shell string.

agent 行程用參數陣列啟動,不是一整串 shell。

  1. Default bind

    From source it listens only on this computer, port 7788, and the desktop app picks a free port there too.

  2. Wider host

    A wider bind leaves the token optional, so anyone who reaches the port can run agents.

  • this computer by default
  • token is optional
  • verify uses the shell
  • This computer
  • Your bind
  • Wider host
If you open it beyond this computer, set a token yourself.
  1. 預設綁定

    從原始碼跑只聽這台電腦的 7788,桌面版也在這台電腦上挑一個空的連接埠。

  2. 對外位址

    綁到所有介面時 token 仍可留空,連得到連接埠的人就能跑 agent。

  • 預設只聽本機
  • token 不是必填
  • 只有驗證走 shell
  • 本機
  • 你怎麼綁
  • 對外
如果開放到這台電腦以外,要自己設 token。
Design decisions設計取捨

Deliberate trade-offs.

刻意做的取捨。

01

Headless runtimes, not web chats

用 headless runtime,不用網頁聊天

Driving each vendor's CLI, app-server, or SDK gives MAT tool events, usage, patches, and exit codes it can normalize into one schema. The cost: every provider needs its own install and sign-in, and the Grok and agy streams carry less detail.

直接驅動各家的 CLI、app-server 或 SDK,MAT 才拿得到工具事件、用量、patch 與結束代碼,並整理成同一套格式。代價是每個 provider 都得各自安裝、登入,而 Grok 與 agy 的串流細節也比較少。

02

Advance degraded, never hang

寧可降級放行,也不卡住

If the orchestrator's answer cannot be parsed, or a stage spends its retry budget, the run moves on and the decision is marked degraded in the UI and the report. A run that finishes with a visible caveat beats one that never finishes.

協調者的回覆無法解析,或某個階段用完重試額度時,執行會繼續往下,這個決策會在介面與報告裡標為降級。帶著明確註記跑完,總比永遠跑不完好。

03

Space out same-provider launches

同一家 provider 錯開啟動

Parallel Codex sessions on one OAuth login can race the single-use refresh token. MAT starts sessions of the same provider at least 1.5 seconds apart, orchestrator included. That narrows the race without removing it, so API keys or serial use stay the durable fix.

同一個 OAuth 登入下平行跑多個 Codex session,可能會搶著輪替只能用一次的 refresh token。MAT 讓同一家 provider 的啟動至少間隔 1.5 秒,協調者也算在內。這只能減少互搶,無法根除,長久的解法還是改用 API key 或依序使用。

04

Follow BAT for runtime handling

runtime 處理跟著 BAT 走

Provider sessions follow Better Agent Terminal: a persistent codex app-server controller and Claude Agent SDK sessions. MAT is not a fork. It ports that pattern to a Node-only server and applies it to grok and agy as well.

provider session 的處理方式參考 Better Agent Terminal:常駐的 codex app-server controller,加上 Claude Agent SDK session。MAT 不是 fork,而是把這套做法移植到純 Node 的伺服器,也套用到 grok 與 agy。

Retry, then move on

重試完還是往下

The budget is 2 retries a stage. When it's gone, the stage advances and is marked degraded.

每個階段預設能重試 2 次。用完還是往下,並標成降級。

  1. Stage result

    You get the candidate's text back, sometimes with a patch or a verification result.

  2. Retry budget

    An answer that can't be parsed advances immediately and is marked degraded.

  • JSON or degraded
  • two retries
  • retry names no slot, moves on
Choosing to stop ends the whole run.
  1. 階段結果

    候選會交回文字,有時還附 patch 或驗證結果。

  2. 重試額度

    回答解析不了,就立刻放行,並標成降級。

  • 有 JSON 或降級
  • 預設兩次
  • 沒指名槽位也往下
選擇停下來,整次執行就結束。
Start here從這裡開始

Desktop app, or from source.

桌面版,或原始碼。

Both need Node.js 20 or newer, Git (2.32+ recommended), and sign-in or an API key for each provider you use. Claude and Codex can use runtimes MAT downloads for you; Grok needs grok, Antigravity needs agy, and OpenRouter needs OPENROUTER_API_KEY.

兩種方式都需要 Node.js 20 以上、Git(建議 2.32 以上),以及你要用的每家 provider 的登入或 API key。Claude 與 Codex 可以使用 MAT 代為下載的 runtime;Grok 需要 grok,Antigravity 需要 agy,OpenRouter 則需要 OPENROUTER_API_KEY。

Desktop app

桌面版

Windows: the setup .exe or the .msi. macOS: the .dmg for Apple silicon or Intel; it is not signed or notarized, so allow the first launch under System Settings, Privacy & Security. Linux: .deb, .AppImage, or .rpm.

Windows 用 setup .exe 或 .msi。macOS 依晶片選 Apple silicon 或 Intel 的 .dmg;它沒有簽章也沒有經過公證,第一次開啟要到「系統設定」的「隱私權與安全性」允許。Linux 可選 .deb、.AppImage 或 .rpm。

From source

從原始碼執行

Clone the repository, then build and start it. The UI and API come up on http://127.0.0.1:7788. Port, host, data directory, and token can each be set with a flag or an environment variable.

先 clone 這個 repo,再建置並啟動。網頁介面與 API 會開在 http://127.0.0.1:7788。連接埠、位址、資料目錄與 token 都可以用參數或環境變數設定。

npm install
npm run build
npm start

First run

第一次執行

In Projects, add a workspace. Back in Launch, pick Planning, Build, Review, or Pipeline, describe the task, and press Start. Open Customize only to change stages or agents.

在「專案」加入工作區,回到「啟動」選規劃、建置、審查或 Pipeline,寫下任務後按「開始執行」。要改階段或 agent 時,才需要打開「進階設定」。

Honest status誠實的現況

Where it stands.

目前的樣子。

Built in six days: 21 releases, from v0.1.0 on July 18, 2026 to v0.2.10 on July 23. No feature changes since. An October 2026 cleanup removed personal notes, fixed the docs, and updated dependencies to close every open Dependabot alert; those updates are on main and not yet in a release.

六天內做出來:從 2026 年 7 月 18 日的 v0.1.0 到 7 月 23 日的 v0.2.10,共 21 個版本。之後沒有新功能。2026 年 10 月整理過一次,移除個人筆記、修正文件,並更新相依套件,關掉所有未處理的 Dependabot 警示;這些更新目前只在 main 上,還沒有包進新版本。

Works today

目前可用

  • Desktop installers for Windows, macOS (Apple silicon and Intel), and Linux (.deb, .AppImage, .rpm)
  • Windows、macOS(Apple silicon 與 Intel)、Linux(.deb、.AppImage、.rpm)都有桌面版安裝檔
  • Five providers (claude, codex, grok, agy, openrouter) plus a deterministic mock for tests
  • 五個 provider(claude、codex、grok、agy、openrouter),另有測試用、結果固定的 mock
  • Custom stages, worktree isolation, patch apply, verification, steering, reports, and debug bundles
  • 自訂階段、worktree 隔離、套用 patch、驗證、追加指示、報告與除錯套件
  • AI-Sister, Midnight, and Daylight themes, in English or Traditional Chinese
  • AI-Sister 紀念版、午夜深色、日光淺色三種主題,介面可切換英文或繁體中文
  • CI passed on Ubuntu and Windows for the release commit: npm test (620 passing on Ubuntu), agent-lane tests, typecheck, a browser smoke test, and 5 evidence scenarios on Linux
  • release 的 commit 在 Ubuntu 與 Windows 的 CI 都通過:npm test(Ubuntu 上 620 項通過)、agent 通道測試、型別檢查、瀏覽器冒煙測試,Linux 另外跑 5 個證據情境

Limits and not yet

限制與尚未完成

  • Grok's stream has no tool events, so grok nodes show thinking and text only
  • Grok 的串流沒有工具事件,grok 節點只看得到思考與文字
  • agy has no JSON mode or session resume: its stream is plain text, and as orchestrator it is re-briefed at every gate
  • agy 沒有 JSON 模式,也不能接續 session:串流只有純文字,擔任協調者時每道關卡都要重新交代背景
  • Desktop builds are not code-signed or notarized, and still need Node.js 20+ on PATH (or MAT_NODE)
  • 桌面版沒有程式碼簽章,也沒有經過 Apple 公證,而且 PATH 上仍需要 Node.js 20 以上(或設定 MAT_NODE)
  • CI and the evidence suite use the mock provider; runs with real signed-in Codex, Claude, or OpenRouter accounts are not part of CI
  • CI 與證據測試用的是 mock provider,實際登入 Codex、Claude 或 OpenRouter 帳號的執行不在 CI 範圍內
  • Parallel sessions on one Codex OAuth login can still race refresh-token rotation; API keys or serial use avoid it
  • 同一個 Codex OAuth 登入下的平行 session,仍可能搶著輪替 refresh token;改用 API key 或依序使用就能避開
  • Do not run a source build while the desktop app is open: both use the same data directory and race its stores
  • 桌面版開著時,不要再從原始碼啟動一份:兩者共用同一個資料目錄,會互相搶寫
  • After a reboot, crash recovery kills stale process groups by saved PID and accepts the risk of PID reuse
  • 機器重開後,崩潰復原會依記錄的 PID 結束殘留的行程群組,並接受 PID 被重複使用的風險
Version版本
v0.2.10
License授權
MIT, theme artwork excludedMIT,主題角色圖除外
Stack技術
TypeScript · Fastify · React · Tauri 2
Platform平台
Windows · macOS · Linux
Requires需求
Node.js 20+ and GitNode.js 20+ 與 Git
Verified查核日期
2026-10-06

More from Ted Huang. Every public project has a page like this one, in English and Traditional Chinese.

Ted Huang 的其他作品。每個公開專案都有一頁像這樣的中英雙語介紹。

All projects →全部專案 →