Agent Workbench: Problems and Reflections from Building an Agent Project
An AI agent workbench built with Bun and TypeScript, supporting multiple models, multiple transports, and a pluggable tool system.
1. Why Build It?
After reading plenty of documentation and examples on AI application development, I had gained a fairly broad understanding of how agents work—from ReAct to Toolformer, and from MCP to Agent Skills. But there is a huge gap between understanding what you read and knowing how to build it.
The best way to learn is to build one yourself.
So I decided to start from scratch, beginning with the protocol layer and gradually building my own agent workbench. The aim wasn't to replace Cursor or Claude Code, but to understand the engineering decisions beneath an agent system by building each module myself: protocol design, task scheduling, tool registration, and event streams.
Before starting, I also studied several existing agent frameworks. I'll use these excellent projects as references while building my own general-purpose agent.
The project's core design goals are:
Model independence: no ties to a particular AI vendor; switch freely between OpenAI, Anthropic, DeepSeek, and others through configuration.
Transport independence: the kernel doesn't care where messages come from. Stdio, Electron IPC, and a possible future WebSocket transport can all connect.
Pluggable tools: built-in filesystem, shell, and task-management tools, plus support for external tools through MCP (Model Context Protocol).
Multi-agent collaboration: the main agent decides whether to delegate subtasks. It handles simple tasks itself and distributes complex ones, scheduling flexibly and parallelizing when needed.
Cascading configuration: global → workspace → session, with three layers of priority-based overrides for flexibility and control.
Platform independence: abstracting the UI and communication protocol (JSON-RPC) lets the frontend run on any platform supporting web technologies, including Electron and browsers. Core logic connects through the same protocol contract. Inspired by projects such as OpenClaw, this makes the kernel a real backend service rather than an embedded component tied to a particular shell.
2. Technology Choices
Technology | Why I Chose It |
|---|---|
Bun 1.3 | Runtime, package manager, and test runner in one, with native TypeScript support and much faster startup than Node.js. |
TypeScript (strict) | End-to-end type safety, with consistent type constraints from the protocol layer to the UI. |
Vercel AI SDK | An abstraction over AI providers, with a unified streaming API supporting OpenAI, Anthropic, and DeepSeek. |
Electron + React 19 | The desktop solution: React in the renderer, with core logic running in a Bun subprocess launched by the main process. |
Drizzle ORM + Bun SQLite | Lightweight persistence without an extra database process; a straightforward, maintainable schema. |
Zod | Runtime validation, with the protocol and kernel sharing the same schema definitions. |
Zustand | Minimal state management without nested providers or boilerplate. |
Biome | Unified formatting and linting, replacing ESLint and Prettier. |
Tailwind CSS v4 | Atomic styles, with CSS custom properties for theming. |
Why Bun Rather Than Node.js?
Bun's isolated linker avoids traditional Node "node_modules hell," and workspace packages in the monorepo can be referenced immediately without a build. Its test runner also offers excellent Jest compatibility and execution speed. Most importantly, in the Electron desktop app, Bun runs the core logic as a subprocess with almost no startup delay.
Why Vercel AI SDK Instead of Calling APIs Directly?
Vercel AI SDK has been widely validated in Anthropic's Agent Skills documentation and community practice. Its unified streamText / generateText interfaces and ToolLoopAgent abstraction avoid writing repetitive adapters for every model provider. I only use it for underlying model calls and tool loops, though; routing, configuration, and task scheduling remain under my own control.
3. Project Design
3.1 Architecture Evolution: From Blindly Pursuing Multiple Agents to a Configurable Agent System
Task execution architecture is where I spent the most time thinking—and encountered the most pitfalls.
First Pitfall: Blindly Pursuing Multiple Agents
Like many people, I initially thought "multi-agent collaboration" sounded impressive: the main agent plans, subagents execute, and everyone has a role. So I built the typical Main Agent + SubAgent arrangement:
@startuml
!theme plain
rectangle "Main Agent\n(Planner)" as main #LightBlue
rectangle "SubAgent A\n(Coder)" as suba #LightGreen
rectangle "SubAgent B\n(Tester)" as subb #LightGreen
rectangle "SubAgent C\n(Reviewer)" as subc #LightGreen
main --> suba : Pass task
suba --> main : Return result
main --> subb : Pass code
subb --> main : Return tests
main --> subc : Pass code and tests
subc --> main : Return review
@endumlDiagram unavailable. Use Show code to inspect the source.
Once it ran, the problem quickly became obvious: a huge number of tokens went into synchronizing context between agents. Subagent A wrote code and returned it to the main agent, which passed it to subagent B for tests. After B finished, the main agent integrated the result and sent it to subagent C for review. Every handoff consumed substantial tokens. Information also degraded at each step of this relay, and the final results were disappointing.
Building multi-agent systems describes this problem well: dividing agents by task type—planner, implementer, tester, reviewer—is the most common mistake. Such pipelines lose context at every handoff, and subagents spend more tokens coordinating than doing useful work. The article gives a direct estimate: multi-agent systems typically consume 3–10 times as many tokens as a single agent. The better approach is context-centric decomposition. An agent that already understands a feature's implementation should also write its tests, rather than handing them to another agent unfamiliar with the details. Multiple agents are worth introducing only when their contexts can truly be isolated, such as parallel searches across different sources or changes to independent modules with clear interfaces.
Returning to One Agent: Much Better Quality, but New Problems
After that experience, I returned to the simplest approach: one agent + tools + skills. The results actually improved dramatically. Without the loss from passing context around, the agent could complete the whole task coherently.
But two new problems emerged:
1. Homogeneous Output
Within one session, the agent has to write code, tests, and documentation. Without a role-switching mechanism, these outputs converge in style and reasoning. This isn't beneficial consistency; it is mental inertia. Tests validate the code using the same assumptions that produced it, while documentation tells the story from the same angle. Independent perspectives are missing.
2. Limits of a Single Agent
One agent clearly struggles in some situations:
Large-scale search: when searches need to span multiple repositories or documentation sources in parallel, a single agent searches sequentially, with time growing linearly with scope.
Parallel changes across modules: when several independent modules need changing, one agent cannot advance them in parallel and must switch context repeatedly.
Degraded tool selection with too many tools: once the tool count grows beyond about 15–20, the model starts selecting tools incorrectly. Adding tools for one domain can even hurt performance in another.
These problems correspond to the three scenarios where Anthropic's multi-agent article says multiple agents are useful: parallelization for large searches and independent modules, context isolation to avoid mental inertia, and specialization to separate tools by responsibility and prevent overload.
The Final Approach: A Configurable Agent System
I didn't hard-code a "main agent → split task → delegate to subagents" workflow. Instead, I designed a mechanism that makes each agent a configurable unit:
@startuml
!theme plain
title Configurable Agent System
rectangle "Agent configuration" as config #LightYellow {
rectangle "Name & description" as name
rectangle "System Prompt\n(Role and responsibilities)" as prompt
rectangle "Tool set\n(Defines the agent's capabilities)" as tools
rectangle "Model selection" as model
}
rectangle "Orchestrator Agent" as orch #LightBlue {
rectangle "Understand tasks and create plans" as plan
rectangle "Has the delegate_task tool" as dtool
rectangle "Decide whether to act or delegate" as decide
}
rectangle "Explorer Agent" as explorer #LightGreen {
rectangle "Search code and documentation" as search
rectangle "Read-only tools" as rotools
}
rectangle "Worker Agent" as worker #LightGreen {
rectangle "Edit code and write tests" as modify
rectangle "Read/write tools" as rwtools
}
config --> orch : Configure
config --> explorer : Configure
config --> worker : Configure
orch --> explorer : delegate_task\n(Pass only the search objective)
orch --> worker : delegate_task\n(Pass only requirements and relevant files)
note bottom of config : Agent capabilities are defined entirely by configuration\nThe delegate_task tool determines whether subagents are available\nNo code changes required
@endumlDiagram unavailable. Use Show code to inspect the source.
The central idea is simple: whether an agent can launch subagents depends on whether it has the delegate_task tool. Give an Orchestrator this tool and an appropriate system prompt. Omit it for a Worker that only executes tasks. Every agent type—Orchestrator, Explorer, Worker—is defined through configuration rather than hard-coded.
The flexibility shows up in several ways:
New agent types require no code changes: want a Code Reviewer Agent? Configure a system prompt and read-only tools. No implementation changes are needed.
Agents are decoupled from models: an Orchestrator can use Opus for complex decisions, a Worker can use Sonnet for efficient execution, and an Explorer can use Haiku for quick searches. Choose models by responsibility and allocate them as needed.
Tools define capability boundaries: what an agent can and cannot do is determined entirely by its tools. delegate_task is just another tool, not fundamentally different from the others.
Simply put: when one agent is enough, it works efficiently on its own. When multiple agents are needed, configure an Orchestrator with delegate_task, and it decides when to split the task and whom to delegate to.3.2 Design Philosophy: Protocol First
The whole project follows a protocol-first philosophy. Before writing any kernel code, define the communication contracts between layers.
The center of that contract is packages/protocol, which defines:
15 JSON-RPC methods covering configuration (config.get/set), workspace CRUD (workspace.create/list/set/delete), sessions (session.create/delete/list/set/detail), task control (task.create/cancel), and system capabilities (kernel.capabilities).
6 server-pushed events for configuration changes (app.config.changed) and the task lifecycle (task.started/delta/succeeded/failed/cancelled).
Complete Zod validation: strict runtime checks for every request's inputs and outputs, and for event data.
With the protocol as the single source of truth, the SDK and kernel can evolve independently. As long as the protocol remains compatible, changes on either side don't affect the other. This also makes the kernel transport-agnostic: it processes protocol-compliant messages without caring whether they arrive through stdio or WebSocket.
3.3 Cascading Configuration: Three Layers of Overrides
A core issue for AI agents is context configuration. LangChain's Context Engineering for Agents emphasizes that context includes not only prompt text but also runtime settings such as model choice, tools, and skills.
I designed three layers of cascading configuration:
@startuml
!theme plain
title Configuration Resolution Priority
rectangle "Global Config\nglobal model & defaults" as global #LightBlue
rectangle "Workspace Config\nper-project override" as workspace #LightGreen
rectangle "Session Config\nper-conversation tweak" as session #LightYellow
rectangle "Agent Presets\nfallback defaults" as agent #LightGray
global -[hidden]down-> workspace
workspace -[hidden]down-> session
session -[hidden]down-> agent
note right of global : modelId, agentId\nskillPaths, toolTitles, mcpNames
global --> workspace : unresolved fields fall through
workspace --> session : unresolved fields fall through
session --> agent : unresolved fields fall through
@endumlDiagram unavailable. Use Show code to inspect the source.
Each layer is a RuntimeConfig defining modelId, agentId, skillPaths, toolTitles, and mcpNames. Fields absent from a lower layer fall back to the layer above, ultimately reaching the agent's preset defaults. This allows:
Global defaults for models and tools.
Workspace overrides for specific projects, such as pinning one project to Claude Opus.
Session overrides that temporarily enable or disable tools or skills.
Looking back at the three elements in Lilian Weng's LLM Powered Autonomous Agents—planning, memory, and tool use—my configuration system essentially makes those elements configurable at runtime.
3.4 Reducing Boilerplate: A Minimal IoC Container
In practice, modules frequently need shared context objects: loggers, configuration, database connections, and so on. Passing them manually adds parameters to every function and quickly creates boilerplate across layers. I didn't want a heavyweight DI framework such as InversifyJS or TSyringe for this small project, so I wrote an IoC container of roughly 50 lines using AsyncLocalStorage:
// 创建 Token(一个具名 Symbol)
const LoggerToken = token<ILogger>("Logger");
// 注册提供者(懒加载工厂函数)
register(LoggerToken, () => createLogger(options));
// 注入依赖(在容器上下文中调用时自动解析)
const logger = inject(LoggerToken);
// 惰性注入——解决循环依赖
const router = inject.lazy(RouterToken);AsyncLocalStorage automatically carries a container context with each request. Callers don't pass it explicitly: inject(LoggerToken) in any function retrieves the current request's Logger, and separate requests don't interfere with each other's contexts. inject.lazy() uses a Proxy to resolve a dependency only on first property access, addressing circular dependencies between providers.
3.5 The Tool System: Built-In Tools and MCP
Tools are an agent's hands. The project includes 14 tools—reading, writing, editing, and deleting files; directory listing; glob and text search; patch application; shell execution; and task creation, updates, listing, and delegation—covering the core needs of a coding agent.
External tools also connect through the standardized Model Context Protocol (MCP). McpManager supports three transports:
HTTP: remote MCP services.
SSE: Server-Sent Events.
Stdio: local process communication, such as connecting a local LSP server.
4. Overall Architecture
4.1 Process Model
@startuml
!theme plain
title Process Model
package "Electron Shell" #LightGray {
rectangle "Renderer\n(React UI + ClientSdk)" as ui #LightYellow
rectangle "Main Process\n(IPC Bridge)" as main #LightGreen
ui <-> main : Electron IPC
}
package "Host Process (Bun)" #White {
rectangle "HostRuntime\n(Request validation + Event broadcasting)" as host #LightBlue
rectangle "Kernel (IoC)\n(Routing + Business logic)" as kernel #LightPink
database "SQLite" as db #LightGray
host -> kernel : handleRequest()
kernel -> db : Drizzle ORM
}
main -> host : stdio\n(JSON-RPC)
@endumlDiagram unavailable. Use Show code to inspect the source.
Key design decisions:
The kernel is a pure logic layer: it contains no transport-specific code. It accepts Request objects, returns Response objects, and broadcasts Event objects through callbacks. That lets it run unchanged inside a stdio process, an Electron main process, or a possible future HTTP server.
The host is the process boundary: it starts the transport, loads the kernel, and validates requests. As the kernel's runtime environment, it owns all process-related responsibilities.
The desktop shell is the UI layer: the Electron renderer handles presentation, while the main process bridges IPC and stdio, forwarding renderer requests to the host process.
4.2 Data Flow
A typical conversation request follows this path:
@startuml
!theme plain
title Conversation Request Data Flow
actor User
participant "React UI\n(ClientSdk)" as ui
participant "Main Process\n(IPC Bridge)" as ipc
participant "HostRuntime" as host
participant "Kernel\n(Router + Handler)" as kernel
participant "TaskExecutor" as executor
User -> ui : Enter a message
ui -> ipc : ElectronIpcTransport\nValidate with req.safeParse()
ipc -> host : StdioJsonRpcTranslator\nJSON + newline
host -> kernel : ReqSchema.safeParse()\nhandleRequest()
kernel -> executor : createTask()
executor -> executor : Stream agent execution\n(streamText + tool loop)
executor --> kernel : emit(task.delta/succeeded...)
kernel --> host : emit(event)
host --> ipc : broadcast()
ipc --> ui : ipcRenderer.on()
ui -> ui : Reduce Zustand state\nRender UI
@endumlDiagram unavailable. Use Show code to inspect the source.
Each step has a clear responsibility, and layers connect through type-safe contracts.
4.3 Context Engineering in Practice
Prompt templates: prompt fields in runtime configuration can be customized by layer. The agent's system prompt is dynamically assembled from the prompt, skill list, and tool list.
Message context management: ContextProvider collects messages from peer tasks in a session and sorts them chronologically into a complete conversation history.
Tool results become context: tool inputs and outputs are stored as structured function_call parts, allowing the agent to review its own action history.
5. Project Structure
agent-workbench/
├── apps/
│ ├── desktop/ # Electron 桌面壳
│ │ ├── src/
│ │ │ ├── main/ # Electron 主进程
│ │ │ │ ├── index.ts # 窗口创建、IPC 注册、Host 启动
│ │ │ │ ├── core-process.ts # Bun 子进程管理
│ │ │ │ ├── ipc.ts # IPC 通道注册
│ │ │ │ └── window-options.ts # 窗口配置(透明、毛玻璃等)
│ │ │ ├── preload/ # 预加载脚本(contextBridge)
│ │ │ ├── renderer/ # React 渲染进程
│ │ │ └── shared/ # IPC 契约常量
│ │ └── electron.vite.config.ts
│ │
│ └── host/ # Host 运行时进程
│ └── src/
│ ├── index.ts # 入口:加载配置、启动 Runtime
│ ├── config.ts # 传输层配置解析
│ ├── runtime.ts # HostRuntime:内核 + 传输层编排
│ └── transports/ # 传输层实现(Stdio)
│
├── packages/
│ ├── kernel/ # 内核(核心业务逻辑)
│ │ └── src/
│ │ ├── kernel/ # IoC 容器、bootstrap
│ │ └── application/
│ │ ├── router/ # JSON-RPC 路由(15 个方法)
│ │ ├── store/ # 持久化层(Drizzle SQLite)
│ │ ├── provider/ # 依赖提供者(模型、Agent、配置等)
│ │ ├── manager/ # 管理器(MCP、Skill、Tool)
│ │ ├── engine/ # 任务引擎(创建、排队、取消)
│ │ └── task/ # 任务执行器(streaming、事件广播)
│ │
│ ├── protocol/ # 共享协议契约
│ │ └── src/
│ │ ├── schemas/ # Zod Schema + 常量定义
│ │ └── types/ # 推断的 TypeScript 类型
│ │
│ ├── sdk/ # 客户端 SDK
│ │ └── src/
│ │ ├── client.ts # ClientSdk:请求/响应/事件管理
│ │ └── transports/ # Stdio + Electron IPC 传输层
│ │
│ ├── ui/ # React UI 组件库
│ │ └── src/
│ │ ├── App.tsx # 根组件 + 路由
│ │ ├── stores/ # Zustand 状态管理
│ │ ├── components/ # UI 组件(工作台、设置、原语)
│ │ ├── pages/ # 页面(工作台页、设置页)
│ │ └── lib/ # 工具函数
│ │
│ └── toolkit/ # 共享工具包
│ └── src/
│ ├── env/ # 环境变量校验(Zod)
│ └── logger/ # 结构化日志
│
├── data/ # 运行时数据(SQLite 数据库)
├── biome.json # 格式化/Lint 配置
├── tsconfig.base.json # TS 严格模式配置
├── bunfig.toml # Bun 配置
└── package.json # Monorepo 根6. Core Modules in Detail
6.1 Kernel
This is the system's brain, with bootstrap() as its main entry point:
export const bootstrap = (options: KernelOptions): IAgentKernel => {
const container = new IocContainer();
container.set(EmitToken, emit);
container.set(EnvToken, loadEnv(KernelEnvSchema, envSource));
runWithContainer(container, () => {
registerProviders(); // 注册 Provider
container.resolveAll(); // 懒加载 → 预加载,提前发现初始化错误
registerRoutes(); // 注册路由
});
return {
handleRequest: async (request) =>
runWithContainer(container, async () => {
return await router.handle(request);
}),
};
};bootstrap() exposes only handleRequest. Routing, business execution, and error handling happen transparently inside it. This pipeline design makes the kernel's implementation a complete black box to outside callers.
The kernel broadcasts events through a callback registered under EmitToken. It doesn't need to know who consumes the events, keeping it fully transport-agnostic.
6.2 Protocol
The protocol layer is the project's most important design decision. Every cross-process or cross-package message undergoes strict type validation:
// 方法常量(编译时检查)
export const MethodMap = {
"config.get": "config.get",
"config.set": "config.set",
"workspace.create": "workspace.create",
"task.create": "task.create",
// ... 15 个方法
} as const;
// Zod Schema 验证每个方法的请求参数
export const ParamsSchemaMap = {
[MethodMap["config.get"]]: z.array(z.string()),
[MethodMap["task.create"]]: z.object({
sessionId: z.string(),
input: taskInputSchema,
}),
// ...
} as const;Every cross-process message is checked by ReqSchema.safeParse(). Invalid requests are rejected before they enter the kernel, so they cannot corrupt internal state.
6.3 SDK: The Client Abstraction
ClientSdk is the UI's only interface. It correlates requests and responses by request ID, dispatches server events, and manages the transport lifecycle:
class ClientSdk implements ClientSdkInterface {
// 发送 RPC 请求,返回 Promise<Response>
async rpc(req: MethodRequest<MethodType>): Promise<Response> {
const request = ReqSchema.parse(req);
return new Promise((resolve) => {
this.pendingRequests.set(request.id, resolve);
this.transport.send(request);
});
}
// 注册事件监听器(支持按事件类型过滤)
onEvent(listener: EventListener): () => void;
onEvent(eventType: EventType, listener: EventListener): () => void;
}This design shields the UI from transport details. Stdin/stdout, Electron IPC, and a possible future WebSocket transport are all transparent to it.
6.4 UI: The Workbench Interface
The UI uses React 19, Zustand, and react-router-dom 7. Its central design is event-driven state reduction:
As a React Context Provider, WorkbenchControllerProvider does three things on mount:
Fetch initial data: workspaces, sessions, and a configuration snapshot.
Subscribe to kernel events: listen for app.config.changed to update global configuration and task.* events to update active task state.
Reduce task deltas: turn streaming task.delta events into a structured list of TaskParts.
The workbench has a sidebar with sessions and drag-and-drop project creation, and a main panel with messages and an input area. Settings provide visual model/provider management, MCP server switches, and tool/skill toggles. All changes persist through the config.set RPC.
6.5 Desktop: The Electron Shell
Traditionally, core logic is bundled into Electron's main process. Bun fits naturally as a subprocess instead: start it with child_process.spawn("bun", ["run", "host"]) and communicate over stdio using JSON-RPC. This also leaves room for other client frameworks later:
Kernel logic is fully separated and can be tested and deployed independently.
A main-process crash doesn't affect the kernel.
CLI and desktop shells can run on the same host and share the same kernel implementation.
7. Summary and Outlook
There is still plenty to improve, but the overall framework is in place. Agent Workbench is my first complete engineering exercise in building an AI agent workbench from the protocol layer to the UI. The project is available at agent-workbench.
Next steps include:
Filesystem tool implementation: protocols for read/write/edit/delete/search and other tools are defined; their core logic is being implemented.
Plugin system: configure MCP services and accompanying skills in one step, making tools and knowledge available out of the box.
More transports: add WebSocket and HTTP alongside stdio and Electron IPC to support remote execution.
Worktree isolation: run each session in its own Git worktree to avoid file conflicts between concurrent tasks.
A richer skill ecosystem: support skill dependencies and integrate mainstream skill marketplaces.