Skip to the content
Georgi DimitrovdaTuzzo

Valox Office

An open-source 3D agent office, stripped of the slop and rebuilt as a harness for agents

Role
Product owner, directing an agent fleet on an MIT fork
Status
In progress
Source
Private repository
Stack
TypeScriptNode.jsthree.jsVitenode-ptyxterm.jsWebSocketzodMCP SDKExcalidrawMermaidthree-mesh-bvhBlender (Python)Playwrightnode:testGitHub Actionssystemd
Third-person view of the author's own 3D avatar standing on the open office floor: a man with dark curly hair and a short beard in a navy blazer with red embroidered cuffs, a black shirt and olive trousers. Behind him are desks with coding agent workers under name labels, the issues board, the task queue, a pull requests board, a whiteboard showing a checkout funnel chart, a potted plant and a gong.
Walking the floor as my own avatar, in my fork of the MIT-licensed AgentSystemLabs/agent-office. Every worker is a coding agent; the project and names are demo data.

In numbers

20

MCP tools the floor manager drives, plus 5 desk tools every worker gets

32

isolated worker accounts in 16 floor groups, each an unprivileged pool user

20.1 to 6.0 ms

median frame time with 12 streaming workers

9

independent verification rounds before the PC runner shipped

The problem

I run Claude Code, Codex and OpenCode across several accounts, many repos and more than one machine, and I wanted one place to see and direct all of it: projects as floors, agents as staff at desks, and the people I work with able to walk in, watch a terminal and take over. I found AgentSystemLabs/agent-office, an MIT project by another developer, that already had the 3D office, shared PTY terminals, per-worker git worktrees, issue and PR boards, WebRTC voice and an Excalidraw whiteboard. It was a good scene with no harness behind it.

The approach

I started from it. I kept the scene, the terminals and the worktrees, stripped what I call the slop, and built the tooling that turns an office into a harness. The fork keeps the upstream MIT notice, and a file named UPSTREAM.md records each upstream PR it reviewed as ported, adapted or skipped. The code was written by an agent fleet I directed. Codex and Claude builders worked in worktrees; an Opus integrator rebased, ran the suite and reviewed each endpoint for auth and floor access; independent verifier and security-review rounds came next; one coordinator session only merged on green and deployed. My part was product direction, play-testing, bug reports and accept or reject calls. Output from the cheaper, faster builder model is re-validated by the strongest model before it merges, a rule that came from a screen preview that passed shallow tests and broke on real Claude Code screens. When a headless browser test clipped my real mouse cursor to its window, the fix went into each agent brief: fake pointer lock in every headless test.

First-person view looking straight down at the floor: the avatar's navy jacket with red embroidered cuffs, dark trousers and brown shoes below the camera, and its two hands held out over beige floor tiles.
First person, looking down at the avatar's own hands and shoes.

How it works

  1. A floor manager that is an ordinary worker

    Each floor runs one Claude Code or Codex worker as its manager, driving a custom MCP server with 20 manager tools (hire_worker, prompt_worker, read_worker_screen, merge_pull, start_meeting, ask_humans and more) plus five desk tools each worker gets. No extra orchestration framework sits above it. The manager acts for whoever hired it, its hires may only use accounts that person can use, and it merges only pull requests whose checks are green. A real-MCP end-to-end test drives each tool for each worker kind. Workers take their default names from my belote bots, and managers are khans and tsars.

  2. Three agent CLIs on one floor

    Claude Code, Codex and OpenCode workers sit side by side, and the manager can hire any of them. Each person brings their own AI accounts per provider, can lend one to a colleague with the owner's approval, and sees a usage meter per account; a running worker can move between accounts. Meetings seat two to five workers in patterns such as debate, map-reduce, red and blue, or review panel, and each seat picks its own provider, model and account, so a Claude, a Codex and an OpenCode worker can argue at one table against a deadline.

  3. Workers isolated by the operating system

    On Linux each AI account runs as its own unprivileged pool user, 32 users in 16 floor groups, reached through one sudo-ruled helper. Hooks arrive over a unix socket, and each read the office makes of a worker-writable file uses a no-follow, owner-checked open. It was denial-tested on a throwaway pool on a real server, mutation-checked, and put through an Opus security review that found a critical root-script injection through a floor name before the PR was opened. It went live, and the running workers restarted under pool users.

  4. Runners on personal and remote computers, failing closed

    A worker can run on someone's own PC, so the logins stay on that machine. The runner opens one outbound WebSocket and accepts a fixed set of named operations (spawn, attach, write, resize, kill, diff, commit, PR, worktree); no executable, shell, environment or working directory is accepted from the socket. Enrolment tokens are hashed, single-use and expire after 10 minutes. It shipped disabled and refuses to run until isolation is on; round five of nine verification rounds found that an office shell could read the session-signing secret and forge an owner cookie. A remote computer joins over SSH instead: workers run in tmux behind a forced command with a pinned host key, and a confidential floor writes none of its work to the office's disk.

  5. Roles, with guests denied by default

    Everyone starts as a guest, and admins and members are granted per floor and per group. The guest role is a message filter: any client message without an explicit rule is refused, and a CI test fails until each new message type has one. The integrator inventoried 85 client messages, 57 server messages and 37 HTTP and WebSocket entry points, then ran a hostile-guest harness of 174 messages and 59 assertions that found no work leaking. A floor you cannot enter answers exactly like a floor that does not exist.

  6. Performance as a measured loop

    PR #81 interleaved base and branch on the same scripted 30-second walk with 12 streaming workers on an RTX 2080 at 1080p. Merged, vertex-colour-batched meshes, BVH raycasting and up-front shader compilation took median frame time from 20.1 ms to 6.0 ms and draw calls from about 1,550 to about 355. An in-game stats overlay with a 30-second recorder followed, so the next regression arrives as a number.

Over the shoulder of the floor manager, a worker figure with an orange spiked head seated at the boss's desk in front of a laptop that shows terminal text. Through the glass wall behind it are the open floor, a Task queue board, a Pull Requests board, a lounge sofa and a lift marked Storefront, and two empty chairs face the desk.
The floor manager at work, with the floor beyond the glass.

What I chose, and what lost

Chose

Fork the MIT upstream and credit it in the repo

Over

Build the office from scratch

It already had the 3D scene, PTY terminals and worktrees, so the work could start on the harness. UPSTREAM.md keeps the lineage auditable.

Chose

One outbound connection and a fixed list of named operations for PC runners

Over

An extra agent framework, a broker and a remote shell endpoint

The named operations cover the workflow, and a socket that accepts no command line has no command line to abuse. The README states that production failure rates have not been measured.

Chose

Move orchestration into the office's own manager and runners

Over

Keep Orca, a third-party MIT tool, as the engine that spawns agent terminals

The office soon became its own harness, with a manager, worktrees and boards, so a second orchestrator duplicated them. The move is partial: the build's coordinator was still a Claude Code session outside the product, and dogfooding through PC runners has begun.

Outcome

The fork's commits sit under my GitHub identity and were written by the agents I directed. Some of that work is ported upstream code: by an estimate from PR titles and bodies, about 17 percent of the merged PRs and 29 percent of the added lines. Upstream kept shipping and later built its own versions of a few things the fork had first, including per-account sign-ins and an MCP server. The office runs on my own server and the repo is private. Not yet verified: a full-suite pass on the latest commit, and failure rates for the runner in daily use.

What comes next

Whiteboard pages on the 3D board, approval prompts for PC workers, and moving my own daily agent work into the office.