Sesamebot
A Discord bridge that makes a whole voice channel sound like one caller to a voice AI
- Role
- Product owner, directing a lead agent and four teammates
- Status
- Paused
- Source
- Private repository
- Stack
- Pythondiscord.pydiscord-ext-voice-recvfaster-whisperffmpegpytestsystemdGitHub Actions
In numbers
64ms
send-pump cadence: one 1,024-sample chunk at 16 kHz, voiced or silent
13
iterations it took to find the silence keep-alive
1.5s
minimum time the elected speaker holds the floor
The problem
The voice assistant is built for one person in a browser tab: one microphone, one voice, and a voice-activity detector that decides when a turn has ended. A Discord channel has several people, and Discord stops delivering audio when nobody speaks. v1 worked for a single speaker on my server. I wanted a v2 that could hold a conversation with a group, remember earlier sessions and be covered by tests.
The approach
The bridge makes a channel look like one well-behaved caller. Audio arrives from Discord as 48 kHz stereo per user, passes an allowlist and an RMS gate, and is downsampled to 16 kHz mono and queued; the assistant's 24 kHz replies are resampled to 48 kHz stereo and played back. For v2 I filed 14 issues with detailed comments and wrote a handoff document, then had a lead agent split the work across four teammates in separate git worktrees that fed one v2 branch.
How it works
A pump that sends silence
When everyone goes quiet, Discord stops calling the bot's audio sink, so the assistant never hears the silence that ends a turn and waits as if the speaker were mid-sentence, indefinitely. A send pump emits one 1,024-sample chunk every 64 ms whatever happens: voiced audio if any is queued, zero-filled silence if not. Finding that took 13 iterations, and a test pins the cadence for voiced and empty queues.
Floor election
On every write the sink measures each user's RMS level. The loudest speaker above the silence threshold claims the floor and holds it for at least 1.5 seconds, renewed while they keep talking, and only their audio is forwarded; the others' packets are dropped, as interruptions are in a real conversation. A server with a single host can switch the election off.
Name markers
The assistant hears one voice at a time and cannot tell whose it is. The first time a person speaks in a session, the bot plays a pre-rendered text-to-speech clip of their name or alias into the stream and drops the packet that triggered it, so the marker lands cleanly before their first words. Markers can be pre-baked for a whole channel or server.
One injection path
Markers, the opening bootstrap preamble and cross-session recall all go through one function: it splits a PCM clip into pump-sized chunks, queues them and gags every user microphone for the clip's length plus a half-second tail, so two streams never overlap in the assistant's input.
A safe word heard locally
A small local speech model listens to the floor holder, and a match on the safe phrase hangs up through the same teardown as the leave command, at a cost of about 5 to 10 percent of one CPU core on the small server.
Review by running
The lead reviewed each teammate's pull request by running the suite itself before squash-merging into v2. A teammate's claimed count of 55 tests did not match the tree, which held 67, and a claimed alias command had to be found in the code before the merge. Teammates negotiated interfaces with each other directly (a lifecycle-hook registry, a queue swap, a bootstrap that takes the sink, the speaker names and a recall clip), and the lead stayed out unless asked to arbitrate.
What I chose, and what lost
Chose
Floor election that drops interjections
Over
Mixing every speaker into one stream
One speaker at a time is how a conversation already works, and the assistant's turn detector expects one voice.
Chose
Dropping whisperx and pyannote
Over
Local speaker diarisation
Both were too heavy for the small VPS, and Discord already labels each packet with its user.
Chose
I merge v2 into master after validation
Over
Agents pushing to master
The lead opened the pull request and I validated it before it merged; no automated push reaches master.
Outcome
v2 ran healthy on my server, signed in to Discord as Jeisan Steisan, the name my agents carry. The bot depends on an unofficial way of signing in to the vendor's service, which is outside the vendor's intended use, and the vendor ended anonymous access, so it stays a private tool. I scoped v3 on a different voice stack, then found that an existing agent framework covered the need, and v3 was not built.