Skip to the content
Georgi DimitrovdaTuzzo

Transcribe service

Long recordings in, a diarised transcript and a screen timeline out, with each run costed

Role
Solo, directing agents
Status
In progress
Source
Private repository
Stack
PythonFastAPIReactTypeScriptffmpegSQLiteOpenRouter
Job page of the transcribe service for a demo recording of a weekly sync: status done, a pipeline strip with speech, video and screenshot stages, a detail pane, a log tail and a list of output files such as transcript, meeting record and meeting notes.
Demo data only.

In numbers

~90min

longest recordings it was hardened on, in three languages including Bulgarian

1

call site for every model request, with key rotation and provider fallback

The problem

Meetings and screen recordings hold decisions that do not reach a document, and a transcript alone loses what was on screen. I needed a record I could audit: who said what, when, and what was shown at that moment.

The approach

I built the pipeline first as a skill that Claude and Codex could run and hardened it on recordings of up to about 90 minutes in three languages, including Bulgarian. Then I wrapped it as a local service with a job page and a ledger.

Jobs ledger listing two demo jobs numbered 001 and 002, with columns for media, status, billing, stage, length, cost and creation date, and buttons for Scan and Import and New Job.

How it works

  1. Parallel stages

    Transcript and video ingestion start together. Each video clip goes to the model as soon as ffmpeg has cut it, and one merge barrier, which can be re-run, assembles the document. Failed items retry at the end of a stage and from a Retry failed button in the UI.

  2. Screenshots against a shared hash set

    Screenshots are picked, grabbed and de-duplicated against a hash set shared across segments, so repeated frames collapse to one image.

  3. One call site for model requests

    All model calls go through one function. It rotates pooled API keys on 402 and 429 responses, retires a key on 401 and 403, and falls back from OpenRouter to Vertex AI with a client-side check for key expiry.

  4. Billing mode per job

    A job must declare a billing mode, and there is no default. The mode decides which providers the job may reach and wins over the environment file.

  5. A test that reads the source

    A test reads the source back to confirm that no key lives in the repository.

New job form: a media browser over a demo recordings folder with seven files, a choice between personal and company billing, an output folder and speaker settings.

What I chose, and what lost

Chose

A billing mode declared on each job

Over

A default taken from the environment file

The mode is visible on the job and a stray environment setting cannot override it.

Outcome

The service has a licence, a contributing guide and three screenshots of its job page and ledger. A speaker-harmonising defect that cost 776 labels over two runs was still open in the skill version it grew from. The recordings it was validated on belong to clients and teams, so none are shown.

What comes next

Run it on a fresh set of recordings and confirm the speaker defect is gone.