Skip to the content
Georgi DimitrovdaTuzzo

Valox Cinema

Making a fully AI feature film, starting with the two tests that could stop it

Role
Founder; directing research and design agents
Status
In progress
Source
Private repository
Stack
TypeScriptNext.jsPostgreSQLPrismapg-bossReact FlowSeedanceClaude CodeCodex
V2V1A1A2

In numbers

41

decisions in the architecture's table, each with its rejected option and a reopen condition

7

generation modalities behind one adapter interface, each model a typed row whose price and verified capabilities are data

~27,000

clips the verdict queue and automatic triage are sized for at feature scale

2

cheap decisive tests that run before full production: a resolution test and a ten-shot test

The problem

The film study's thesis is short: the last Bulgarian medieval epic was made in 1981 with state money, and a historical feature about Simeon the Great does not otherwise get made. Then Higgsfield made a 114-minute AI feature and published its whole production method. Read against my pipeline, their craft was ahead in asset discipline, spatial continuity and prompt structure. My automation was ahead: the layer that turns a script into 540 shot specifications took them 28 people and two chat windows, and in ValoxVSL that layer is the product.

The approach

I had agents tear down their method and my pipeline side by side. The study was merged into the pipeline repo, and each gap it found became an engineering issue, from a per-location spatial map to Bulgarian lip-sync from pre-recorded audio. I then started an autonomous run to design the successor pipeline that would make the film. Two parallel workflows read the old repo and gathered verified research, a background agent wrote an operator corpus that challenged my brief on eight points with file and line evidence, and three independent architecture proposals went to a judging panel. The result is a full set of blueprints, decision records and research files, and no code until I validate it.

How it works

  1. Assets before shots

    Characters, locations, props, costume and damage states, lighting rules and continuity constraints are designed before any expensive generation. Scenes are planned with spatial maps and master shots, so a reverse angle lands in the same room. The editor works while generation runs and orders missing coverage while it is still cheap to make.

  2. Voices first, picture second

    The picture is generated end to end; the performances are not. Actors record the lead voices first, in controlled sessions with repeated takes and overlapping dialogue, and the model animates to an existing waveform, so it never has to generate Bulgarian speech. That inverts casting: the best performance gets the part and the face is generated separately. A blueprint covers what follows an accepted clip: subtitles, dubbing, picture-preserving lip-sync, and voice cloning as a consent-gated capability.

  3. Two tests that can stop it

    No source verifiably confirmed the video model above 720p, and a cinema master is 2K. So the first test is a live call to the provider: the real maximum output resolution and concurrency confirmed in writing; a weak answer points the film at streaming instead of cinemas. The second is ten representative shots: intimate dialogue, an establishing view, hands, firelight, an emotional face, cavalry, a melee, a phone-shot motion reference, shallow focus and a prop insert, each at draft and at final resolution. It counts generations to an accepted take and whether cheap drafts predict the final ranking.

  4. One seam for every provider

    A single adapter interface covers seven generation modalities. Each model is a typed row whose price and verified capabilities are data, so cost becomes a routing input and a provider's unverified claim cannot be routed to. The same tool registry serves MCP, a command line and the UI routes, so an agent and a person call the same tools.

  5. Spend limits written as code

    In the current pipeline the money guardrails live in skill markdown and the spend counters block nothing. The design enforces per-job and per-scene limits, estimate-before-spend and hard stops in the application, where a prompt cannot argue past them. The study filed the same rule against today's tool as an issue: a budget governor enforced in the MCP tool layer.

  6. Every clip with a lineage and a verdict

    Each generation carries its lineage and a persisted verdict, and a verdict queue with automatic triage is sized for the tens of thousands of clips a feature needs. Video QA starts with a deterministic optical-flow and SSIM prefilter, then a frame-grid review at a sampling rate I set, so the cheap checks run before the expensive ones.

What I chose, and what lost

Chose

Asset-first generation, with live motion reference only for physically hard shots

Over

Shooting the film live, or regenerating a motion that keeps failing until it works

A motion that fails the same way twice is a structural failure, and more rolls do not fix structure. A phone-shot plate supplies the body mechanics, and the model replaces costume, setting and light.

Chose

Run the two cheap tests before full production

Over

Starting production on the assumptions in the study

Every downstream figure in the study depends on two experiments that have not run, and the resolution test decides whether this is a cinema film at all.

Chose

Capabilities held as verified data

Over

Trusting a provider's published claims

A claim nobody has verified cannot be routed to, so a wrong claim costs nothing.

Chose

Design the successor before building it

Over

A greenfield rewrite started straight from the brief

Writing the design first put the spend rule into the requirements before any code could repeat the old flaw, and the design records the rejected option and a reopen condition for each of its 41 decisions.

Outcome

No film exists. The study is merged, and the issues it filed against the pipeline are all still open. The successor design is complete, critiqued and fixed, and it has no application code, no pilot and no commit. Neither decisive test has run. The design run hit a usage limit once and resumed from its journals, and a verifier agent refused to touch the project until it was added to its allowed project list. The project's review notes carry the warning: a dossier this thick can support action or replace it, and the tests decide which.

What comes next

Run the resolution test, then the ten shots, then cut a 90-second proof reel from the best of them.