LayerV  /  Video Production System Mission report

Academy video series

Video Production System Mission Report

A model-agnostic AI video production system, customised for the LayerV Academy. It runs with Claude, ChatGPT, open-source models and any other generative model. Built from scratch around the Academy in under 3 weeks: your lessons, your character, your brand, your cost structure. It runs without its creator, is fully customised, and is meant to evolve with the team.

Production period: 29 July to 17 August 2026

4Stages, only 1 spends
~$50Total generation spend to build and prove the system
10:1Cost ratio across the series, with the system against without
465Automated checks protecting the series rules

Contents

  1. Before and after
  2. What this gives LayerV
  3. How the chain works
  4. The 4 skills
  5. The guardrails
  6. Built to evolve, and to be made yours
  7. Where it stands, and what is next
  8. Questions you are likely to have

1Before and after


The Academy lessons were already written. What did not exist was a repeatable route from a written lesson to a finished film.

Before
  • No route from a written lesson to a finished AI video.
  • Every correction meant paying to generate again, and the price was only discovered after committing to it.
  • AI-generated text warped and misspelled inside the frame — a limit of the models themselves, not something a better prompt alone fixes.
  • Captions drifted from the voice, because nothing measured the real recording to line them up.
  • The character looked different from one generated shot to the next, with nothing tying one generation to the last.
  • Hiring an expensive video maker who charged hundreds per video.
  • Every shot was a one-off. Nothing compounded.
After
  • An end-to-end chain: a lesson goes in, a finished short AI video comes out, voiced, captioned, scored and levelled. The operator sends it through the four stages and the chain runs the rest.
  • Revising a finished film needs no new generation. The rebuild works through composited layers and settings, and takes a few minutes rather than a new invoice.
  • Every word on screen is composited locally in the brand typeface after generation — never drawn by the model, so it never warps.
  • Captions are recut against the real voice recording, not a theoretical pace, and ship as a separate file too: searchable, translatable, reachable by a screen reader.
  • Shots accumulate in a reusable library, so a shot's look repeats exactly on reuse; new shots stay locked to whatever the character palette settles on next.
  • The system prices the work before submitting and stops until that exact amount is approved in writing.
  • 465 automated checks enforce the series rules, each proved by reintroducing the defect it catches.

2What this gives LayerV


The economics of a video series are decided by a few structural choices. Here is what those choices saved, and how each figure was established.

Choice What it saves How it was established
The system itself $35 with it against $348 without A ratio close to 10:1 that holds at any provider's price point
Reusable shot library $30.48 already saved, against $25.50 of footage generated Reuse has returned more than the footage cost. Measured, and revised downward from $48.73 by an audit
Generate at 720p, composite type locally at full resolution 56 to 60% off every generated shot Measured on tight crops rather than assumed: no visible loss on the character, and none at all on the type
Revisions rebuild locally No generation cost, however many rounds it takes Type, panels, captions, transitions and sound are never generated, so changing them cannot cost credits

What an episode actually costs

The compounding part. Reuse has already returned more than the entire footage spend, and that happened while generation was deliberately held to the strict minimum. The saving grows on its own once the character palette is settled, because every shot then stops being a one-off and becomes stock.

3How the chain works


4 stages, each with an explicit contract: what it reads, what it writes, what it may not touch, and where it stops. The chain halts wherever a human decision belongs, and nowhere else.

Stage What it does Waits for Spends
1. Script Lesson page to manifest and script, with burned captions and their sync timings Script approval No
2. Storyboard Manifest to shot list and prompts, rendered as real images with teaching panels composed at the true geometry of the film Storyboard approval No
3. Produce Approved storyboard to stills, footage, voice and finished file Written approval of an exact amount Yes
4. Revise Notes on a finished cut to a local rebuild, with before and after at each note Nothing No
Contact sheet of a finished episode, every beat with its timestamp and caption
What the chain produces. A finished episode laid out beat by beat, each frame with its timestamp and the caption playing over it. This sheet is generated automatically with every build, and it is how a reviewer checks a film without scrubbing through it.

It can be run by an operator who has never seen it. An operator with no prior knowledge of the chain fed a lesson in at one end and took a finished, voiced, captioned episode out at the other. The command layerv-academy-video makes the workflow run end to end.

The first cut is a draft, not a delivery. No generative pipeline lands the intended result on the first pass, and this one is no exception: expect rounds on pacing, emphasis and wording before an episode is right. That is precisely what the architecture is for. Those rounds rebuild locally, they cost no generation, and they are the cheapest part of making the film.

3 of the 4 stages cost nothing to run. A lesson can be written, storyboarded, reviewed, rejected and rewritten as often as needed. Money enters once, at a single stage, after an exact figure has been approved.

4The 4 skills


Each stage is a skill with a written contract. What makes them reliable is what they do and what they refuse to do: ownership of each field is assigned to exactly 1 skill, so no stage can silently overwrite another's work.

5The guardrails


Every guardrail below exists because something went wrong at least once. Each was proved able to fail before it was trusted: run against the broken code first, watched to fail, then adopted.

A · Spend

The amount approved is the amount billed

  • The approval is a ceiling, checked to the cent against the manifest at submission time. Over it, the run refuses and says why.
  • 4 ways of understating a quote were found and closed in a single audit, including one that announced $0.02 for an episode priced at $3.20.
  • The most expensive possible mistake was made impossible: nothing designated which image opens a shot, so an episode could have gone to generation with a character reference sheet as its first frame. Refused by name now.
B · Quality

Defects are caught before the money moves

  • Reused shots cannot teach the wrong lesson. Shots with one lesson's text baked into the image were being reused elsewhere, so a film could display one thing while the voice taught another. Refused.
  • Generative models cannot count, so a prompt containing a figure or a countable object is refused outright. The first cut drew 4 cards at the exact moment the caption said 3.
  • The validator catches what an eye would not. A speaking beat showing no panel, no figure and no word used to pass everything: sweeping the catalogue found 12 of them where 4 were expected. An episode could also ship silent without a single check objecting, because the audio inspection stopped as soon as there was no audio.
  • 3 quality checks were found reporting success without measuring anything. Once repaired, they revealed 10 sound effects out of 11 were inaudible while production reported green.
C · Revisions without regeneration

Notes in plain language, answered locally

  • Write the note as you would say it: the cuts are too slow, he moves too slowly in shot 2, that caption is wrong, that sound effect is too loud.
  • Motion that reads too slowly replays up to 1.35× faster without regenerating anything.

6Built to evolve, and to be made yours


This is version 1 of a system shaped around the Academy as it exists today. It was built to be changed by the team that runs it, not frozen at handover.

What the operator changes, without touching code

2 ways to put words on screen

Every word the viewer reads arrives by one of 2 routes, and the choice belongs to the operator. It is worth seeing the difference, because it decides both what a change costs and what the letters look like.

Drawn by the model · paid, fixed Panel text generated by the model, with distorted letterforms
Composited locally · free, editable Teaching panel composited locally in the brand typeface
Left: the words sit inside the generated picture. The model drew them, so the letterforms wander, and changing a word means paying to generate the shot again. Right: the same kind of teaching panel, composed locally in the brand typeface after generation. Rewording it, re-timing it or restyling it costs nothing.

Look at the bottom of the left frame: its caption is crisp, because that one was composited too. Both methods sit in the same image, and the difference shows without being pointed at.

3 layers, and how each one can be swapped

Layer What it does Changing it
The assistant Reads the method files and writes the script, the storyboard and the revision notes The rules are plain markdown, not code, so they are readable by any capable assistant. Operated so far with a single one
The assembly machinery Turns approved material into a finished film: cuts, transitions, type, teaching panels, captions, sound mix, loudness Calls no AI at all. Standard code and standard media tooling. This is why revisions cost no generation, and why no vendor sits between you and a finished film
The generation models Produce the footage, the stills and the voice Named in configuration. The registry declares 8 video endpoints and 3 voice options; change the identifier and the spend gate reprices automatically from the live rate

Today the series is generated with the frontier Seedance video model, named in configuration rather than written into the code.

7Where it stands, and what is next


What exists today: The method files and the framework reference page are already in your hands.

The system is in service, and like any version 1 it has a roadmap. What follows is the open list as the system itself keeps it, rather than a list reconstructed for this report.

Open item Why it matters Owner
Character palette not approved Every shot generated before it is settled stays provisional. A palette change would invalidate the $25.50 library. This is the single decision that unlocks the reuse economics at scale Design lead
First run on the consolidated spend gate The 4 approval steps were merged into 1 command. That consolidated path has not yet carried a paid run through to a finished film, so the first one is worth walking step by step on the cheapest episode Operator, with the reviewer approving the amount
Prompt grammar migration 3 of the 4 existing episodes still carry the earlier grammar. A guard refuses those shots by name, so the gap surfaces before money moves rather than after Operator
Two prompt-level checks still to build The validator confirms that a figure is spoken somewhere in the episode, but not which card it sits on. It also does not yet refuse a prompt that reserves space while scripting a light effect across it. Both gaps have produced errors a check would have caught Operator
Loudness of the earliest films The first films play 4 to 5 dB quieter than the current chain produces. Levelling them is free; redelivering is an editorial call Content reviewer

The 3 things worth doing first

The layer after that

Once the team has run an episode or 2, 2 pieces are worth building on top of what exists.

The rule above the others. Nothing is generated without a written authorisation naming the exact amount, calculated at the moment of asking. A revision is not a generation: revising a finished cut is local, costs no credits, and takes minutes.

That distinction is the entire economics of the series. Holding it is what keeps the cost per episode where it is.

8Questions you are likely to have


Can it spend money without my approval?

No. The approval names an exact amount and works as a ceiling, checked to the cent at submission. Over it, the run refuses and says why.

Will the first cut be what we wanted?

Rarely, and that is true of any generative pipeline. Plan on a few rounds of notes on pacing, emphasis and wording. The system was built around that expectation: those rounds rebuild locally and cost no generation, so iterating is the cheap part.

I do not like a finished film. What now?

Write your notes the way you would say them. The rebuild works through composited layers and settings rather than a new generation, and takes a few minutes rather than a new invoice. A note that genuinely needs new footage is named as such and priced separately, never absorbed silently.

What is free of generation cost, and what is paid?

No generation: scripts, storyboards, every caption, panel, transition and sound mix, and all revisions. Paid: new footage, new stills, new voice lines. Nothing else.

Do we need to write code to use it?

No. The chain runs on documented commands, stops at every point where a decision belongs to you, and takes review notes in plain language.

How do we make the animation better?

Through the prompt. Each shot's motion comes from a written prompt, and how carefully that prompt is written is what decides how Vega moves and presents. The system supplies the structure, the brand constraints and the automatic refusals; the craft of the prompt itself belongs to the operator, and it is where the difference between an average shot and a good one is made.

Can we change the look, the narrative or the editorial rules?

Yes, in the method file of the stage that applies them. A change there applies to every future episode, and the automated checks protect the rest of the series while you change it.

How long does an episode take?

Machine time varies with how many shots are generated versus reused, and with the provider's queue: roughly 15 to 45 minutes from an approved storyboard to a finished file is a realistic range. The calendar time on top of that is review: how fast the script and the storyboard get approved, which is where the schedule actually lives.

Are we tied to one AI provider?

No, and that is deliberate. The assembly machinery calls no AI at all, so nothing proprietary sits between a lesson and a finished film. The rules each stage follows are plain markdown rather than code, so any capable assistant can read them; the series was authored with Claude, and nothing in the chain depends on that choice. The video model and the voice are named in configuration, so either can be swapped for a more capable or a cheaper one without touching the chain.