A working prompt formula for building a coherent series without copying television, films, or people. Pick a visual grammar, give every reference an explicit job, then generate an original scene. This is the bridge between a lawful source dataset and the worlds we actually own.
Build a shot brief
The model gets a job description, not a pile of vibes. This produces a structured video prompt with picture, camera, sound, constraints, and reference roles.
10 original synthetic recipes
These are ready for a 5.167-second / 124-frame / 24fps generation wave once the existing renderer endpoint is reachable. Click one to load it into the formula.
Dataset contract & rights trail
Sources are private reference material. Synthetic outputs are original transformations. Every layer is linkable back to an item-level source declaration.
Important: We train or reference visual grammar, not someone else's show. A clip being on the internet is not permission. We collect only sources with documented rights declarations and retain the evidence.
Current source foundation
5 source records · 10 normalized private candidates · industrial/machine, crowd, home, travel, and research-machine grammars.
candidate / needs review
H3-ready target
1152×640, 24.000 fps, 124 frames, H.264/AAC, paired flowing caption. A small proof set comes before a larger training corpus.
124 = 17×7 + 5
Public exhibits
Only approved synthetic exports go public, with a fictional/AI-generated disclosure. Raw source clips and training pairs remain private.
human gate
Training sequence — don’t burn money early
The first goal is a reliable generation language, not a 5,000-step adapter.
Phase 0 · prove the pipe
10 source candidates → 10 original synthetic clips. Validate duration, fps, audio, captions, and provenance. Reject aggressively.
zero ambiguous rights
Phase 1 · curate
Build 50–80 diverse clips across the seven grammars. Keep 10–15% held out. Remove black frames, watermarks, near-duplicates, and slow-motion bias.
hand-reviewed
Phase 2 · smoke train
Rank 16, 1,000–1,500 steps, 24fps, 73 or 124 frame grid. Use one trigger strategy only: baked into captions or trainer field, never both.
paid run / human review
Phase 3 · A/B evaluate
Same prompts and seeds, adapter scale 0 vs 1. Score identity stability, camera language, sound, original-world adherence, and unwanted source resemblance.