Problem
What was needed?
Publishing explainer videos regularly meant an unsustainable manual chain of research, narration, editing, captions, thumbnails, and uploads.
From a topic file to a publish-ready package: a human-approved YouTube automation producing narration, storyboards, captions, thumbnails, and Shorts.
Platforms
Automation
Role
Architecture and end-to-end development of the pipeline.
Tech stack
Node.js · FFmpeg · OpenCLIP
Status
Launch-ready

Problem
Publishing explainer videos regularly meant an unsustainable manual chain of research, narration, editing, captions, thumbnails, and uploads.
Solution
A single-command pipeline: narration and storyboard are generated from an editorial brief, stock footage is scored against the script by a local OpenCLIP model, Whisper produces word-timed captions, Shorts are derived from the long cut — and nothing reaches publish without passing a hard quality gate.
Outcomes
Visual-script alignment via local OpenCLIP scoring, no cloud vision cost
Quality report with LUFS/true-peak checks + mandatory human approval
SQLite institutional memory preventing topic and footage repetition
Automatic Shorts derivation and timezone-aware scheduling
Build note
The pipeline has shipped 9 long-form videos and 36 Shorts for a YouTube channel. Its defining trait is discipline rather than autonomy: scripts stay editorially human-made, the system accelerates production, and publishing requires a passing quality report, a selected thumbnail, and explicit human approval together. Built-in cost accounting reports per-video token, character, and infrastructure spend.
One tracking system, delivered through two interfaces shaped around each role’s daily workflow.