Field notes/Announcements

Introducing VideoCue - the AI video tool we wanted to use

Five minutes of cinematic, narrated, on-brand video from a single prompt. Here's what we built, why, and who it is for.

June 5, 2026·5 min read·1,149 words
Introducing VideoCue - the AI video tool we wanted to use

We spent eighteen months trying to publish a single long-form AI video without seven open tabs, three subscriptions, and a folder of half-broken Python scripts. VideoCue is what we wished existed.

The state of "AI video" right now

The category is louder than it is useful. Every week another text-to-clip tool launches with a teaser reel of perfect six-second loops. Try to make a real five-minute narrative with one of them and the cracks show up fast:

  • Voice quality is "uncanny TED-Talk" by default - you spend an hour wrestling SSML to get a single line that sounds human.
  • Image consistency falls apart at frame seven. A character's eye colour changes. A coffee cup becomes a wine glass. The palette drifts from warm-tungsten to neon-Vegas in twenty seconds.
  • The render is somebody else's opinion. You get a thirty-second clip with branded transitions baked in, no timeline to edit, no audio stems to remix.
  • Pricing punishes the actual workload. "Unlimited" plans cap you at four minutes of monthly output. Per-second pricing makes a real video cost more than hiring a freelancer.

We talked to forty-three creators and agencies during the private beta. Every single one had built their own duct-tape rig: Claude for script, ElevenLabs for voice, Midjourney for stills, Runway for B-roll, ffmpeg for assembly, Descript for cleanup. Each rig was different. Each one took three to six hours to ship one video.

VideoCue replaces the rig.

What it actually does

You give it a prompt. It hands you back an MP4.

In between, here is the pipeline that runs on every render:

  1. Script generation - a structured-output LLM call drafts a scene-by-scene script with cue boundaries, image prompts, and voice markup baked in. You can edit any of it inline before generation continues.
  2. Voice synthesis - ElevenLabs V3 with Bradford as the default narrator (configurable). The script is pre-marked with [contemplative], [pause], [exhales] tags where the rhythm needs them. We chunk on character boundaries to respect V3's 4,500-char window without losing pitch continuity between chunks.
  3. Image generation - Gemini 2.5 Flash Image (Nano Banana) for the primary hero frames, with a reference-anchor protocol that keeps the visual language locked across forty-plus images. A separate post in this set walks through how.
  4. Music + ambience - three royalty-cleared stems mixed under the voice with automatic side-chain ducking.
  5. Render - a three-pass ffmpeg pipeline (image-to-clip, concat with crossfades, overlay-plus-audio-plus-watermark) that exports 1080p H.264 in roughly two times realtime on a single mid-tier worker.

The whole thing runs as a queued Laravel job, shells out to Python where Python is the right tool, and streams progress back to the browser over WebSockets.

The video that proves it works

The reference render in our seed library is called "Your Life as One Coffee Bean" - a five-minute, twenty-seven second POV narrative that follows a single bean from a hillside in the Yirgacheffe region of Ethiopia to a barista's tamper in Brooklyn.

Some numbers from that render:

  • 8,420 characters of V3-marked script across five narrative beats (Seedling, Cherry, Mill, Roast, Cup)
  • 40 hero images generated against a single reference frame
  • 3 music stems, all ducked under the VO automatically
  • 56 individual ffmpeg fragments stitched into a single timeline
  • $2.10 in provider pass-through costs end to end

It is not a tech demo. It is a video we would actually publish on a real channel. That distinction is the entire bet of the company.

Who this is for

Three audiences, in this order:

Solo creators who want to publish more long-form video without becoming editors. The format that works on YouTube right now - explainer, narrative, listicle, micro-doc - is technically achievable for one person, but operationally exhausting. VideoCue flattens the ops.

Small agencies who need to ship volume for clients. The agencies we beta-tested with are running six to twelve client channels in parallel. They do not need bespoke per-shot crafting; they need consistent, on-brand video on a weekly cadence. The API plan is for them.

Developers building video features into their own products. The same pipeline behind the editor is available as a REST API. You send a script, you get back an MP4 URL and a JSON of the underlying assets. The coffee video took 14 seconds of code to call.

If you are looking for Hollywood-grade VFX, this is the wrong tool. If you are looking for "type a prompt, get a TikTok" - also wrong tool, the cycle time is too high. If you are in the middle - narrated, formatted, on-brand long-form - you are exactly who we built this for.

What we got wrong on the way here

A few honest notes from the build.

We spent two months trying to do everything in Laravel. PHP is great for HTTP and queues, terrible for media pipelines that need to call libavfilter at sample-rate precision. The pipeline now shells out to a Python worker pool that wraps ffmpeg, and Laravel just owns orchestration and state. Three to four times faster, half the bugs.

We over-invested in fancy transitions early. Crossfades and a tasteful "speed-ramp on cut" are it. Everything more elaborate looked AI-generated within a week and aged badly.

We assumed people would want music selection. They want defaults that sound good. We added a one-line "music mood" prompt instead of a library picker and 92% of users never change it.

What is shipping today, what is shipping next

Today:

  • The full text-to-MP4 pipeline, 1080p, watermarked on the free tier
  • Five templates: POV narrative, explainer, listicle, micro-doc, recipe
  • Voice library of 14 ElevenLabs voices, three default music moods
  • REST API with a 120 RPM ceiling on Studio plans
  • Auto-generated YouTube metadata (title, description, tags, chapter timestamps)

Next 90 days:

  • Voice cloning on Studio (one-time enrollment, then any script in your voice)
  • Stems export so editors can remix in Premiere or Resolve
  • A Figma plugin for storyboarding before render
  • Long-form support (15+ minute pieces, multi-narrator)
  • A Showcase page that pulls real customer renders, not just our own

Try it

Start free - 3 videos, 1 minute each, no credit card. The first render is on us; if you publish it, send us a link and we will watch it. We are at the size where every paying user gets a direct line to whoever is on call.

See the pricing page when you are ready to upgrade past the free tier.

Tagged
announcementlaunchvideocue
Share

Your first cut
is on us.

+ New videoFree plan · no card · 1080p export