In the is voice the AI slop fix article we managed to automate the generation of the short explainer video fully. The biggest achievement was creating it entirely using AI with no human involvement at all. It is not perfect, but it is exponentially better than what was possible twelve months ago.
Why bother? I believe that with this tooling we should aspire to create ever more multi-modal articles. The idea is to have a short two-minute video that plays with the article, giving readers a sense of what's contained and helping them decide whether to read further, save for later, or take another action. I believe there is little value in having a brand-new technology and simply doing old things faster.
The solutions used: Eleven Labs, Agent Opus, OpenAI, GPT Image, two prompts, and one custom agent. You can see the outcome below.
The big human involvement was in creating the script in the first place, and that script was then put into Eleven Labs to create the voice file. This voice file was then used inside Agent Opus with a few prompts to create the video that you can see above.
It is a manifestation of how much the diffusion models have evolved just in the space of 18 months. We were trying to get to this type of outcome 18 months ago, and it was a disproportionate amount of effort to get to it.
Now the effort is lower, but in my view, there is still a bit of work to add to the quality and the storytelling. I think for two-minute overviews this is fine, but for anything in more detail, you again want a bit more granularity. For that, we are going to be experimenting with Pictory to see whether that can help us.
As with any new tooling, one of the main things we are validating by logging a large number of data points is whether this type of video overview enhances engagement and clarifies the messaging. That information will emerge over time as more people use the site.
We must remember that just because you can do something doesn't mean you should. We have to be able to draw the dots to a CVO.
