A promo spot without a camera crew: how a whole spot is made with AI

A promo spot classically costs three things: money, time and compromises. The budget trims the idea, the shoot day eats a week, the voice you want does not fit the fee. What is left is a clip nobody watches to the end.

That connection has come apart over the past eighteen months. Not in theory, in practice. This piece uses a real example to show how a finished spot is made, what the workflow really costs and where the technology still fails.

AI generated
The complete spot, 42 seconds, generated entirely with AI

The test case

The example video deals with an uncomfortable subject: being insured. In several short scenes I come to grief in absurd, exaggerated ways in everyday situations. The message behind it is meant seriously. Whoever is covered does not leave their family a second problem.

No camera crew. No actors. No shoot day. No sound studio.

The subject was interchangeable. It could just as easily have been a trade, software, property or a dental practice. The mechanics stay identical: an idea that sticks, and a voice that carries it.

The workflow in five steps

1. Concept before technology

The most common mistake: people open an AI tool and start prompting. The result looks like a tool, not like advertising.

Before the first prompt comes the question every agency asks first: what should the viewer think, feel or do after 30 seconds? For the test case the answer was: "I ought to check my policy." Everything after that is only carrying out that one line.

2. Script and shot list

A spot of 30 to 60 seconds consists of five to eight shots. More gets hectic, fewer gets sluggish.

Every shot needs three entries before anything is generated:

  • Picture content, meaning what actually happens in the frame
  • Camera angle, meaning wide, close, low angle, static or moving
  • Audio, meaning which part of the text sits on this shot

That table is the real work. Build it cleanly and everything after it is generated in a fraction of the time.

3. Picture and motion

Current video models such as Seedance or comparable systems generate clips of a few seconds from text or from a starting image. That is exactly the length of one shot. For advertising that is not a drawback, it is the normal case.

Two rules from practice:

Picture first, then motion. A precisely generated still as the starting image gives noticeably more even results than a pure text-to-video prompt. You control look, colour and composition before motion comes into it.

Consistency across repeats. The same figure in eight scenes is the hardest discipline. Reference images, identical descriptions of clothing and face and one uniform style suffix in every prompt solve about eighty per cent of the problem. The edit solves the rest.

4. The voice: the underrated lever

This is where it is decided whether a spot feels cheap or expensive. Pictures are something most people manage well enough by now. Timing and delivery are not.

Modern speech models deliver studio quality in German, including stress, pauses and tone. The fault almost never lies with the model, it lies with the text. Written language sounds wrong when it is spoken.

What helps in practice:

  • Cut sentences down to speaking length. One thought, one sentence.
  • Set pauses explicitly instead of hoping commas will do it.
  • Generate three to five versions of the same line and take the best. It costs seconds.
  • Choose the voice to fit the message, not your own taste. Dry humour needs a calm voice, not an excited one.

5. Edit, music, sound design

The edit turns clips into a spot. Three things carry more weight here than any model:

  • Timing on the voice. The cut lands on the stress, not next to it.
  • Sound design. A footstep, a rustle, an impact off screen sells a scene that stays incomplete in the picture.
  • Music bed. Generative music tools supply royalty-free beds in the right length and mood.

With humour especially: the cut before the punchline is often funnier than the punchline on screen.

What that costs

That depends entirely on how long your clip is meant to be, whether you already have a script and how much reworking there is. Just ask me and we will work it through on your case.

Two legal points

Labelling. For advertising with AI-generated content the requirements in the EU are still developing. Anyone publishing commercially should check the current state rather than rely on a blog post. An open declaration never hurts in practice and heads off arguments.

Platform rules. Exaggerated, dark humour is permitted on most platforms as long as no clear violence is visible. The cut before the impact is not only the better punchline, it is also the safer route. Paid ads are held to far stricter rules than ordinary posts. Planning a toned-down version for ads saves trouble later.

Conclusion

The technology is no longer the bottleneck. Anyone can generate pictures and voices today that would have needed a large budget not long ago.

The bottleneck is the idea. An AI spot without a concept looks like an AI spot. An AI spot with a concept looks like advertising somebody paid a lot for.

Go through the whole workflow once and you quickly understand where your own work actually sits: not only in the prompting, but in the thinking before it and the editing after it.

Want a spot like that for your business? Write to me.

Back to the blog

Portrait photo of the owner of KInnovativ

Not sure what you need?

Write me two sentences about what you are planning. I tell you honestly what makes sense, even if it is less than you expected. It costs you nothing.