A client asked me to produce a walkthrough video for their company store a few weeks ago. When I sent it over, they were elated. Then they asked me something that could only be asked in 2026: "Could we use ElevenLabs or something for the voiceover next time?"
That question is the whole reason for this post. Before I got into print and promo, my career was in video production, so this kind of work is home turf for me. But you don't need a video background to pull it off. The framework carries most of the weight, and the tools have gotten good enough to close the rest of the gap.
I've been making these videos since my distributor days. Company store tutorials, demos of what a store could actually do, that kind of thing. The goal hasn't changed. You want something clean a customer can follow. What's changed is how much of the work you have to do yourself.
The framework is the same as it's always been: talking script, then shot list, then VO track, then video capture, then a final edit. That's it. It looks like a lot written out, but none of the steps are daunting, and every tool I'll mention below exists to make one of them faster.
Here's how I used to run that framework, and how I'd run it today.
The way it used to work
Everything started with a talking script. Once I knew what the video needed to communicate, I'd write the script by hand and build a shot list right alongside it. The two grew together. If I was writing the part about logging in, I'd make a note in the shot list to get a clean capture of the login screen. Script on one side, the shots I needed on the other.
Then I'd send that script to a voiceover artist I worked with regularly. This is the part people are surprised by. It wasn't expensive. Maybe $100 for a two-minute script, and I was really happy with the results. He had the voice, he had the studio, and he did this for a living. For a hundred bucks, it added so much to the production value.
When the VO track came back, I'd record my screen. I'd play the track (usually on a slower speed) and drive through the app doing whatever the narration described. Sometimes that was one clean session, sometimes a couple, depending on how many screens I needed and how much I fumbled. Then I'd pull the footage into Final Cut or Premiere for light cleanup. Trimming dead gaps, retiming a section here and there. Nothing heavy. At the end I'd drop in a logo bumper I'd built in After Effects for the distributorship.
That was the whole thing. Script and shot list, voiceover, capture, light edit, logo. The order mattered, and it still does.

Where AI can help
Start with the script. I don't stare at a blank page anymore. I'll talk through the goal of the video the same way I'm talking to you here and get a solid first draft back in a couple of minutes.
There's a better trick I've started leaning on, though. Instead of writing anything first, do a rough run of the demo and narrate it out loud as you click, like you're showing a coworker over their shoulder. Screen record that scratch pass and don't worry about it being clean. Then hand the transcript to an AI like Claude and ask it to turn your rambling into a tight talking script and pull a shot list out of it. Because you narrated what you were doing, the AI can read off which screens you moved through and build the shot list from that. You get the script and the shot list out of one messy pass, which is a lot easier than composing both from a blank page.
Whichever way you start, I don't ship the AI's draft as is. I run it through Humanizer (a Claude skill) so it sounds like us and not like a robot read the manual. The AI gets you off zero. You still do the shaping.
Voiceover is where that client question comes in. You've got three real options, and they're for three different situations. Your own voice when personality matters and you want people to know it's you. A pro when it's going in front of the market and needs to be polished. And ElevenLabs when you're short on time or budget. For natural-sounding narration that's fast to revise, ElevenLabs is a good fit. The revise part is underrated, too. Change one line of script and you regenerate just that line, instead of booking another recording session.
ElevenLabs will also clone a voice. You give it one to five minutes of audio for a quick clone, or a longer clean recording for a better one. For walkthrough videos specifically, a clone of your own voice is a pragmatic middle ground. You keep a consistent sound across a whole library of videos, and you can patch a changed line without setting everything up again. I'd still want a real voice on anything where the personality is the point, but I like the idea of cloning myself for the routine stuff.
For capture, I use Screen Studio, and I'd honestly use it for almost everything if I could. It zooms in on your button clicks on its own, highlights the part of the screen you're pointing at, and it'll mask confidential info or a logo when you need it to. That removes a whole category of manual keyframing that used to eat my time. I run my mic through Krisp while I'm at it, because its virtual mic does a scary good job of making a normal room sound like a studio.
You'll almost always want a light edit at the end. On my most recent video I stayed in Screen Studio for the capture but dipped into Premiere for some finishing tweaks and timing. That's normal. Premiere and Descript both do text-based editing now, where you edit the video by editing a transcript of it. Delete a sentence in the text and the matching video comes out with it. For the light trimming and gap-cutting I described earlier, that's a real time saver.
Last piece is the logo motion at the end. I built those in After Effects, but most distributors don't need to. Go to Envato, find a logo animation you like, drop your logo in, done.
The part AI doesn't change
AI speeds up making each piece. It doesn't change the order you make them in.
I've tried to shortcut that order, and it doesn't work for me. What I really want is to record the voiceover and capture the screen at the same time, but I can't do both well at once. Reading a script off the screen while driving the mouse and paying attention to what the app is doing is genuinely hard. I have a teleprompter, and it doesn't solve this, because when I overlay the script on top of my screen I can't click the applications behind it. The teleprompter swallows the clicks. That's an OS limitation, not a me problem, and I haven't found a clean way around it.
So I do what I always did. Run the framework in order: script, shot list, VO, capture, final edit. Separate steps, one after the other. AI just makes each step faster to produce.
Where the time goes now
AI speeds up making each piece, but it doesn't change the order you make them in. Script, shot list, VO, capture, edit, same as it ever was. What's actually changed is where my hours go. I spend more time on the story and less time pushing buttons. Figuring out what the video needs to say, and in what order, is the part that makes a walkthrough good. The mechanics used to take up most of my time. Now they don't.
If you want to make demos for your own company store, don't get stuck shopping for the perfect stack. Write a strong script and a shot list to match. Everything else follows from that.



