
Table of Contents
Hey, It's Me, Your Favorite Conference Chaos Agent 👋
Quick confession: I was supposed to be the "expert" demoing AI video workflows at Foxwell Founders in Portland last week.
Plot twist: Google dropped Veo 3.1 literally 24 hours before my session.
So there I was, in front of 100+ people, frantically Googling features while pretending I knew what I was doing. Professional? Debatable. Effective? Somehow, yes.
The result: I generated a fully-produced airport B-roll in under 10 minutes. One guy literally asked if I'd fly to France to make videos for his brand full-time.
(Still thinking about it, honestly.)
Here's everything I learned while winging it in real-time.

Here is the end result
TL;DR - The 30-Second Version
If you're short on time, here's what you need to know:
Veo 3.1 is Google's video generator that just got start/end frames + extensions beyond 8 seconds
You can now create professional B-roll in 20 minutes for the cost of lunch
The workflow: Stock photo → Midjourney → Nano Banana → Seedream → Veo 3.1 → Done
Total monthly cost for unlimited requests: ~$200-300 vs. $10k+ for a production team
Still here? Let's dive in!
WTF is Veo 3.1 and Why Am I Yelling About It?
Veo 3.1 is Google's video generation model, and it just got a major update. Like, the day before the conference. Thanks, Google.
Here's what's new:
You can now use start AND end frames (finally, some control)
You can extend videos beyond 8 seconds (game changer)
Veo 3 fast is still unlimited
NOT in Gemini. Everyone asked about Gemini. It's in there, but you can’t volume produce so it makes no sense.
At the conference, someone asked if I had a subscription to "all these tools" and I had to admit I'm basically just hemorrhaging $10-20 every time my Nano Banana credits dip below $10. Living the dream.
The Full Workflow (Or: How I Made Airport B-Roll While People Judged My File Organization)
Alright, this is the part where I actually looked semi-competent.
Here's the full process I demoed live:
Step 1: Find Your Starting Image (Or Make One Like a Fancy Person)

Staring Image
Start with a stock photo. Shutterstock, whatever. Or if you want to be extra, generate one in Midjourney.
For the demo, I used a basic stock photo of a woman in an airport terminal. Nothing special.
Pro move: Upload it to imageprompt.org
This tool reverse-engineers Midjourney prompts from any image. It's trained on Midjourney's style, so when you paste that prompt back into Midjourney, your success rate is way higher than just asking ChatGPT.
Could you skip this and go straight to ChatGPT? Yeah, absolutely.
Will it work as well? Probably not. Your call.
Step 2: Generate Your Hero Image in Midjourney
Take that prompt from imageprompt.org and drop it into Midjourney.

Midjourney Render
Generate your base image, person, setting, lighting, vibe, all of it.
This is your foundation. If this looks weird, everything downstream will look weird. Spend the time here.
I generated a clean shot of a woman walking through an airport terminal with a suitcase. Looked pretty good. Definitely not AI-weird.
Step 3: Product Swap with Nano Banana
Now you need to swap out the generic suitcase for your actual product. This is where Nano Banana comes in.
Here's the process:
Take your image and put a red rectangle around the product

Prompt GPT5 with both that image and the product shot

(continued) Tell it: "I need a Nano Banana prompt to swap out the suitcase in the red rectangle with [your product]. Please include all copy from the product. I am not including the red rectangle in the prompt."
Copy that prompt, paste it into Nano Banana
Add your image from Nano Banana + the product image
CRITICAL: DO NOT include the red rectangle in the image you upload to Nano Banana. The rectangle is just for ChatGPT to understand what you're replacing.

Boom. Now you have your actual product in the image.
Step 4: Make the Person Look Less Like a Fever Dream
Here's the dirty secret: Nano Banana is amazing at products, but sometimes the people come out looking... funky.
So you take that product-swapped image and run it through Seedream to fix the person.
Upload your image, use it as a reference, generate a few variations, and boom suddenly your person looks human again.
Here is the prompt to use (works 60% of the time, all the time):IMG_3984.CR2. please upscale these people, showcase pores etc and make them feel less plastic

After

Before
This step is optional if your Nano Banana output already looks good, but like... it usually doesn't. So don't skip it.
Step 5: Generate Prompt in GPT5 and Then Video in Veo 3.1
Okay, this is where it gets fun.
Take your image and put into that same ChatGPT chat you’ve been using, and ask it for a veo 3 JSON prompt with one or two actions you are after.
In the below example, I just asked to have her walking down the tarmac.

From there:
Upload your final hero image as the starting frame
Paste your JSON prompt.
Hit generate (I used Veo 3.1 fast because it's free and honestly good enough)
Here's the thing nobody tells you: You're going to reprompt this 3-4 times, with four outputs per request.
In my experience 1 out of 10 videos are solid.
Don’t worry so much about prompt engineering: I am sure your prompt is fine!
At the conference, I got a couple weird outputs where the video was wonky, but then one came through that was chef's kiss perfect.
Time investment: Maybe 5 minutes if you're reprompting a few times.
Step 6: Start + End Frame Magic (The New Feature That Actually Slaps)
This is the new Veo 3.1 feature that makes it legitimately useful.
You can now give it a start frame and an end frame, and it'll generate the video that transitions between them.
So I wanted to show the woman putting her suitcase in the overhead bin. Here's how:
Using the first frame photo you generated, create an end frame within Nano Banana. In this example, I used the same woman, same suitcase, but requested that she put it in the overhead bin.

Go back to Veo 3.1 (I didn’t go back to Seedance here, because I thought the quality of Nano Banana was sufficient enough to use as an end frame).
Upload both images: start frame on the left, end frame on the right
Use ChatGPT to generate the transition prompt: "Please give me a JSON prompt to get from this starting frame to this ending frame"

Paste your JSON into Veo 3.1 with the start and endframe and generate!
The result? A smooth video of her walking through the tarmac AND putting the suitcase away.

So Here's The Thing...
I'm writing this from the Skamandia Lodge in Portland, about to fly back to NYC in a few hours, and I keep thinking about that France offer.
Not because I actually want to move to France (but I also totally would...).
But because it perfectly captures what just happened at Foxwell.
I showed up with a workflow I'd barely tested.
Demoed it live while Googling half the features.
Made airport B-roll in 10 minutes.
And someone immediately tried to hire me.
Here's what I know:
Your competitors aren't doing this yet. That's the whole point.
While everyone else is stuck in production meeting hell or waiting on assets that will never come, you could be shipping 10 variations before lunch.
This isn't about keeping up.
This is about getting so far ahead they don't know what hit them.
The window is now. In six months, everyone will be doing this. And you'll be the person who waited for permission while someone else ate your lunch.
If you try this and get stuck, just reply. I have an embarrassing number of screenshots and zero shame about my messy process.
Party on Garth,
Will
P.S. At one point during the demo I couldn't find my files because I forgot they were organized alphabetically. A room full of marketers watched me discover how the alphabet works in real time. And they STILL asked for the workflow. That should tell you something.
P.P.S. Boarding soon. If you're on my flight and want to talk AI workflows, I'll be the guy frantically generating B-roll on airplane wifi. Come say hi.