Every frame is real. It's the wrong person.
Most of the argument about AI and video is about whether the pictures are real. Generated or filmed. Synthetic or archive.
I’ve spent this year on something that sits on the archive side of that. You give it a subject. At several points along the way it goes to a news archive and pulls real footage that bears on it, and that footage is the factual basis for what it writes at that point. The synopsis, the scenes, the narration. Once that’s settled it works out what to actually show on screen while the narration plays. Every clip in the finished video is real material, with a date, a location, and a reference number you can look up.
The team went further than just using real footage. There’s a refusal gate. If there’s nothing factual in the archive about the subject, it stops and says so. It doesn’t produce a video anyway. I won’t go into how it decides, but the behaviour is the point: the default everywhere else is to give you something rather than nothing, and this gives you nothing. It’s there to stop the thing writing a story when the archive holds nothing factual to write one from. That part works.
I was pleased with that. Then I watched a video where the narration was about one public figure and the footage was of a different person with the same name.
The example is simplified. The failure is real and I’ve watched it happen more than once.
Nothing was wrong in its world
I described it at the time as the pictures pulling the script. That was backwards. Nothing went wrong in the script at all.
Here’s the bit I had wrong. Each stage goes and gets its own footage when it runs. That footage becomes the factual basis for what that stage writes, it stays attached to that stage, and it gets handed forward to the next one. No stage re-runs a search that’s already been done, and nothing goes back and throws anything out.
Then, at the end, it works out what to put on screen. That might include all of the footage that formed the factual basis. It might include some of it. It might include none of it, and go and find other real footage that fits the narration more closely. The point isn’t that those are two different sets of video. The point is that nothing requires them to be the same set.
So it read the name in the narration. It went and found real archive footage of a person with that name. That footage fits the words extremely well, because the name matches. It used it.
Nothing was wrong in its world. It had its story. It had the narration that was going to be read out. It had found real video of the person it thought it needed to serve. Every stage did the job it was given, so nothing downstream had any reason to object.
One phrase doing two jobs
What I’d missed is that “factually based” means two different things depending on which end of the process you’re standing at.
For the text it’s a strong claim. This sentence is supported by material we actually hold. That’s the thing the refusal gate is protecting, and it’s the reason the narration doesn’t invent anything.
For the picture it’s a much weaker claim. This clip is real archive material. That’s it. It’s real, it’s dated, it’s attributable. It says nothing about whether the person in it is the person the sentence is about.
A name isn’t an identity. Two people share one, the archive holds footage of both, and a search on the name can’t tell you it picked the wrong one, because by its own measure it didn’t.
There is metadata that bears on this. Tags, descriptions of what’s visible, shot types. It’s useful and it isn’t nothing. It just isn’t enough to settle the question, because settling it needs context the metadata was never carrying.
Stopping the text from inventing things pushed the problem into the pictures, where it’s harder to see, because everything about the pictures is verifiably true.
I don’t think that’s a bug we’re one fix away from. What an archive gives you is provenance, and provenance is worth having. I’d been treating it as the thing generated footage can’t offer, and that isn’t right either. Generated media can carry provenance about how it was made, and there are standards for exactly that. What it can’t give you is a camera-made record of an event, tied to a place and a date. That’s the real difference and it’s narrower than the one I started with.
Either way it’s a guarantee about where a clip came from. It says nothing about whether that clip belongs next to this sentence.
Where the person goes
I should be straight about something. I don’t own the selection logic. I didn’t build it and it isn’t my area. My work is the review surface, the place where the system’s decisions get put in front of a person who has to accept them or reject them. So this is the view from the seat where you watch the output, and it’s a partial view.
From that seat the question is when someone looks.
If you think the risk is fabrication, you review at the end. Watch the finished thing, check nothing’s made up. Someone reviewing at the end will catch a wrong-person error too, if they happen to know the subject. It’s just the expensive place to catch it. By then it all looks right. Real footage, factual script, in sync, nothing obviously broken. Putting it right means going back and running the work again.
The cheap place is where the clip gets picked. One line of narration, the shot chosen for it, and why. That’s a small question a person can answer fast, and it’s the point where a wrong answer is still cheap to undo.
Whether the system can be made to ask that question itself, I don’t know. It’s a fair thing to put to the people who built that end of it, and I have been. Can it tell two people with the same name apart, and if it can, what would it cost to make it check before it fills the slot?
What I do know is that the version where nobody looks doesn’t work, and it doesn’t fail in a way that announces itself. It fails by handing you something that looks finished.