Started with the wrong problem, caught it before building
Scoped this first as batch generation, photo and audio in, video file out, competing with D-ID's Creative Studio product. Realized that's a different system entirely from 'a photo that talks back in a live conversation.' A rendered clip and a real-time reply don't share an architecture: one is request-then-wait, the other needs a GPU held for the call's duration and a model fast enough to render as the audio plays.