Two of us have spent three weeks making FAUST, a 3–5 minute short set in Manhattan high finance in 1987. All the visuals are AI-generated video, the voice and score are AI audio, and every spoken line is Goethe's own verse in Bayard Taylor's 1870 translation (public domain), subtitled. Trailer is in the comments. What we learned, in the order it mattered:
- Character sheets before anything else. Lock the face and the voice follows. Skipping this is what makes AI shorts look like a slideshow of strangers.
- Write the cut before generating. Every shot in the trailer existed as a line in a shot list first. "Generate and hope" was our most expensive mistake.
- One color grade per scene, one grain pass over the whole film. Clips made a week apart never sit together otherwise.
- Short spoken blocks. Verse lipsyncs surprisingly well, but only in pieces of a few seconds. Whole stanzas fall apart.
- Use the LLM as a librarian, not a writer. We didn't let it write a single line. We used it to mine the 1870 translation for the passages we needed and to verify every line against the source with line numbers, because a paraphrase that sounds like Goethe is worse than nothing.
- Silence is a cue. The best moment in the trailer is two seconds of nothing.
What it didn't solve: the pronunciation of archaic words drifts between takes ("thou", "hath", names), and once a scene is generated to a voice line, re-recording that line means re-generating the scene.
Open question for anyone who's tried period language or verse in AI film: did it survive contact with an audience that has to read?
Source: r/ChatGPT · by /u/gluecksbaerchie1