郭立 (leeguoo)

Having Claude Opus 5.5 Code a 3D Promo Video: 1,250 Lines of Code, Not a Single Manually Animated Frame

With a single request, Opus 5.5 wrote the director’s notes, three.js scenes, a music synthesizer, and subtitles in Claude Code; reviewed extracted frames itself to fix bugs; then used the open-source motion-use to render a 102-second, 6,120-frame finished video locally.

ON THIS PAGE

First, watch the finished video: 102 seconds, 60 fps, with Chinese voiceover and subtitles.

Every object on screen, every camera movement, and every note was computed by code. No AE, no keyframe animation, no asset library, and no text-to-video model. The thing that did it was Claude Opus 5.5 running in Claude Code, with rendering handled by our own open-source motion-use.

What I Gave It

One request: use three.js to make a one- to two-minute promo video for motion-use, with voiceover, explaining how it works and what it can do. During the process, I only gave feedback once: the first frame was black, and the README player would use it as the cover, which would look bad. It changed frame 0 into a complete opening shot.

What It Did

Verify first, then write the script. For every claim about motion-use in the voiceover, it first looked through the repository source code for supporting evidence: how the renderer seeks frame by frame, how fonts are subsetted, and what the acceptance checks test. The sources are recorded in the director’s notes.

The voiceover determined the timeline. It first wrote 9 voiceover segments, generated speech, measured the actual duration of each segment, and then set the shot windows based on those durations. motion-use has strict rules: if the voiceover exceeds the shot, rendering fails directly. It does not truncate it for you.

One line tells the whole story. The opening request falls to the ground and melts into a glowing timeline. Then storyboard cards stand up along it, a playhead scrubs back and forth across it, frames rise from it and are captured in parallel, voiceover blocks are fitted into its shot windows, and the finished video passes through acceptance scanning rings along it. The little felt monster from the previous work serves as the guide.

The 3D scene is about 1,000 lines of three.js: a mirror floor with a black lacquer layer, felt material using sheen plus a translucent fuzzy shell layer, Bloom for glow, and every pose as a pure function of time, so any frame at any moment can be computed independently.

The music and sound effects are Python-synthesized waveforms: a 96 BPM base, with typing, shutter, error, and pass notification sounds aligned to the timing of on-screen events. Melodies only appear in gaps between voiceover segments.

Problems It Caught Itself

This was the most interesting part. It did not just hand things over after writing them; it repeatedly extracted frames, inspected images, and made fixes:

  • In the first version, almost every segment was framed too far away, the text was too small, and the storyboard cards were blown out into white blocks by Bloom.
  • Frames on the right half of the filmstrip disappeared, but they were still visible in the floor reflection. It judged that the sorting of the translucent backing panel was covering the frames, and fixed it by locking the draw order.
  • The camera paused at every keyframe, causing the opening shot to sit still for 0.8 seconds. It changed the interpolation to be continuous.
  • In the previous video, around 30 frames in the draft had black squares. It traced the issue to individual pixels producing NaN values, which Bloom then blurred and spread into an entire block, so it added a bitwise detection cleanup step at the very start of the post-processing chain. That fix was also reused in this video.

Before delivery, motion-use decodes the entire finished video, measures the actual duration, frame rate, and loudness, and checks whether the voiceover was actually mixed in by comparing waveform correlations per shot. This video: 0 faulty frames, all 9 voiceover segments passed detection, loudness -14.5 LUFS.

What It Did Not Do

  • The voiceover is AI speech, generated with edge-tts.
  • It cannot hear audio. For the music, it only measured spectrum and loudness; the listening experience still needs to be judged by a human.
  • For now, there is only a Chinese version.

Try It Yourself

motion-use is open source. After installing it, have your coding agent read its skill, and it can use the same workflow to make videos:

bash
curl -fsSL https://raw.githubusercontent.com/leeguooooo/motion-use/main/install.sh | sh
motion-use doctor

The full source code for this video is in examples/motion-use-promo, and the Bilibili version is here.

More in this category · AI & Agent

next →
Making Claude Code and Codex talk to each other — even across two computers

Comments

Replies are public immediately and may be moderated for policy violations.

Max 1000 characters.