Think of the inputs as a casting plan: two images for identity and one video for performance. Set the roles deliberately, then inspect how they interact in the generated clip.
Inspect the Two-Person Reference
Choose the Hotel Lobby clip from the reference video options. Note which performer stands on each side and where their gestures overlap; those details matter when assigning the images.
Assign an Image to Each Performer
Upload the intended left character in the first image field and the intended right character in the second. Prefer one clearly visible subject per image over collages or group photos.
Confirm the Mapping and Output Length
Review the preset replacement instructions and adjust the performer descriptions when needed. Selecting the reference option syncs its duration to the model control; review the available length and quality settings before generating.
Render and Inspect the Interaction
Generate the Rap Duo AI clip, then watch the speaking turns, identity consistency, and overlapping movement. If a role drifts, change that input or clarify the mapping before another attempt.
Inspect the Two-Person Reference
Choose the Hotel Lobby clip from the reference video options. Note which performer stands on each side and where their gestures overlap; those details matter when assigning the images.
Assign an Image to Each Performer
Upload the intended left character in the first image field and the intended right character in the second. Prefer one clearly visible subject per image over collages or group photos.
Confirm the Mapping and Output Length
Review the preset replacement instructions and adjust the performer descriptions when needed. Selecting the reference option syncs its duration to the model control; review the available length and quality settings before generating.
Render and Inspect the Interaction
Generate the Rap Duo AI clip, then watch the speaking turns, identity consistency, and overlapping movement. If a role drifts, change that input or clarify the mapping before another attempt.
A duet asks the model to track two identities across one changing performance. Rap Duo AI works best when the references make the relationship between each image and performer explicit.

AnyMimic gives the character images and reference video different jobs. The replacement prompt uses the pictures for appearance while asking the model to preserve the video setting and performance.

Use the reference as a controlled starting point and change one casting choice at a time. This makes it easier to compare outputs and identify which input needs improvement.