The Markup
· collected 2026-09-05 · by Mohamed Al Elew, Khari Johnson, and Levi Sumagaysay
A team of researchers from CalMatters and The Markup tested four commercially available AI video generation models, producing a total of 36 dance videos in different styles such as folklorico, popular TikTok dances, and the Macarena. The models, which included Sora 2 by OpenAI and Veo 3.1 by Google, struggled to accurately generate videos of people dancing, with about a third of the generated videos exhibiting inconsistencies in appearance or movement. Despite improvements since initial testing in late 2024, the AI models failed to produce convincing videos of humans performing prompted dance moves, but were able to create lifelike video footage overall. The researchers used default settings for each model and submitted prompts to ChatGPT for editing before finalizing their list.
Written by the local model on 2026-09-06,
using this article's own text rather than the other coverage of the
same event.
Artificial intelligence models can produce lifelike video footage with a simple text prompt.
asserted
models → produce → prompt
But these tools still struggle with generating realistic videos of complex natural movements, like human dance.
asserted
tools → struggle → dance
When CalMatters and The Markup asked dancers and choreographers about whether AI could disrupt their industry, most concluded that human dancers could not be replaced.
uncertain
dancers → ask → industry
For the most part, we found that they were right.
asserted
they → find → part
We tested nine different cultural, modern and popular dance styles using four commercially available generative AI video models, generating a total of 36 videos.
asserted
We → test → videos
We found that the latest commercially available AI video generation models produced convincingly lifelike videos of people dancing — but none produced a figure performing the prompted dance.
asserted
none → find → dance
About a third of the generated videos exhibited inconsistencies in a subject’s appearance from frame to frame, along with abnormalities in movement and limbs.
asserted
third → generate → movement
The frequency and magnitude of issues observed were a significant improvement compared to initial testing in late 2024.
asserted
frequency → observe → 2024
CalMatters and The Markup tested four commercial video generation models produced by major tech companies to create video clips of traditional and popular dance.
asserted
CalMatters → test → dance
We limited our tests to consumer-facing, closed-source generative video tools because they are the most readily available for everyday users and tend to perform better than open-source models.
asserted
they → limit → models
We tested Sora 2 by OpenAI, Veo 3.1 by Google, Kling 2.5 by Kuaishou, and Hailou 2.3 by MiniMax.
asserted
We → test → MiniMax
We drafted nine video prompts testing a variety of dances in different settings, such as dance floors, stages, bedrooms, studios, cultural events, public squares and classrooms.
asserted
We → draft → floors
We tested for popular, modern and traditional cultural dance styles, including the Macarena, the Mashed Potato, folklorico and popular TikTok dances.
asserted
We → test → Macarena
We varied the level of specificity to test whether identifying the dance by name was enough to generate a video of the desired motion, or whether explicitly specifying the exact physical movements improved output.
asserted
specifying → vary → output
Before finalizing the list of prompts, we submitted them to ChatGPT for edits based on the Sora 2 Prompting Guide.
asserted
we → finalize → Guide
Each prompt was submitted once, using each model’s default settings for generating landscape-oriented videos.
asserted
prompt → submit → videos
Three prompts submitted to Sora 2 were edited to remove words that triggered OpenAI’s filter, blocking prompts that may have violated “guardrails concerning similarity to third-party content.”
uncertain
that → submit → content
For example, Sora 2 flagged prompts referencing specific years, popular music artists and banned words.
asserted
Sora → flag → years
One blocked prompt was for a video of a politician dancing the Macarena.
asserted
prompt → block → Macarena
For that prompt, replacing “politician in a suit” with “man in a suit” bypassed the guardrail.
asserted
replacing → replace → guardrail
Veo 3.1 flagged similar prompts when we submitted via Gemini or Flow but did not when we submitted directly to the Veo 3.1 API.
asserted
we → flag → API
We evaluated the generated videos on six different criteria related to prompt alignment and video consistency:
- Did the main subject dance in any way?
- Did the main subject perform the specific dance we prompted for?
asserted
we → evaluate → dance
- Did the main subject maintain the same physical appearance throughout the video?
- Did the main subject produce realistic motions based on human physiology?
asserted
subject → maintain → physiology
- Did the scene and setting match the prompt?
asserted
scene → match → prompt
- Did the camera match the prompted camera angle and position?
asserted
camera → match → angle
Each of the above criteria was assessed as a pass or a fail by a single reviewer, with the assistance of a second reviewer when needed.
asserted
Each → assess → reviewer
The generated videos of cultural dances were reviewed for accuracy by dancers familiar with them.
asserted
videos → generate → them
Of the 36 videos generated, all but one showed a figure dancing.
asserted
all → generate → figure
The one video that did not show a dancing figure — produced by Kling 2.5 — instead presented the bottom half of a figure performing side lunges.
asserted
that → show → lunges
No video produced the actual dance we prompted for.
asserted
we → produce → dance
For the Cahuilla Band of Indians bird dance, tribal member Emily Clarke said, “None of these depictions are anywhere close to bird dancing, in my opinion.”
asserted
None → say → opinion
The videos for the Horton dance did not show the specific dance movement we prompted for, but choreographer Emma Andre said she found the depiction by Veo 3.1 to be “staggeringly lifelike.”
asserted
depiction → show → Veo
For the remaining pop culture dances, we compared the generated videos to videos we found on YouTube to evaluate whether the dance was accurate.
asserted
dance → remain → YouTube
11 out of 36 videos exhibited issues with either motion or appearance consistency.
asserted
videos → exhibit → consistency
This included sudden changes in clothing, hair or limb structure, such as heads rotating on separate axes from their bodies and limbs liquefying and reconstituting.
asserted
bodies → include → axes
We did not use images to prompt the models.
asserted
We → use → models
Image-to-video generation involves uploading a static image along with a text prompt, producing a dynamic video from both.
asserted
generation → involve → both
Image-to-video generation is an advertised use case for models that produce dance videos from user-submitted images.
asserted
that → advertise → images
We did not prompt for videos with multiple dancers, even though some of the dances are often performed in groups.
asserted
some → prompt → groups
We limited our video prompts to showcase a single dancer to avoid ambiguity around whether a failed evaluation was due to issues with generating complex human movement or a realistic multi-subject video.
asserted
evaluation → limit → movement
…and 34 more, not listed.