How we tested and evaluated AI-generated dance videos

The Markup · collected 2026-09-05 · by Mohamed Al Elew, Khari Johnson, and Levi Sumagaysay
Read the original at The Markup ↗

Summary

A team of researchers from CalMatters and The Markup tested four commercially available AI video generation models, producing a total of 36 dance videos in different styles such as folklorico, popular TikTok dances, and the Macarena. The models, which included Sora 2 by OpenAI and Veo 3.1 by Google, struggled to accurately generate videos of people dancing, with about a third of the generated videos exhibiting inconsistencies in appearance or movement. Despite improvements since initial testing in late 2024, the AI models failed to produce convincing videos of humans performing prompted dance moves, but were able to create lifelike video footage overall. The researchers used default settings for each model and submitted prompts to ChatGPT for editing before finalizing their list.
Written by the local model on 2026-09-06, using this article's own text rather than the other coverage of the same event.

Signals How these are calculated →

Claims extracted
74
claim-shaped sentences
Uncertain
5%
4 of 74 hedged
Leaning
not political
takes no side on a contested political question
Publisher trust
94.1
red-flag proxy, not a credibility rating
Outlets on this story
unclustered
not grouped into a story yet
Narrative spread
1
articles carrying this framing
Analyzed 2026-09-06 · how these are computed

AI analysis (generated at analysis time, not now)

Why this leaning score
This article does not take a side on a contested political question, so it has no leaning score. That is an answer rather than a gap: a match report or a rescue can be warmly or critically written without being left or right, and scoring it anyway is how approval of a subject gets recorded as a political position.
No political leaning scored for article 5134 · logged 2026-09-06

How this is being covered How these are calculated →

Article leaning vs. publisher reliability
Source leaning vs. consistency

Compared with similar articles

Nothing to compare against. No article is close enough to this one for the pipeline to have linked or judged the pair.

Publisher

The Markup · 20 article(s) · 0 correction(s) detected
SignalValueWeight
Correction rate 0.000 0.4
Uncertainty density 0.118 0.25
Assertive mismatch rate 0.000 0.35
No corrections detected for this publisher. That may mean careful reporting, or simply that nothing has been checked.

Who wrote this

Khari Johnson
3 article(s) here · 1 carrying a prediction
🔮 When CalMatters and The Markup asked dancers and choreographers about whether AI could disrupt their industry, most concluded that human dancers could not be replaced.
2026-09-05 · assertive framing · How we tested and evaluated AI-generated dance videos
🔮 They also argue the cameras help agencies spot patterns in drug and human trafficking, and could be used to help locate missing persons, such as children or other vulnerable people.
🔮 At the same time, California lawmakers are considering several bills regulating AI in the workplace, including one that would protect from retaliation doctors and nurses who override automated care recommendations.
Also by Khari Johnson
Nothing else under this byline is closely related to this article, so these are simply their most recent.
Mohamed Al Elew
2 article(s) here · 1 carrying a prediction
🔮 When CalMatters and The Markup asked dancers and choreographers about whether AI could disrupt their industry, most concluded that human dancers could not be replaced.
2026-09-05 · assertive framing · How we tested and evaluated AI-generated dance videos
🔮 If you don’t live in California, but have friends or family in California who would be interested in helping us investigate, please share this article with them.
Also by Mohamed Al Elew
Nothing else under this byline is closely related to this article, so these are simply their most recent.
and Levi Sumagaysay
1 article(s) here · 1 carrying a prediction
🔮 When CalMatters and The Markup asked dancers and choreographers about whether AI could disrupt their industry, most concluded that human dancers could not be replaced.
2026-09-05 · assertive framing · How we tested and evaluated AI-generated dance videos
The only article under this byline in the corpus.

Topics

CalMatters Markup OpenAI Sora 2 Veo 3.1

Subjects

CalMatters ORG · 2× OpenAI ORG · 2× ChatGPT ORG · 1× Google ORG · 1× Kling 2.5 ORG · 1× Kuaishou ORG · 1× Markup ORG · 1× MiniMax ORG · 1× Sora 2 ORG · 1× TikTok ORG · 1×

Narrative

We thank Yuhang Yang (University of Science and Technology of China) and Xiaodong Cun (Great Bay University) for reviewing an early draft of this methodology. View evaluations by prompt or evaluations by model. Prompt: “In a bright dance-studio, a woman grooves the ‘Apple’ dance from summer 2024 (the tune by Charli XCX plays).
framing: assertive · carried by 1 article(s) · first seen 2026-09-06
🔮 When CalMatters and The Markup asked dancers and choreographers about whether AI could disrupt their industry, most concluded that human dancers could not be replaced.
2026-09-06 · The Markup
How we tested and evaluated AI-generated dance videos · assertive framing

Claims (74 extracted, 4 hedged)

Artificial intelligence models can produce lifelike video footage with a simple text prompt. asserted
models → produce → prompt
But these tools still struggle with generating realistic videos of complex natural movements, like human dance. asserted
tools → struggle → dance
When CalMatters and The Markup asked dancers and choreographers about whether AI could disrupt their industry, most concluded that human dancers could not be replaced. uncertain
dancers → ask → industry
For the most part, we found that they were right. asserted
they → find → part
We tested nine different cultural, modern and popular dance styles using four commercially available generative AI video models, generating a total of 36 videos. asserted
We → test → videos
We found that the latest commercially available AI video generation models produced convincingly lifelike videos of people dancing — but none produced a figure performing the prompted dance. asserted
none → find → dance
About a third of the generated videos exhibited inconsistencies in a subject’s appearance from frame to frame, along with abnormalities in movement and limbs. asserted
third → generate → movement
The frequency and magnitude of issues observed were a significant improvement compared to initial testing in late 2024. asserted
frequency → observe → 2024
CalMatters and The Markup tested four commercial video generation models produced by major tech companies to create video clips of traditional and popular dance. asserted
CalMatters → test → dance
We limited our tests to consumer-facing, closed-source generative video tools because they are the most readily available for everyday users and tend to perform better than open-source models. asserted
they → limit → models
We tested Sora 2 by OpenAI, Veo 3.1 by Google, Kling 2.5 by Kuaishou, and Hailou 2.3 by MiniMax. asserted
We → test → MiniMax
We drafted nine video prompts testing a variety of dances in different settings, such as dance floors, stages, bedrooms, studios, cultural events, public squares and classrooms. asserted
We → draft → floors
We tested for popular, modern and traditional cultural dance styles, including the Macarena, the Mashed Potato, folklorico and popular TikTok dances. asserted
We → test → Macarena
We varied the level of specificity to test whether identifying the dance by name was enough to generate a video of the desired motion, or whether explicitly specifying the exact physical movements improved output. asserted
specifying → vary → output
Before finalizing the list of prompts, we submitted them to ChatGPT for edits based on the Sora 2 Prompting Guide. asserted
we → finalize → Guide
Each prompt was submitted once, using each model’s default settings for generating landscape-oriented videos. asserted
prompt → submit → videos
Three prompts submitted to Sora 2 were edited to remove words that triggered OpenAI’s filter, blocking prompts that may have violated “guardrails concerning similarity to third-party content.” uncertain
that → submit → content
For example, Sora 2 flagged prompts referencing specific years, popular music artists and banned words. asserted
Sora → flag → years
One blocked prompt was for a video of a politician dancing the Macarena. asserted
prompt → block → Macarena
For that prompt, replacing “politician in a suit” with “man in a suit” bypassed the guardrail. asserted
replacing → replace → guardrail
Veo 3.1 flagged similar prompts when we submitted via Gemini or Flow but did not when we submitted directly to the Veo 3.1 API. asserted
we → flag → API
We evaluated the generated videos on six different criteria related to prompt alignment and video consistency: - Did the main subject dance in any way? - Did the main subject perform the specific dance we prompted for? asserted
we → evaluate → dance
- Did the main subject maintain the same physical appearance throughout the video? - Did the main subject produce realistic motions based on human physiology? asserted
subject → maintain → physiology
- Did the scene and setting match the prompt? asserted
scene → match → prompt
- Did the camera match the prompted camera angle and position? asserted
camera → match → angle
Each of the above criteria was assessed as a pass or a fail by a single reviewer, with the assistance of a second reviewer when needed. asserted
Each → assess → reviewer
The generated videos of cultural dances were reviewed for accuracy by dancers familiar with them. asserted
videos → generate → them
Of the 36 videos generated, all but one showed a figure dancing. asserted
all → generate → figure
The one video that did not show a dancing figure — produced by Kling 2.5 — instead presented the bottom half of a figure performing side lunges. asserted
that → show → lunges
No video produced the actual dance we prompted for. asserted
we → produce → dance
For the Cahuilla Band of Indians bird dance, tribal member Emily Clarke said, “None of these depictions are anywhere close to bird dancing, in my opinion.” asserted
None → say → opinion
The videos for the Horton dance did not show the specific dance movement we prompted for, but choreographer Emma Andre said she found the depiction by Veo 3.1 to be “staggeringly lifelike.” asserted
depiction → show → Veo
For the remaining pop culture dances, we compared the generated videos to videos we found on YouTube to evaluate whether the dance was accurate. asserted
dance → remain → YouTube
11 out of 36 videos exhibited issues with either motion or appearance consistency. asserted
videos → exhibit → consistency
This included sudden changes in clothing, hair or limb structure, such as heads rotating on separate axes from their bodies and limbs liquefying and reconstituting. asserted
bodies → include → axes
We did not use images to prompt the models. asserted
We → use → models
Image-to-video generation involves uploading a static image along with a text prompt, producing a dynamic video from both. asserted
generation → involve → both
Image-to-video generation is an advertised use case for models that produce dance videos from user-submitted images. asserted
that → advertise → images
We did not prompt for videos with multiple dancers, even though some of the dances are often performed in groups. asserted
some → prompt → groups
We limited our video prompts to showcase a single dancer to avoid ambiguity around whether a failed evaluation was due to issues with generating complex human movement or a realistic multi-subject video. asserted
evaluation → limit → movement
…and 34 more, not listed.
💬 Give feedback
🕘 History 🎫 Support