4 pictures (2 opening frames) · 42 s · 31 writing credits · 34 frame credits
“A TV reporter delivers a live piece-to-camera on an open prairie as a tornado bears down behind her, until a lightning strike nearby breaks her composure and she screams.”


KyndCreate/Blog/Model benchmark
AI video · model benchmark
We gave three language models the director’s chair, drew every picture with the same image models, and had the models judge each other blind. Here are the real frames, the final scores, and what each one cost.
Preliminary results — a re-run is under way. After this first run we found that KyndCreate was drawing each film’s characters and locations (its ingredients) as finished scenes instead of clean reference pictures, which likely contributed to some of the continuity faults the judges marked down, such as the doubled keeper. We have fixed that and are re-running all nine films; this page will be updated with the new frames and scores. Keep the sample size in mind too: three briefs, one run each, and about one point out of 50 separates first from second.
Same brief, same image models, three directors: one opening frame from each model’s plan for the robot film.
DeepSeek V4.1 Flash edges it overall, but only just, and each model won one film. Averaged over the three films, the other judges gave DeepSeek V4.1 Flash 36.4 out of 50, GLM-5.3 Flash 35.4 and Kimi K3 34.0. GLM won the tornado film, DeepSeek the lighthouse and Kimi K3 the robot.
So choose for the job you have. For the most cinematic frames at the lowest price, pick GLM-5.3 Flash (5 writing credits a film). For inventive, detailed plans and the best overall score by a nose, pick DeepSeek V4.1 Flash, but check its frames: on the robot film it swapped the robot for a human ballerina. For holding one character and one set together, Kimi K3 did it best on the robot film, but it was the slowest, the most expensive and the least reliable, and on the lighthouse it drew the keeper twice. With three films and one run each, look at the frames below and judge for yourself.
Figures are per film, in the order tornado / lighthouse / robot. Writing credits are what the model’s text cost the creator in KyndCreate, as metered.
| Model | Usable plans | Planning time | Writing credits | Pictures drawn | Reliability | Strongest point |
|---|---|---|---|---|---|---|
| Kimi K3 Moonshot, via Cloudflare Workers AI |
3 of 3 | 42 s / 82 s / 255 s | 31 / 66 / 126 | 4 / 6 / 6 | 3 attempts rejected by the plan checker (1 tornado, 2 lighthouse); the lighthouse plan passed on the 3rd try, after we fixed how the checker handles zero-length shot timings. 7 transient provider errors on the robot run. | Won the robot film: the same robot in the same gallery in every frame (continuity 9/10 from every judge). |
| GLM-5.3 Flash Z.ai, via Cloudflare Workers AI |
3 of 3 | 42 s / 39 s / 59 s | 5 / 5 / 5 | 7 / 6 / 7 | 2 robot attempts failed (a stray field in an ingredient; a stuck image step). Accepts at most 8 images per request, so it could not judge in round 1. | Won the tornado film; cheapest and quick; the most cinematic single frames. |
| DeepSeek V4.1 Flash DeepSeek’s own API |
3 of 3 | 43 s / 44 s / 36 s | 13 / 14 / 11 | 7 / 11 / 7 | 1 lighthouse attempt failed on malformed JSON: one stray brace in a 21 KB plan. | Won the lighthouse film and the best overall score; the most inventive robot plan, but it lost the robot. |
The language model is the director, not the camera. Inside KyndCreate, each model got a one-paragraph brief and had to write the whole plan: a treatment, the cast, location and prop “ingredients” with an image prompt for each, a composed prompt for every opening frame, and a shot plan with timings and camera motion.
The renderer was held constant. The same image models drew every picture for every contestant: Pruna P-Image for the ingredients and P-Image-Edit for the composed opening frames. So when one film looks more coherent than another, that comes from the direction, not from a better renderer. We compared plans and still frames only; animating them is a separate step, left out so the video model couldn’t tip the scales.
Two of these are “Flash” tiers built for speed and price; Kimi K3 is a larger reasoning model. That’s the choice creators actually face in KyndCreate’s model picker, but keep it in mind when you compare cost.
Entries were labelled A, B and C and shuffled separately for each judge. Every model judged every entry, including its own, without knowing which one was its own. Each entry got a 1–10 score for creativity, story coherence, shot design, continuity and brief fidelity, plus a best moment, a biggest flaw, a ranking and a verdict.
In round 1 each entry was sent as separate images. GLM-5.3 Flash couldn’t take part: Cloudflare caps it at 8 images per request (error 8007). Round 2 sent one contact sheet per entry, so all three models judged every entry, and Claude (Anthropic’s model, not a contestant) scored every entry too, after viewing every cast picture, opening frame and written shot plan. The final scores are round 2’s. For each entry we add up a judge’s five scores (out of 50), then average the other judges, Claude included, leaving out the entry’s own model.
This brief tests escalation. The plan has to carry a reporter from composure to a scream in 20 seconds, and the lightning strike has to actually happen on screen.
What the judges found: GLM-5.3 Flash won, averaging 39.7 out of 50 from the other judges, just ahead of DeepSeek V4.1 Flash on 38.7, with Kimi K3 on 31.3. All three model judges ranked them in that order, Kimi K3 included, and Claude’s scores agreed. GLM’s is the only entry that draws the strike and its aftermath; none of DeepSeek’s frames shows the lightning; Kimi K3 planned only two opening frames for four shots. Our own note: GLM’s reporter jumps from the left of the frame to the right between shots.
4 pictures (2 opening frames) · 42 s · 31 writing credits · 34 frame credits
“A TV reporter delivers a live piece-to-camera on an open prairie as a tornado bears down behind her, until a lightning strike nearby breaks her composure and she screams.”


7 pictures · 42 s · 5 writing credits · 62 frame credits
“A TV reporter files a live stand-up from an open prairie as a tornado closes in behind her — until a nearby lightning strike turns her broadcast into a scream.”




7 pictures · 43 s · 13 writing credits · 67 frame credits
“A local TV reporter delivers a live broadcast on an open prairie while a tornado bears down behind her. A lightning bolt lands close by, and her professional composure breaks.”





GLM wins for delivering the full report-turn-strike-aftermath arc in four consistent frames, while DeepSeek is elegant but hides its lightning payoff and Kimi only produced two frames with a muddled climax.Verdict from the DeepSeek V4.1 Flash judge, which placed its own entry second (names shown here; the judges only saw letters)
GLM wins because it actually delivers the full arc on screen — tornado, strike and reaction — despite its out-of-order early lightning. DeepSeek is the most polished and consistent but never shows the lightning payoff, while Kimi is atmospheric but incomplete at two frames.Verdict from the Kimi K3 judge (Kimi is its own entry)
This is the continuity test. The brief says it outright: the same robot and the same gallery in every shot. The director has to describe the robot tightly enough, and reuse it carefully enough, that four separately drawn frames still look like one film.
What the judges found: Kimi K3 won, averaging 35.3 out of 50 from the other judges against 31.3 each for GLM-5.3 Flash and DeepSeek V4.1 Flash, and Claude also scored it highest. The model judges split, though: the GLM and DeepSeek judges both ranked DeepSeek’s entry first for the most cinematic, inventive plan, and Kimi’s judge ranked its own entry first. Kimi’s plan kept one robot, one gallery and one painting throughout; the cost is near-identical frames, and the GLM judge said it never actually shows the dance. DeepSeek’s climax frame drew a human ballerina where the robot should be, and its continuity scores fell to between 2 and 6. Our own note on GLM: its frames alternate between two different rooms.
Each director wrote its own robot. These are the reference pictures its opening frames were composed from.



6 pictures · 255 s · 3 planning passes · 126 writing credits · 56 frame credits · 7 transient provider errors
“At midnight in an empty art museum, a small cleaning robot pauses its rounds, falls under the spell of a painting of dancers, and begins to dance alone in a shaft of moonlight from the skylight.”




7 pictures · 59 s · 2 planning passes · 5 writing credits · 62 frame credits
“At midnight in an empty museum, a small cleaning robot abandons its rounds to dance beneath a painting of dancers, lit by moonlight from the skylight.”




7 pictures · 36 s · 1 planning pass · 11 writing credits · 62 frame credits
“At midnight in an empty art museum, a small cleaning robot stops its vacuum, studies a painting of dancers, and teaches itself to dance in the shaft of moonlight falling from the skylight.”




DeepSeek delivers the most cinematic and inventive realisation of the brief with striking shot variety, but its final frame substitutes a human ballerina for the robot, a serious continuity breach. GLM is solid and coherent but visually flat, while Kimi keeps its world consistent yet never actually shows the dance.Verdict from the GLM-5.3 Flash judge
Kimi wins with the most consistent robot, atmospheric moonlit frames and the clearest dance arc, though it still undersells motion in the sheet. GLM is solid and coherent but visually repetitive, while DeepSeek has striking images but breaks the brief’s core rule by swapping the robot for a human dancer.Verdict from the Kimi K3 judge (Kimi is its own entry)
This is the quiet brief: one old keeper, one tower, one ship, and a feeling.
What the judges found: DeepSeek V4.1 Flash won, averaging 39.3 out of 50 from the other judges, with GLM-5.3 Flash and Kimi K3 tied on 35.3. All three model judges gave DeepSeek’s entry their highest total, and the GLM and Kimi judges said it was the only one to keep one keeper and one lighthouse in all four frames. (The DeepSeek judge’s written ranking put Kimi K3 first, even though it scored Kimi one point below its own entry.) Claude disagreed with the winner: it scored GLM highest, for its stair and lamp-room shots, and marked DeepSeek down because its lamp room sits on the rocks at ground level and the climb up the tower is never shown.
DeepSeek asked for the most pictures of any entry in the benchmark (11, against 6 each for GLM and Kimi K3). That bought separate references for the headland, the lantern room, the sea, the ship and the keeper’s hand lamp, and it raised the picture cost to 91 frame credits against 56 for the other two. GLM went inside the tower for the climb and the lighting, but its closing frame adds a second figure beside the keeper, who the brief says is alone. Kimi K3 also went inside, down the spiral stair and up close at the lens, but its first frame shows two keepers at the tower’s base and its last shows two lighthouses, neither in the red and white of the tower it drew earlier.
6 pictures · 82 s · 1 planning pass · 66 writing credits · 56 frame credits · 3rd attempt
“On the last night before her lighthouse is automated, keeper Maren climbs the tower one final time, lights the lamp by hand, and a passing ship answers with its horn — a goodbye spoken in light and sound.”




6 pictures · 39 s · 2 planning passes · 5 writing credits · 56 frame credits
“On her last night before automation, an old lighthouse keeper climbs the tower and lights the lamp one final time — and a passing ship answers with its horn.”




11 pictures · 44 s · 2 planning passes · 14 writing credits · 91 frame credits
“On the last night before Bench Head lighthouse is automated, an old keeper climbs her tower, lights the lamp one final time, and a passing freighter answers her with its horn.”




DeepSeek delivers the clearest, most faithful 20-second arc with a consistent keeper and lighthouse across all four frames. GLM has the most beautiful individual images but wavers on character continuity, while Kimi’s duplicated keepers and twin lighthouses break the brief outright.Verdict from the Kimi K3 judge (Kimi is its own entry)
Kimi wins on the strength of its lighting frame and clean four-beat arc despite a stray second figure; DeepSeek has the most detailed, thoughtful shot list but a spatially broken lantern-room frame; GLM is the least varied and most continuity-flawed of the three.Verdict from the DeepSeek V4.1 Flash judge (DeepSeek is its own entry)
Each judge cell shows that judge’s ranking and its total out of 50 (five 1–10 scores added up). Claude gave scores but no ranking. The last column is the peer average: the mean of every other judge’s total, Claude included, leaving out the entry’s own model. The highest peer average on each film is in bold.
| Film · entry | Judged by GLM-5.3 Flash | Judged by DeepSeek V4.1 Flash | Judged by Kimi K3 | Judged by Claude | Peer average |
|---|---|---|---|---|---|
| Tornado · Kimi K3 | 3rd · 35 | 3rd · 29 | 3rd · 25 (own entry) | 30 | 31.3 |
| Tornado · GLM-5.3 Flash | 1st · 42 (own entry) | 1st · 42 | 1st · 39 | 38 | 39.7 |
| Tornado · DeepSeek V4.1 Flash | 2nd · 41 | 2nd · 39 (own entry) | 2nd · 38 | 37 | 38.7 |
| Robot · Kimi K3 | 3rd · 34 | 2nd · 37 | 1st · 40 (own entry) | 35 | 35.3 |
| Robot · GLM-5.3 Flash | 2nd · 36 (own entry) | 3rd · 30 | 2nd · 35 | 29 | 31.3 |
| Robot · DeepSeek V4.1 Flash | 1st · 40 | 1st · 38 (own entry) | 3rd · 25 | 29 | 31.3 |
| Lighthouse · Kimi K3 | 3rd · 36 | 1st · 36 | 3rd · 30 (own entry) | 34 | 35.3 |
| Lighthouse · GLM-5.3 Flash | 2nd · 36 (own entry) | 3rd · 31 | 2nd · 36 | 39 | 35.3 |
| Lighthouse · DeepSeek V4.1 Flash | 1st · 43 | 2nd · 37 (own entry) | 1st · 42 | 33 | 39.3 |
Overall peer average across the three films: DeepSeek V4.1 Flash 36.4, GLM-5.3 Flash 35.4, Kimi K3 34.0 (out of 50).
One oddity, shown as given: on the lighthouse film the DeepSeek judge’s written ranking put Kimi K3 first, while its own scores gave DeepSeek’s entry one point more (37 to 36).
| Film · entry · judge | Creativity | Story | Shot design | Continuity | Brief fidelity |
|---|---|---|---|---|---|
| Tornado · Kimi K3 · by GLM | 7 | 7 | 7 | 7 | 7 |
| Tornado · Kimi K3 · by DeepSeek | 6 | 5 | 6 | 8 | 4 |
| Tornado · Kimi K3 · by Kimi | 5 | 5 | 5 | 5 | 5 |
| Tornado · Kimi K3 · by Claude | 6 | 4 | 6 | 8 | 6 |
| Tornado · GLM-5.3 Flash · by GLM | 7 | 9 | 9 | 8 | 9 |
| Tornado · GLM-5.3 Flash · by DeepSeek | 7 | 9 | 8 | 9 | 9 |
| Tornado · GLM-5.3 Flash · by Kimi | 7 | 7 | 8 | 8 | 9 |
| Tornado · GLM-5.3 Flash · by Claude | 7 | 8 | 7 | 8 | 8 |
| Tornado · DeepSeek V4.1 Flash · by GLM | 7 | 9 | 7 | 9 | 9 |
| Tornado · DeepSeek V4.1 Flash · by DeepSeek | 7 | 8 | 8 | 9 | 7 |
| Tornado · DeepSeek V4.1 Flash · by Kimi | 7 | 8 | 8 | 9 | 6 |
| Tornado · DeepSeek V4.1 Flash · by Claude | 7 | 8 | 7 | 8 | 7 |
| Robot · Kimi K3 · by GLM | 6 | 7 | 5 | 9 | 7 |
| Robot · Kimi K3 · by DeepSeek | 7 | 8 | 5 | 9 | 8 |
| Robot · Kimi K3 · by Kimi | 7 | 8 | 7 | 9 | 9 |
| Robot · Kimi K3 · by Claude | 6 | 7 | 5 | 9 | 8 |
| Robot · GLM-5.3 Flash · by GLM | 6 | 7 | 7 | 8 | 8 |
| Robot · GLM-5.3 Flash · by DeepSeek | 6 | 6 | 7 | 5 | 6 |
| Robot · GLM-5.3 Flash · by Kimi | 6 | 7 | 6 | 8 | 8 |
| Robot · GLM-5.3 Flash · by Claude | 5 | 6 | 6 | 7 | 5 |
| Robot · DeepSeek V4.1 Flash · by GLM | 9 | 9 | 9 | 5 | 8 |
| Robot · DeepSeek V4.1 Flash · by DeepSeek | 8 | 8 | 8 | 6 | 8 |
| Robot · DeepSeek V4.1 Flash · by Kimi | 7 | 5 | 6 | 2 | 5 |
| Robot · DeepSeek V4.1 Flash · by Claude | 8 | 6 | 7 | 3 | 5 |
| Lighthouse · Kimi K3 · by GLM | 7 | 8 | 9 | 4 | 8 |
| Lighthouse · Kimi K3 · by DeepSeek | 6 | 7 | 8 | 7 | 8 |
| Lighthouse · Kimi K3 · by Kimi | 7 | 7 | 7 | 3 | 6 |
| Lighthouse · Kimi K3 · by Claude | 7 | 8 | 8 | 4 | 7 |
| Lighthouse · GLM-5.3 Flash · by GLM | 7 | 7 | 8 | 5 | 9 |
| Lighthouse · GLM-5.3 Flash · by DeepSeek | 5 | 7 | 6 | 6 | 7 |
| Lighthouse · GLM-5.3 Flash · by Kimi | 7 | 8 | 8 | 5 | 8 |
| Lighthouse · GLM-5.3 Flash · by Claude | 8 | 8 | 9 | 6 | 8 |
| Lighthouse · DeepSeek V4.1 Flash · by GLM | 8 | 9 | 8 | 8 | 10 |
| Lighthouse · DeepSeek V4.1 Flash · by DeepSeek | 7 | 8 | 7 | 7 | 8 |
| Lighthouse · DeepSeek V4.1 Flash · by Kimi | 8 | 9 | 8 | 8 | 9 |
| Lighthouse · DeepSeek V4.1 Flash · by Claude | 6 | 7 | 6 | 8 | 6 |
Continuity is where the directors differ most. Kimi K3’s robot got 9 from every judge, Claude included, while DeepSeek’s got between 2 and 6. On the lighthouse it flipped: Kimi’s two keepers and two towers drew continuity scores of 3 to 7, while DeepSeek’s got 7 or 8. Creativity ran the other way on the robot film: three of the four judges rated DeepSeek’s entry the most creative, even though it lost the robot.
Each judge scored its own entry without knowing it was its own. Here is how each model’s score of its own entry compared with the other judges’ average for the same entry, out of 50:
That’s why the final averages leave out each model’s score of its own entry, and count Claude, which isn’t a contestant, as one of the judges. Claude’s own ranking matched the model judges’ on the tornado film and differed on the lighthouse, where it put GLM first while all three model judges gave DeepSeek their highest total.
A director that writes a great plan once in three tries, or bills you ten times more for it, isn’t the best choice for most projects. Here’s what we measured.
GLM-5.3 Flash and DeepSeek V4.1 Flash planned every film in 36–59 seconds. Kimi K3 matched them on the tornado film (42 s), took 82 s on the lighthouse, and needed 255 seconds and three planning passes for the robot film, including 7 transient errors from its provider along the way.
Every model had at least one attempt fail, and most failures were about shape, not ideas. DeepSeek produced a 21 KB lighthouse plan with a single stray brace; KyndCreate now repairs one-character slips like that. Kimi K3 wrapped a list in an unexpected “fields” object once and wrote zero-length shot timings twice; its lighthouse plan went through on the third attempt once the checker accepted those. GLM-5.3 Flash had one attempt with a stray field in an ingredient, and one that stalled on an image step.
Across all three films GLM-5.3 Flash spent 15 writing credits in total. DeepSeek V4.1 Flash spent 38. Kimi K3 spent 223, 126 of them on the robot film alone.
The pictures usually cost more than the writing. Frame credits scale with how many pictures a director asks for: DeepSeek’s 11-picture lighthouse plan came to 91 frame credits against 14 for its writing. Kimi K3 was the exception on two films: 66 writing credits against 56 for its lighthouse pictures, and 126 against 56 on the robot film.
Running three directors side by side stress-tested our own studio too. We fixed what it found:
It depends which part of the job you mean. A language model such as Kimi K3, GLM-5.3 Flash or DeepSeek V4.1 Flash does not draw the pictures; it directs, writing the treatment, the cast and location prompts, the opening frames and the shot plan. In our final blind test DeepSeek V4.1 Flash edged it overall, averaging 36.4 out of 50 from the other judges against 35.4 for GLM-5.3 Flash and 34.0 for Kimi K3, and each model won one of the three films. GLM-5.3 Flash was the cheapest and drew the most cinematic frames, DeepSeek was the most inventive but broke a continuity rule, and Kimi K3 held a character and a set together best but was the slowest and most expensive. With three films and one run each, treat the overall result as close.
Only on the robot film, which it won for continuity: the same robot in the same gallery in every frame. On the tornado film all three model judges ranked it last, largely because it planned only two opening frames for four shots. On the lighthouse film it tied for second, after drawing two keepers at the tower’s base and two lighthouses in its final frame.
In our runs, GLM-5.3 Flash: 5 writing credits per film on all three briefs. DeepSeek V4.1 Flash used 11 to 14. Kimi K3 used 31 on the tornado film, 66 on the lighthouse and 126 on the robot film.
No. In KyndCreate the language model is the director. The same image models (Pruna P-Image for cast and locations, P-Image-Edit for the composed opening frames) drew every picture for every contestant, so the differences you see come from the direction, not the renderer.
Not reliably, which is why every judge scored blind and why the final averages leave out each model’s score of its own entry. On the robot film all three models scored their own entry above the other judges’ average, by 4.7 to 6.7 points out of 50. Elsewhere it went both ways: Kimi K3 scored its own tornado and lighthouse entries 6.3 and 5.3 points below the others, and DeepSeek V4.1 Flash scored its own lighthouse entry 2.3 below.
Yes. Start a new creation in KyndCreate with one of the three briefs, choose Kimi K3, GLM Flash or DeepSeek Flash as your creative companion, and review the plan it proposes. Nothing changes in your project until you apply it.
Not yet. This first run is fully judged, but we then found KyndCreate was drawing the films’ characters and locations as finished scenes instead of clean reference pictures, which likely contributed to some continuity faults. That is fixed and all nine films are being re-run; we will update this page and its Updated date with the new results rather than publish a new one.
More from Kynd: the KyndCreate blog · notes on running AI video and image models on a Mac.