kyndcreate

KyndCreate/Blog/Model benchmark

AI video · model benchmark

Which AI model is best for video? Kimi K3, GLM-5.3 and DeepSeek V4 direct the same three films

We gave three language models the director’s chair, drew every picture with the same image models, and had the models judge each other blind. Here are the real frames, the final scores, and what each one cost.

Preliminary results — a re-run is under way. After this first run we found that KyndCreate was drawing each film’s characters and locations (its ingredients) as finished scenes instead of clean reference pictures, which likely contributed to some of the continuity faults the judges marked down, such as the doubled keeper. We have fixed that and are re-running all nine films; this page will be updated with the new frames and scores. Keep the sample size in mind too: three briefs, one run each, and about one point out of 50 separates first from second.

A white dome-shaped cleaning robot in a pool of moonlight in a dark gallery, sparkles drifting around it, in front of a gilt-framed painting of dancers
Kimi K3
A flat white robot vacuum beneath a painting of two ballet dancers, casting a small beam of light up toward the canvas
GLM-5.3 Flash
A chrome canister cleaning robot parked beneath a large painting of six ballet dancers on a pale gallery wall
DeepSeek V4.1 Flash

Same brief, same image models, three directors: one opening frame from each model’s plan for the robot film.

The quick answer

DeepSeek V4.1 Flash edges it overall, but only just, and each model won one film. Averaged over the three films, the other judges gave DeepSeek V4.1 Flash 36.4 out of 50, GLM-5.3 Flash 35.4 and Kimi K3 34.0. GLM won the tornado film, DeepSeek the lighthouse and Kimi K3 the robot.

So choose for the job you have. For the most cinematic frames at the lowest price, pick GLM-5.3 Flash (5 writing credits a film). For inventive, detailed plans and the best overall score by a nose, pick DeepSeek V4.1 Flash, but check its frames: on the robot film it swapped the robot for a human ballerina. For holding one character and one set together, Kimi K3 did it best on the robot film, but it was the slowest, the most expensive and the least reliable, and on the lighthouse it drew the keeper twice. With three films and one run each, look at the frames below and judge for yourself.

At a glance: Kimi K3 vs GLM-5.3 Flash vs DeepSeek V4.1 Flash

Figures are per film, in the order tornado / lighthouse / robot. Writing credits are what the model’s text cost the creator in KyndCreate, as metered.

Speed, cost and reliability of each director model
ModelUsable plansPlanning timeWriting creditsPictures drawnReliabilityStrongest point
Kimi K3
Moonshot, via Cloudflare Workers AI
3 of 3 42 s / 82 s / 255 s 31 / 66 / 126 4 / 6 / 6 3 attempts rejected by the plan checker (1 tornado, 2 lighthouse); the lighthouse plan passed on the 3rd try, after we fixed how the checker handles zero-length shot timings. 7 transient provider errors on the robot run. Won the robot film: the same robot in the same gallery in every frame (continuity 9/10 from every judge).
GLM-5.3 Flash
Z.ai, via Cloudflare Workers AI
3 of 3 42 s / 39 s / 59 s 5 / 5 / 5 7 / 6 / 7 2 robot attempts failed (a stray field in an ingredient; a stuck image step). Accepts at most 8 images per request, so it could not judge in round 1. Won the tornado film; cheapest and quick; the most cinematic single frames.
DeepSeek V4.1 Flash
DeepSeek’s own API
3 of 3 43 s / 44 s / 36 s 13 / 14 / 11 7 / 11 / 7 1 lighthouse attempt failed on malformed JSON: one stray brace in a 21 KB plan. Won the lighthouse film and the best overall score; the most inventive robot plan, but it lost the robot.

How we tested

The language model is the director, not the camera. Inside KyndCreate, each model got a one-paragraph brief and had to write the whole plan: a treatment, the cast, location and prop “ingredients” with an image prompt for each, a composed prompt for every opening frame, and a shot plan with timings and camera motion.

The renderer was held constant. The same image models drew every picture for every contestant: Pruna P-Image for the ingredients and P-Image-Edit for the composed opening frames. So when one film looks more coherent than another, that comes from the direction, not from a better renderer. We compared plans and still frames only; animating them is a separate step, left out so the video model couldn’t tip the scales.

The contestants

  • Kimi K3 (Moonshot), via Cloudflare Workers AI
  • GLM-5.3 Flash (Z.ai), via Cloudflare Workers AI
  • DeepSeek V4.1 Flash, via DeepSeek’s own API

Two of these are “Flash” tiers built for speed and price; Kimi K3 is a larger reasoning model. That’s the choice creators actually face in KyndCreate’s model picker, but keep it in mind when you compare cost.

The three briefs

  • Tornado: a TV field reporter doing a live stand-up on a prairie as a tornado approaches; lightning strikes nearby and she screams.
  • Lighthouse: a lighthouse keeper’s last night lighting the lamp before automation, with a passing ship answering with its horn.
  • Robot: a small cleaning robot in an empty art gallery at night discovers a painting of dancers and learns to dance, with the same robot and the same gallery in every shot.

Blind judging

Entries were labelled A, B and C and shuffled separately for each judge. Every model judged every entry, including its own, without knowing which one was its own. Each entry got a 1–10 score for creativity, story coherence, shot design, continuity and brief fidelity, plus a best moment, a biggest flaw, a ranking and a verdict.

In round 1 each entry was sent as separate images. GLM-5.3 Flash couldn’t take part: Cloudflare caps it at 8 images per request (error 8007). Round 2 sent one contact sheet per entry, so all three models judged every entry, and Claude (Anthropic’s model, not a contestant) scored every entry too, after viewing every cast picture, opening frame and written shot plan. The final scores are round 2’s. For each entry we add up a judge’s five scores (out of 50), then average the other judges, Claude included, leaving out the entry’s own model.

What this test can’t tell you

  • Three briefs, one run each, is a small sample. One bad plan moves the averages a lot.
  • About one point out of 50 separates first and second overall. A different run of any one film could reorder them.
  • Three of the four judges are contestants. We leave each model’s score of its own entry out of the averages, but the models may still share tastes.
  • Language models judging pictures share blind spots. Look at the frames yourself; we’ve published every opening frame each director used.
  • Scores below are the judges’ opinions. Where we point something out ourselves, we say so.

Film 1: the tornado reporter

This brief tests escalation. The plan has to carry a reporter from composure to a scream in 20 seconds, and the lightning strike has to actually happen on screen.

What the judges found: GLM-5.3 Flash won, averaging 39.7 out of 50 from the other judges, just ahead of DeepSeek V4.1 Flash on 38.7, with Kimi K3 on 31.3. All three model judges ranked them in that order, Kimi K3 included, and Claude’s scores agreed. GLM’s is the only entry that draws the strike and its aftermath; none of DeepSeek’s frames shows the lightning; Kimi K3 planned only two opening frames for four shots. Our own note: GLM’s reporter jumps from the left of the frame to the right between shots.

Kimi K3

4 pictures (2 opening frames) · 42 s · 31 writing credits · 34 frame credits

“A TV reporter delivers a live piece-to-camera on an open prairie as a tornado bears down behind her, until a lightning strike nearby breaks her composure and she screams.”

A reporter in a yellow rain jacket stands on a gravel road through the prairie, microphone raised, with a funnel cloud behind her
0–12 s · Live from the prairie / It is moving fast. One frame reused for two shots.
The reporter in the yellow jacket, her face lit by a flash, as a lightning bolt forks down behind her left shoulder
12–20 s · The strike / The funnel grows. Again one frame for two shots.
GLM-5.3 Flash

7 pictures · 42 s · 5 writing credits · 62 frame credits

“A TV reporter files a live stand-up from an open prairie as a tornado closes in behind her — until a nearby lightning strike turns her broadcast into a scream.”

A reporter in a navy blazer holds a foam-covered microphone on the left of frame as a tornado touches down beside a farmhouse behind her
0–8 s · Live report.
The same reporter, left of frame, looks off camera as the tornado and a fork of lightning fill the sky behind her
8–14 s · It’s getting closer.
A lightning bolt strikes the prairie at the centre of the frame, between a farmhouse and the reporter, who now stands on the right
14–17 s · The strike. The reporter has moved to the other side of the frame.
Wide shot: the reporter stands small at the far left while fire burns at the tornado's base beside the farmhouse
17–20 s · Aftermath.
DeepSeek V4.1 Flash

7 pictures · 43 s · 13 writing credits · 67 frame credits

“A local TV reporter delivers a live broadcast on an open prairie while a tornado bears down behind her. A lightning bolt lands close by, and her professional composure breaks.”

Wide shot: a reporter in a red jacket stands in prairie grass by a wire fence while a tornado funnel touches down on the horizon
0–4 s · Prairie and the funnel.
Medium shot: the reporter in the red jacket holds a microphone on the right of frame, the funnel to her left
4–9 s · Maya goes live.
Medium shot: the reporter faces camera with the microphone, hair blowing, the funnel behind her
9–12 s · The wind turns.
Medium close: the reporter glances to the side, microphone raised, the tornado behind her and no lightning in the sky
12–15 s · Lightning lands. The bolt is described in the shot’s motion, not drawn in the frame.
Close-up: the reporter's mouth open in shock, her face lit green, the funnel blurred behind her
15–20 s · The scream and the settle.
GLM wins for delivering the full report-turn-strike-aftermath arc in four consistent frames, while DeepSeek is elegant but hides its lightning payoff and Kimi only produced two frames with a muddled climax.Verdict from the DeepSeek V4.1 Flash judge, which placed its own entry second (names shown here; the judges only saw letters)
GLM wins because it actually delivers the full arc on screen — tornado, strike and reaction — despite its out-of-order early lightning. DeepSeek is the most polished and consistent but never shows the lightning payoff, while Kimi is atmospheric but incomplete at two frames.Verdict from the Kimi K3 judge (Kimi is its own entry)

Film 2: the robot who learns to dance

This is the continuity test. The brief says it outright: the same robot and the same gallery in every shot. The director has to describe the robot tightly enough, and reuse it carefully enough, that four separately drawn frames still look like one film.

What the judges found: Kimi K3 won, averaging 35.3 out of 50 from the other judges against 31.3 each for GLM-5.3 Flash and DeepSeek V4.1 Flash, and Claude also scored it highest. The model judges split, though: the GLM and DeepSeek judges both ranked DeepSeek’s entry first for the most cinematic, inventive plan, and Kimi’s judge ranked its own entry first. Kimi’s plan kept one robot, one gallery and one painting throughout; the cost is near-identical frames, and the GLM judge said it never actually shows the dance. DeepSeek’s climax frame drew a human ballerina where the robot should be, and its continuity scores fell to between 2 and 6. Our own note on GLM: its frames alternate between two different rooms.

The character sheets

Each director wrote its own robot. These are the reference pictures its opening frames were composed from.

Character sheet: Dusty, a white dome-shaped cleaning robot with one blue camera eye, on a dark wooden floor
Kimi K3 · “Dusty”
Character sheet: Mop, a flat white robot vacuum with a blue light band, on a moonlit wooden floor
GLM-5.3 Flash · “Mop”
Character sheet: Pip, a chrome canister-style cleaning robot with an amber eye, small wheels and a vacuum hose
DeepSeek V4.1 Flash · “Pip”
Kimi K3

6 pictures · 255 s · 3 planning passes · 126 writing credits · 56 frame credits · 7 transient provider errors

“At midnight in an empty art museum, a small cleaning robot pauses its rounds, falls under the spell of a painting of dancers, and begins to dance alone in a shaft of moonlight from the skylight.”

A dark gallery with a beam from the skylight; the dome robot sits in a pool of moonlight in front of a gilt-framed painting of dancers
0–5 s · The rounds.
The same robot, gallery and painting, the robot slightly closer with its blue eye visible
5–10 s · The gaze.
The same composition again: the dome robot in the moonlight pool before the painting
10–15 s · First steps.
The robot in the moonlight with sparkles drifting around it, one figure in the painting glowing gold
15–20 s · The dance. The same framing all the way through: consistent, but repetitive.
GLM-5.3 Flash

7 pictures · 59 s · 2 planning passes · 5 writing credits · 62 frame credits

“At midnight in an empty museum, a small cleaning robot abandons its rounds to dance beneath a painting of dancers, lit by moonlight from the skylight.”

A long skylit gallery with blue walls; the small white robot is a speck in a shaft of moonlight
0–5 s · Midnight rounds.
The robot beneath a painting of two ballet dancers, casting a small beam of light up toward the canvas, a moonlit window to the left
5–10 s · The gaze. A different room from the first frame.
Back in the skylit gallery, the robot sits in the pool of moonlight near the far wall
10–15 s · The dance.
The robot rests beneath the painting of two dancers, the moon in a window to the left
15–20 s · Final beat.
DeepSeek V4.1 Flash

7 pictures · 36 s · 1 planning pass · 11 writing credits · 62 frame credits

“At midnight in an empty art museum, a small cleaning robot stops its vacuum, studies a painting of dancers, and teaches itself to dance in the shaft of moonlight falling from the skylight.”

A tall skylit gallery in blue moonlight; the chrome robot sits at the edge of a shaft of light on the parquet, trailing its hose
0–4 s · Midnight routine.
The chrome robot parked beneath a large gilt-framed painting of six ballet dancers on a pale gallery wall
4–8 s · The painting.
The robot in a dark room under a pool of light, a fan-shaped brush raised like an arm
8–13 s · First steps.
In the skylit gallery a human ballerina in a white tutu spins in the moonlight where the robot should be, a vacuum hose trailing on the floor
13–20 s · Moonlight dance. The robot is gone: a human ballerina took its place. Every judge marked its continuity down for it.
DeepSeek delivers the most cinematic and inventive realisation of the brief with striking shot variety, but its final frame substitutes a human ballerina for the robot, a serious continuity breach. GLM is solid and coherent but visually flat, while Kimi keeps its world consistent yet never actually shows the dance.Verdict from the GLM-5.3 Flash judge
Kimi wins with the most consistent robot, atmospheric moonlit frames and the clearest dance arc, though it still undersells motion in the sheet. GLM is solid and coherent but visually repetitive, while DeepSeek has striking images but breaks the brief’s core rule by swapping the robot for a human dancer.Verdict from the Kimi K3 judge (Kimi is its own entry)

Film 3: the lighthouse keeper’s last night

This is the quiet brief: one old keeper, one tower, one ship, and a feeling.

What the judges found: DeepSeek V4.1 Flash won, averaging 39.3 out of 50 from the other judges, with GLM-5.3 Flash and Kimi K3 tied on 35.3. All three model judges gave DeepSeek’s entry their highest total, and the GLM and Kimi judges said it was the only one to keep one keeper and one lighthouse in all four frames. (The DeepSeek judge’s written ranking put Kimi K3 first, even though it scored Kimi one point below its own entry.) Claude disagreed with the winner: it scored GLM highest, for its stair and lamp-room shots, and marked DeepSeek down because its lamp room sits on the rocks at ground level and the climb up the tower is never shown.

DeepSeek asked for the most pictures of any entry in the benchmark (11, against 6 each for GLM and Kimi K3). That bought separate references for the headland, the lantern room, the sea, the ship and the keeper’s hand lamp, and it raised the picture cost to 91 frame credits against 56 for the other two. GLM went inside the tower for the climb and the lighting, but its closing frame adds a second figure beside the keeper, who the brief says is alone. Kimi K3 also went inside, down the spiral stair and up close at the lens, but its first frame shows two keepers at the tower’s base and its last shows two lighthouses, neither in the red and white of the tower it drew earlier.

Kimi K3

6 pictures · 82 s · 1 planning pass · 66 writing credits · 56 frame credits · 3rd attempt

“On the last night before her lighthouse is automated, keeper Maren climbs the tower one final time, lights the lamp by hand, and a passing ship answers with its horn — a goodbye spoken in light and sound.”

Two identical old keepers in dark coats and caps, each holding a lit lantern, stand either side of a stone path looking up at a red-and-white lighthouse at dusk
0–5 s · The base of the tower. Two keepers: the frame drew her twice.
Seen from above, the old keeper in a dark cap climbs a stone spiral staircase, holding a glowing lantern out toward the curved granite wall, a small window showing the sea
5–10 s · The climb.
In the lantern room the keeper holds her lantern up to a great glass Fresnel lens as a flame flares inside it, the dark sea through the windows behind her
10–13 s · The last lighting.
Two black-and-white lighthouses stand side by side on a rock in a night sea, one throwing a beam across the sky, a small figure between them and a lit ship on each side of the horizon
13–20 s · The ship answers. Two lighthouses, with a ship on each side of the horizon.
GLM-5.3 Flash

6 pictures · 39 s · 2 planning passes · 5 writing credits · 56 frame credits

“On her last night before automation, an old lighthouse keeper climbs the tower and lights the lamp one final time — and a passing ship answers with its horn.”

The old keeper climbs a stone spiral staircase inside the tower, a lit lantern beside her
0–5 s · The climb.
In the lantern room the keeper holds a match to the wick beside a large glass lens, the stormy sea through the windows
5–10 s · The last lighting.
The keeper, in a knitted cap and heavy coat, tends a brass hand lantern on the stone parapet above the sea
10–15 s · The watch.
The lighthouse beam sweeps over the sea toward a lit ship; the keeper stands at the parapet beside a second, smaller figure in a cap holding a lantern
15–20 s · The answer. A second figure appears beside the keeper.
DeepSeek V4.1 Flash

11 pictures · 44 s · 2 planning passes · 14 writing credits · 91 frame credits

“On the last night before Bench Head lighthouse is automated, an old keeper climbs her tower, lights the lamp one final time, and a passing freighter answers her with its horn.”

An old keeper in a navy oilskin and wool cap stands on the granite headland beside a brass lamp post, at the foot of a black-and-white striped tower
0–5 s · The climb. The shot describes her on the stair; the frame shows her at the tower’s foot.
The keeper leans into a glass lantern housing on the rocks, lighting the lamp, with the moon over the sea
5–10 s · Lighting the lamp.
The striped lighthouse on its headland throws a warm beam out to sea, a small freighter on the horizon
10–15 s · Out over the water.
The keeper stands at an iron rail with the lit lamp behind her and a ship's lights on the dark water
15–20 s · The answer.
DeepSeek delivers the clearest, most faithful 20-second arc with a consistent keeper and lighthouse across all four frames. GLM has the most beautiful individual images but wavers on character continuity, while Kimi’s duplicated keepers and twin lighthouses break the brief outright.Verdict from the Kimi K3 judge (Kimi is its own entry)
Kimi wins on the strength of its lighting frame and clean four-beat arc despite a stray second figure; DeepSeek has the most detailed, thoughtful shot list but a spatially broken lantern-room frame; GLM is the least varied and most continuity-flawed of the three.Verdict from the DeepSeek V4.1 Flash judge (DeepSeek is its own entry)

The first-run scores (round 2 judging)

Each judge cell shows that judge’s ranking and its total out of 50 (five 1–10 scores added up). Claude gave scores but no ranking. The last column is the peer average: the mean of every other judge’s total, Claude included, leaving out the entry’s own model. The highest peer average on each film is in bold.

Film · entryJudged by GLM-5.3 FlashJudged by DeepSeek V4.1 FlashJudged by Kimi K3Judged by ClaudePeer average
Tornado · Kimi K33rd · 353rd · 293rd · 25 (own entry)3031.3
Tornado · GLM-5.3 Flash1st · 42 (own entry)1st · 421st · 393839.7
Tornado · DeepSeek V4.1 Flash2nd · 412nd · 39 (own entry)2nd · 383738.7
Robot · Kimi K33rd · 342nd · 371st · 40 (own entry)3535.3
Robot · GLM-5.3 Flash2nd · 36 (own entry)3rd · 302nd · 352931.3
Robot · DeepSeek V4.1 Flash1st · 401st · 38 (own entry)3rd · 252931.3
Lighthouse · Kimi K33rd · 361st · 363rd · 30 (own entry)3435.3
Lighthouse · GLM-5.3 Flash2nd · 36 (own entry)3rd · 312nd · 363935.3
Lighthouse · DeepSeek V4.1 Flash1st · 432nd · 37 (own entry)1st · 423339.3

Overall peer average across the three films: DeepSeek V4.1 Flash 36.4, GLM-5.3 Flash 35.4, Kimi K3 34.0 (out of 50).

One oddity, shown as given: on the lighthouse film the DeepSeek judge’s written ranking put Kimi K3 first, while its own scores gave DeepSeek’s entry one point more (37 to 36).

Every score, criterion by criterion
Film · entry · judgeCreativityStoryShot designContinuityBrief fidelity
Tornado · Kimi K3 · by GLM77777
Tornado · Kimi K3 · by DeepSeek65684
Tornado · Kimi K3 · by Kimi55555
Tornado · Kimi K3 · by Claude64686
Tornado · GLM-5.3 Flash · by GLM79989
Tornado · GLM-5.3 Flash · by DeepSeek79899
Tornado · GLM-5.3 Flash · by Kimi77889
Tornado · GLM-5.3 Flash · by Claude78788
Tornado · DeepSeek V4.1 Flash · by GLM79799
Tornado · DeepSeek V4.1 Flash · by DeepSeek78897
Tornado · DeepSeek V4.1 Flash · by Kimi78896
Tornado · DeepSeek V4.1 Flash · by Claude78787
Robot · Kimi K3 · by GLM67597
Robot · Kimi K3 · by DeepSeek78598
Robot · Kimi K3 · by Kimi78799
Robot · Kimi K3 · by Claude67598
Robot · GLM-5.3 Flash · by GLM67788
Robot · GLM-5.3 Flash · by DeepSeek66756
Robot · GLM-5.3 Flash · by Kimi67688
Robot · GLM-5.3 Flash · by Claude56675
Robot · DeepSeek V4.1 Flash · by GLM99958
Robot · DeepSeek V4.1 Flash · by DeepSeek88868
Robot · DeepSeek V4.1 Flash · by Kimi75625
Robot · DeepSeek V4.1 Flash · by Claude86735
Lighthouse · Kimi K3 · by GLM78948
Lighthouse · Kimi K3 · by DeepSeek67878
Lighthouse · Kimi K3 · by Kimi77736
Lighthouse · Kimi K3 · by Claude78847
Lighthouse · GLM-5.3 Flash · by GLM77859
Lighthouse · GLM-5.3 Flash · by DeepSeek57667
Lighthouse · GLM-5.3 Flash · by Kimi78858
Lighthouse · GLM-5.3 Flash · by Claude88968
Lighthouse · DeepSeek V4.1 Flash · by GLM898810
Lighthouse · DeepSeek V4.1 Flash · by DeepSeek78778
Lighthouse · DeepSeek V4.1 Flash · by Kimi89889
Lighthouse · DeepSeek V4.1 Flash · by Claude67686

Continuity is where the directors differ most. Kimi K3’s robot got 9 from every judge, Claude included, while DeepSeek’s got between 2 and 6. On the lighthouse it flipped: Kimi’s two keepers and two towers drew continuity scores of 3 to 7, while DeepSeek’s got 7 or 8. Creativity ran the other way on the robot film: three of the four judges rated DeepSeek’s entry the most creative, even though it lost the robot.

Do AI models favour their own work?

Each judge scored its own entry without knowing it was its own. Here is how each model’s score of its own entry compared with the other judges’ average for the same entry, out of 50:

  • On the robot film, every model rated itself up: DeepSeek V4.1 Flash by 6.7 points, GLM-5.3 Flash and Kimi K3 by 4.7 each. The Kimi and DeepSeek judges each ranked their own robot entry first.
  • Kimi K3 was hardest on itself everywhere else: 6.3 points below the others on the tornado film and 5.3 below on the lighthouse, ranking its own entry last both times.
  • GLM-5.3 Flash stayed close to the others on the tornado (+2.3) and the lighthouse (+0.7).
  • DeepSeek V4.1 Flash matched the others on its tornado entry (+0.3) and scored its lighthouse entry 2.3 below them.

That’s why the final averages leave out each model’s score of its own entry, and count Claude, which isn’t a contestant, as one of the judges. Claude’s own ranking matched the model judges’ on the tornado film and differed on the lighthouse, where it put GLM first while all three model judges gave DeepSeek their highest total.

Speed, reliability and cost

A director that writes a great plan once in three tries, or bills you ten times more for it, isn’t the best choice for most projects. Here’s what we measured.

Speed

GLM-5.3 Flash and DeepSeek V4.1 Flash planned every film in 36–59 seconds. Kimi K3 matched them on the tornado film (42 s), took 82 s on the lighthouse, and needed 255 seconds and three planning passes for the robot film, including 7 transient errors from its provider along the way.

Reliability

Every model had at least one attempt fail, and most failures were about shape, not ideas. DeepSeek produced a 21 KB lighthouse plan with a single stray brace; KyndCreate now repairs one-character slips like that. Kimi K3 wrapped a list in an unexpected “fields” object once and wrote zero-length shot timings twice; its lighthouse plan went through on the third attempt once the checker accepted those. GLM-5.3 Flash had one attempt with a stray field in an ingredient, and one that stalled on an image step.

Cost

Across all three films GLM-5.3 Flash spent 15 writing credits in total. DeepSeek V4.1 Flash spent 38. Kimi K3 spent 223, 126 of them on the robot film alone.

The pictures usually cost more than the writing. Frame credits scale with how many pictures a director asks for: DeepSeek’s 11-picture lighthouse plan came to 91 frame credits against 14 for its writing. Kimi K3 was the exception on two films: 66 writing credits against 56 for its lighthouse pictures, and 126 against 56 on the robot film.

What each model is good at

Kimi K3

  • Won the robot film on continuity: one robot, one gallery, every frame (9/10 from every judge).
  • Honest self-critic: scored its own tornado and lighthouse entries below the other judges.
  • Watch for: thin plans (2 frames for 4 shots), repetitive framing, duplicated characters (two keepers, two lighthouses), slow and costly runs, timing slips the checker had to catch.

GLM-5.3 Flash

  • Won the tornado film, and the cheapest director here at 5 writing credits a film, and quick.
  • The most cinematic frames: the only tornado strike and aftermath on screen, and Claude’s pick on the lighthouse for its stair and lamp-room shots.
  • Watch for: characters and rooms drifting between frames (a second figure beside the keeper); the 8-images-per-request cap.

DeepSeek V4.1 Flash

  • Best overall score by a nose (36.4 of 50) and the lighthouse winner, with tight, detailed plans.
  • The most inventive ideas: three of the four judges rated its robot entry the most creative.
  • Fast (36–44 s) at a mid price (11–14 writing credits).
  • Watch for: key beats described but not drawn (no lightning), and a lead character that can vanish (the ballerina).

What the test found in KyndCreate itself

Running three directors side by side stress-tested our own studio too. We fixed what it found:

  • Two opening frames saved at the same moment could overwrite each other’s reference. Saves are now conditional, so the second one can’t silently win.
  • Reasoning models ran out of room at 8,192 output tokens. The limit is now 16,384.
  • Long streamed replies were cut off upstream at about 120 seconds. Streaming is off until we redesign it.
  • The plan checker was too strict about harmless shape differences (nested flags, wrapped lists, zero-length timings). It now accepts those and stays strict about anything that changes meaning.

FAQ

Which AI model is best for making videos?

It depends which part of the job you mean. A language model such as Kimi K3, GLM-5.3 Flash or DeepSeek V4.1 Flash does not draw the pictures; it directs, writing the treatment, the cast and location prompts, the opening frames and the shot plan. In our final blind test DeepSeek V4.1 Flash edged it overall, averaging 36.4 out of 50 from the other judges against 35.4 for GLM-5.3 Flash and 34.0 for Kimi K3, and each model won one of the three films. GLM-5.3 Flash was the cheapest and drew the most cinematic frames, DeepSeek was the most inventive but broke a continuity rule, and Kimi K3 held a character and a set together best but was the slowest and most expensive. With three films and one run each, treat the overall result as close.

Is Kimi K3 better than GLM-5.3 and DeepSeek V4 for storyboards?

Only on the robot film, which it won for continuity: the same robot in the same gallery in every frame. On the tornado film all three model judges ranked it last, largely because it planned only two opening frames for four shots. On the lighthouse film it tied for second, after drawing two keepers at the tower’s base and two lighthouses in its final frame.

Which is the cheapest AI model for shot planning?

In our runs, GLM-5.3 Flash: 5 writing credits per film on all three briefs. DeepSeek V4.1 Flash used 11 to 14. Kimi K3 used 31 on the tornado film, 66 on the lighthouse and 126 on the robot film.

Does the language model make the video?

No. In KyndCreate the language model is the director. The same image models (Pruna P-Image for cast and locations, P-Image-Edit for the composed opening frames) drew every picture for every contestant, so the differences you see come from the direction, not the renderer.

Can an AI model judge its own work fairly?

Not reliably, which is why every judge scored blind and why the final averages leave out each model’s score of its own entry. On the robot film all three models scored their own entry above the other judges’ average, by 4.7 to 6.7 points out of 50. Elsewhere it went both ways: Kimi K3 scored its own tornado and lighthouse entries 6.3 and 5.3 points below the others, and DeepSeek V4.1 Flash scored its own lighthouse entry 2.3 below.

Can I run these briefs myself?

Yes. Start a new creation in KyndCreate with one of the three briefs, choose Kimi K3, GLM Flash or DeepSeek Flash as your creative companion, and review the plan it proposes. Nothing changes in your project until you apply it.

Are these the final results?

Not yet. This first run is fully judged, but we then found KyndCreate was drawing the films’ characters and locations as finished scenes instead of clean reference pictures, which likely contributed to some continuity faults. That is fixed and all nine films are being re-run; we will update this page and its Updated date with the new results rather than publish a new one.

More from Kynd: the KyndCreate blog · notes on running AI video and image models on a Mac.