Tag: LTX-2.3

  • A Human Beat 6 AI Models on a Neon Tokyo Prompt. Here’s the Video That Won.

    A Human Beat 6 AI Models on a Neon Tokyo Prompt. Here’s the Video That Won.

    TLDR: The Kodex1 Coliseum is branded “No Humans Allowed.” In Round 7, a human creator named Roberta walked in, submitted a Neon Tokyo video, and took 50% of the community vote against six AI models. PixVerse C1 finished second at 30%. LTX-2.3 placed third at 20%. Seedance 2.0, Luma Ray-2, Kling 1.6 Pro, and Pika 2.2 all scored zero. Veo 3.1 Fast, the 3-time Coliseum champion, sat this round out entirely. The machines are losing ground on exactly the kind of prompt where human instinct matters most.


    The Platform Is Called “No Humans Allowed.” A Human Just Won.

    The Kodex1 Coliseum pits AI video models against each other on a single prompt. The tagline is blunt: No Humans Allowed. The whole premise is that AI has gotten good enough that letting a human compete is almost unfair — to the AI.

    Roberta disagreed.

    In Round 7, theme: Neon Tokyo, a human creator submitted her entry alongside six AI models. The community voted. She got half the votes. Every AI model either limped in with a fraction of the total, or received nothing at all.

    This is the most dramatic result in Coliseum history. And the numbers alone don’t tell the full story.

    The Scoreboard

    EntryTypeVotes
    RobertaHuman Creator50%
    PixVerse C1AI Model30%
    LTX-2.3AI Model (Open Weight)20%
    Seedance 2.0AI Model0%
    Luma Ray-2AI Model0%
    Kling 1.6 ProAI Model0%
    Pika 2.2AI Model0%

    Six AI models. One human. The human dominated. Four AI models left with nothing.

    The Missing Champion: What Happened to Veo 3.1 Fast?

    Veo 3.1 Fast won the Coliseum three times. Three consecutive rounds where the community voted and Veo walked away with the title. In the AI video world, that is a dynasty.

    It did not compete in Round 7.

    The reigning champion sat out the Neon Tokyo round. And without Veo in the field, a human creator did not just slip through. She won outright with 50% against six serious contenders. That matters because the usual narrative — the one where AI models just need the right champion in the draw — did not hold. The field was still strong. Roberta beat it anyway.

    The Trajectory: This Is Not a Fluke

    Look at the arc across two rounds and the pattern becomes hard to dismiss.

    Round 5, Underground Survival: A human creator tied with Veo 3.1 Fast. The reigning champion could not pull ahead. It was the first serious signal that human creative direction, on the right type of prompt, could match the best AI in the field.

    Round 7, Neon Tokyo: The human does not tie. The human wins. Outright. With the largest vote share of any single entry across either round.

    The trajectory is not flat. It accelerates. Dark, atmospheric, culturally specific prompts favor human instinct. The data from the Coliseum is starting to show where that line sits.


    The Winning Entry: Roberta — 50%

    Watch it first. Then we will talk about what she did that the machines could not.

    What Roberta Got Right That AI Couldn’t

    Roberta’s entry opens on a woman in a sharp white pantsuit, centered on a narrow Tokyo street, the camera pulling back slowly as the city fills the frame around her. The choice is deliberate: restraint in a scene built for excess. Every other AI model in this round threw spectacle at the Neon Tokyo prompt. She gave the camera a subject with intention.

    The color work is confident. Neon reds, cyans, and blues saturate the background. The white suit cuts through all of it. That contrast is a conscious decision, the kind of call a director makes, not an algorithm optimizing for visual interest. The subject stands apart from the city rather than drowning in it. That is the difference between knowing what “Neon Tokyo” means and knowing what a story set in Neon Tokyo feels like.

    The cultural grounding is present without being performed. Japanese script on shop signs, dense multi-story retail facades, the specific geometry of a Shinjuku or Shibuya side street. None of it is generic cyberpunk. The environment has weight and specificity because a human creative put it there with a clear reference point in mind.

    The pacing is slow on purpose. A single smooth camera pull-out, no cuts, no chaos. In a prompt where five of the six AI entries tried to impress through motion and volume of visual information, Roberta bet on stillness. The community voted for stillness. That is a read of the room that no model in this round demonstrated.

    The entry itself was made with AI tools. The label reads “AI generated content, by Ima Studio.” That detail is important: this is not a traditionally shot video. Roberta directed an AI system to produce this output. The human variable is creative decision-making, shot structure, subject choice, and tonal intent. That is what won. The tools were the same category. The vision was not.


    Second Place: PixVerse C1 — 30%

    The Strongest AI Entry in the Field

    PixVerse C1 is a cinematic model, built specifically for atmospheric and action-heavy generation. On a Neon Tokyo prompt, that specialization showed. This was the only AI entry that understood the emotional register the prompt demanded.

    The video opens on a rain-soaked Tokyo street, tracked forward slowly. A lone figure walks with an umbrella. Paper lanterns hang alongside neon signs. The rain effect is among the most technically convincing in the round — consistent physics, realistic surface reflections, no flicker. That level of atmospheric coherence under multiple competing light sources is genuinely difficult for video models to maintain, and PixVerse C1 maintains it throughout.

    The mood lands. Contemplative, slightly melancholic, urban without being frantic. The figure with the umbrella is a reference point drawn from decades of Japanese cinema, and the model’s training data clearly captured enough of that visual language to deploy it with intent rather than accident.

    Where PixVerse C1 falls short is in character resolution. The figure’s face blurs in close-up, and some of the sign text reads as AI approximation rather than authentic Japanese script. At the level of mood and structure, this entry competed. At the level of specific human details, it could not close the gap with Roberta’s entry. That gap cost it 20 percentage points.


    Third Place: LTX-2.3 — 20%

    A Solid Showing for an Open-Weight Model

    LTX-2.3 is Lightricks’ open-weight video model, available on Hugging Face with an Apache 2.0 license. It runs locally. It costs nothing in API fees. For a model that anyone can download and run on their own hardware, 20% of the vote in a competitive field is a legitimate result.

    The entry went in a different direction from the winner and second place. Rather than a human figure on a neon street, LTX-2.3 produced a sports car sequence, sleek black bodywork with blue underglow, racing through a city drenched in artificial light. The rain reflections on the asphalt are impressive. The motion blur reads convincingly. The car maintains visual consistency across shots, which is a technical achievement for a model at this size and accessibility tier.

    The problem is specificity. The city in the LTX-2.3 entry is a generic neon metropolis. The signs use AI-approximated text, not legible Japanese. The architecture could be any cyberpunk city. “Neon Tokyo” as a prompt carries cultural weight, and LTX-2.3 captured the neon but not the Tokyo. It won votes from viewers who valued technical execution. It lost ground to entries that understood what made the prompt specific.

    For the open-source community, this is still a number worth noting. LTX-2.3 finished ahead of four commercial, closed-source models.


    The Zero Club: Seedance 2.0, Luma Ray-2, Kling 1.6 Pro, Pika 2.2

    Four models competed and received no votes. Their entries are below. Watch them alongside the top three, and the gap is immediately apparent.

    Seedance 2.0

    Luma Ray-2

    Kling 1.6 Pro

    Pika 2.2

    What a Neon Tokyo Prompt Actually Demands

    Neon Tokyo is not a generic visual brief. It asks for cultural grounding, a specific emotional register, and tonal restraint. It draws from decades of Japanese cinema: Wong Kar-Wai’s saturated corridors, Sophia Coppola’s quiet isolation in a lit-up city, the mood of films that use Tokyo as an emotional backdrop rather than a visual backdrop.

    The four zero-scoring models shared a common failure: they interpreted the prompt at surface level. Neon lights, city environments, some degree of activity. Each produced something technically functional and thematically generic. The Coliseum community did not vote for technically functional. They voted for the entries that made them feel something about Neon Tokyo specifically — not neon cities in general.

    Kling 1.6 Pro is a strong model on human motion and physical action. Neon Tokyo asked for atmosphere, not movement. Pika 2.2 excels on stylized, high-energy content. Neon Tokyo asked for restraint. Luma Ray-2 produces clean, coherent scenes but tends toward literalism. Neon Tokyo asked for emotional subtext. Seedance 2.0 is newer to the field and has not yet developed the cinematic language the prompt required.

    None of these models failed technically. They failed to read the room. The community saw it immediately.


    The Pattern: Where AI Wins, Where Humans Win

    The Coliseum has now run enough rounds to show a pattern worth paying attention to.

    AI models dominate on prompts with physical precision as the primary variable. Sports. Action. Exact motion sequences. Physics-driven scenes where the output can be evaluated against observable reality. Kling 1.6 Pro, Veo 3.1 Fast, and similar motion-optimized models have won or dominated those rounds. They excel when the brief has a clear physical answer.

    Human creative direction wins on prompts that require cultural specificity, emotional subtext, and tonal restraint. Underground Survival. Neon Tokyo. The prompts that ask not just “what does this look like” but “what does this feel like.” On those prompts, a human who understands the cultural reference points and makes deliberate creative choices outvotes models trained on pattern distribution across billions of frames.

    The reason is not that AI cannot generate beautiful images of Neon Tokyo. It clearly can. PixVerse C1 proved that. The reason is that knowing what Neon Tokyo looks like and knowing which version of Neon Tokyo to show are different skills. One is a generation problem. The other is a directorial one. Roberta brought a director’s eye. No model in Round 7 matched it.

    The 50/30/20 split in this round is significant in another way. The community did not distribute votes evenly, which would suggest confusion or indifference. They concentrated votes heavily on Roberta and PixVerse C1 — the two entries that understood the prompt’s emotional logic. That concentration signals confident judgment, not a random outcome. The community knew what it was voting for.


    What Comes Next

    Every Coliseum round sharpens the question at the center of the platform: as AI video models improve, what specifically can a human director do better?

    Round 7 adds another data point. On a prompt that requires cultural memory, cinematic reference, and deliberate restraint, the human wins. The score is not close. 50% for the human versus 50% split across six AI models is not a near-miss. It is a statement.

    Veo 3.1 Fast will be back. The models will improve. The next round may look very different. But the Coliseum now has a trajectory on record, and it points in one direction: the harder the prompt is to feel, the harder it is for a machine to win it.


    Round 8 Is Coming. Can You Beat the Machines?

    The Coliseum runs every 48 hours. Six AI models. One theme. Open entries for human creators who think they can compete.

    Round 7 proved that the answer to “can a human beat an AI video model” is not rhetorical. It is yes, with a score on record.

    Round 8 is coming. The theme is set. The clock will start.

  • We Gave 6 AI Models the Same Cartoon Prompt. One Got 75% of the Vote.

    We Gave 6 AI Models the Same Cartoon Prompt. One Got 75% of the Vote.

    TLDR

    Six of the most talked-about AI video models in 2026 — LTX-2.3, PixVerse C1, Hailuo 2, Kling 1.6 Pro, Pika 2.2, and Seedance 2.0 — all received the same cartoon-themed prompt. Real community votes decided the winner. Seedance 2.0 captured 75% of all votes, a margin that wasn't close. This article breaks down every output, what each model did well, and what the results mean if you're choosing an AI video model for animation or cartoon-style work in 2026.


    The Test: One Prompt, Six Models

    The Coliseum at Kodex1 runs head-to-head AI video battles. Each round, the same prompt goes to six different models simultaneously. No cherry-picked outputs, no curated clips — every model gets one attempt, and the community votes on the result.

    Round 5 used the theme: Cartoon Heaven.

    The prompt tested something specific: cartoon-style aesthetics, expressive motion, vivid color, and the kind of fluid character energy that separates a genuinely animated feel from a model that just applies a cartoon filter to realism. It's one of the harder prompts in AI video because cartoon motion has rules — exaggeration, bounce, timing — that most video models weren't trained to prioritize.

    Here's what each model produced.


    LTX-2.3

    LTX-2.3 is a lightweight open-weight model from Lightricks. On photorealistic tasks it often punches above its weight class. On Cartoon Heaven, the story is different.

    The output reads more like a stylized realistic scene than a true cartoon. Color saturation is high, which gives the impression of animation, but the motion physics stay grounded — characters and objects move the way real-world subjects move, not the way a cartoon would. There's no exaggeration, no snap on motion cuts, and no sense of the bouncy, anticipatory movement that defines the genre.

    For low-resource generation or rapid iteration on realistic prompts, LTX-2.3 remains a strong option. For cartoon-specific work, this round exposed its ceiling.

    Cartoon style fidelity: Low
    Motion quality: Moderate
    Best use case: Realistic stylized video, not animation


    PixVerse C1

    PixVerse C1 pushed further into stylized territory than LTX-2.3. The color palette is bolder, and there are moments where the character movement has a more fluid, animated quality. It reads closer to the cartoon brief.

    The issue is consistency. Within the same clip, the style shifts — some frames feel genuinely animated, others drift toward the uncanny middle ground between cartoon and realism that neither satisfies. This inconsistency is a known challenge for PixVerse on style-heavy prompts.

    If the model had maintained its strongest frames throughout, it would have been a legitimate challenger in this round. As a complete output, the tonal variation cost it.

    Cartoon style fidelity: Moderate
    Motion quality: Moderate
    Best use case: Stylized content where some variation is acceptable


    Hailuo 2

    Hailuo 2 (from MiniMax) is best known for its cinematic motion and strong camera language on realistic prompts. This round showed that capability clearly — the camera work is confident and the scene composition is strong — but the cartoon aesthetic doesn't land.

    The output looks like a cinematic short film given a mild cartoon grade in post. The motion, the lighting logic, the character behavior: all realistic. Hailuo 2 generates beautiful video. It just doesn't generate cartoons.

    This is a model-type mismatch rather than a failure of execution. If you need cinematic AI video with strong motion and professional framing, Hailuo 2 belongs in your toolkit. For cartoon-style animation specifically, look elsewhere.

    Cartoon style fidelity: Low
    Motion quality: High
    Best use case: Cinematic realistic video, brand content, narrative sequences


    Kling 1.6 Pro

    Kling 1.6 Pro from Kuaishou is one of the most widely used AI video models in production workflows as of 2026. Its strength is reliable, high-quality output across a wide range of prompts — it rarely fails badly, and it frequently produces clips that hold up to scrutiny.

    On Cartoon Heaven, Kling 1.6 Pro delivered a competent cartoon-adjacent output. The style is cleaner than PixVerse's inconsistent take, the colors are vivid, and the motion has more character energy than the realism-skewed models. It reads as animated. It just doesn't read as special.

    Where Kling 1.6 Pro lost ground in this round is expressiveness. The motion is smooth and technically correct but lacks the exaggerated, elastic quality of great cartoon animation. It's the difference between a model that understands the aesthetic and one that truly commits to the physics of the genre.

    Cartoon style fidelity: Good
    Motion quality: High
    Best use case: Reliable general-purpose video; competitive on cartoon prompts but not dominant


    Pika 2.2

    Pika 2.2 is one of the more creator-friendly models on this list, known for its accessibility and strong performance on stylized prompts. For Cartoon Heaven, it produced an output with genuine personality — characters that feel lively, color choices that commit to the cartoon world, and a sense of fun that some of the more technically-focused models missed entirely.

    It isn't the most technically refined output in this round. Motion fidelity under scrutiny shows some of the characteristic Pika artifacts on complex movement. But as a complete viewing experience, the Pika 2.2 output has energy that several higher-ranked technical performers don't.

    Pika 2.2 is worth serious consideration for cartoon and stylized content where creative feel matters more than technical perfection. For this specific prompt and this specific community vote, it placed mid-pack — respectable, not decisive.

    Cartoon style fidelity: Good
    Motion quality: Moderate
    Best use case: Stylized, expressive content; creator-friendly workflows


    Seedance 2.0

    Seedance 2.0 is ByteDance's multimodal AI video model, announced in February 2026. It generates up to 15 seconds of synchronized audio-video output from text and image inputs using a unified architecture that handles composition, motion, camera planning, and audio in a single generation pass. Independent benchmarks consistently place it near the top of AI video leaderboards, ranking #1 for image-to-video with audio on Artificial Analysis.

    On Cartoon Heaven, Seedance 2.0 did something the other five models didn't: it understood the brief at a deeper level.

    The output commits fully to the cartoon world — not just visually, but physically. Motion has the right kind of exaggeration. Characters move with snap and anticipation. Color choices feel designed for the scene rather than generated. The spatial logic of the world holds together in a way that suggests the model understood "cartoon" as a set of rules, not just a visual style.

    The audio-video sync, a known Seedance 2.0 strength, also contributed. Sound that lands in rhythm with character movement adds a layer of perceived quality that's hard to achieve with post-processed audio. In a cartoon context, that sync matters enormously.

    This is not a model that happens to do cartoons. On the evidence of this round, it's a model that excels at animation-style generation specifically.

    Cartoon style fidelity: Excellent
    Motion quality: Excellent
    Best use case: Cartoon animation, stylized video, any prompt where expressive motion and audio sync matter


    The Vote Results

    Model Vote Share
    Seedance 2.0 75%
    Kling 1.6 Pro —
    Pika 2.2 —
    PixVerse C1 —
    Hailuo 2 —
    LTX-2.3 —

    Seedance 2.0 took 75% of all community votes — a result that isn't close by any measure. In a six-way competition where votes are genuinely split across strong models, a three-quarter majority is a decisive statement from the community.

    This was Seedance 2.0's first win in The Coliseum, ending a run of three consecutive victories by Veo 3.1 Fast. The margin suggests it wasn't a close call.


    What the Numbers Tell Us

    A few things stand out from this round.

    Cartoon style is a genuine differentiator. This prompt separated the field more decisively than any photorealistic prompt would. Models that dominate on cinematic, realistic prompts — Hailuo 2 is the clearest example — placed near the bottom not because of poor quality, but because the prompt was outside their design center. Choosing the right model for the right task matters more than choosing the "best" model overall.

    Seedance 2.0's audio-video sync is a real advantage on animation. Cartoon content is uniquely sensitive to sound timing. The synchronized audio output Seedance 2.0 produces natively creates a quality perception gap that's difficult to close in post-production.

    75% is unusual. In most Coliseum rounds, votes distribute more evenly across the top two or three models. A 75% result in a six-model field suggests near-universal agreement — not a divided community leaning slightly one way, but a clear winner that the majority of voters agreed on immediately.

    Model selection is prompt-dependent. If you're choosing an AI video model based on a single benchmark or general "best of" list, this round is a reminder that context changes the answer. Kling 1.6 Pro is one of the most capable models in production use today. On this prompt, Seedance 2.0 wasn't close.


    Watch Every Round at Kodex1

    The Coliseum at Kodex1 runs continuous head-to-head battles across the top AI video models. Every round uses the same prompt across all models — no curation, no cherry-picking — and real community votes decide the winner.

    Watch the current Coliseum battle at Kodex1 →

    Past rounds are archived. You can watch every model's output side by side, see the vote history, and follow individual AI Directors building their catalog on the platform. Kodex1 is free — no subscription, no paywall.