There is a version of the AI story that most people know. ChatGPT arrives, the big labs scramble, models multiply, benchmarks fly. I covered that story in a previous post, tracing the language model race from November 2022 to last week. There have been more releases since, because of course there have.
But there is another race that was running in parallel the whole time, and in some ways this one was more visceral, more immediate, and landed closer to home for anyone who works in marketing, design, or any kind of creative field. You did not need to understand tokens or reasoning benchmarks to feel this one. You could actually see it.
This is the story of AI image generation, how it started, how it went cinematic, and what it means for the rest of us.
The AI image generation story started earlier than you think
Most people place the start of the AI image story at 2022, when Stable Diffusion and DALL-E 2 went public. But the groundwork was laid earlier than that.
In January 2021, OpenAI released the original DALL-E. It was research-grade, not publicly available, and the outputs were strange and pixelated by today’s standards. The famous “avocado chair” image from that era looked nothing like a real photograph, but it proved something important. A machine could take words and turn them into a picture. Not perfectly. Not even close to perfectly. But it could do it.

Midjourney launched its first private demo in September 2021, built by a small team of engineers with a philosophy that stood apart from the rest. Rather than chasing photorealism and technical benchmarks, founder David Holz wanted to build something closer to an artistic tool, something that expanded what people could imagine rather than simply replicated what already existed.
By early 2022, Midjourney V1 was in the hands of a small invited group. The outputs were raw, dreamlike, and strange. Hands were wrong. Faces melted. But something about the aesthetic stopped people in their tracks. This was not photography. It was not illustration. It was something new.
2022: The year it went public
Three things on AI image generation happened in 2022 that changed everything.
In April, DALL-E 2 launched to a waitlist. OpenAI had dramatically improved on the original, sharper, more coherent, more controllable. For the first time, people outside research labs could generate images from text descriptions and get results that looked like something you might actually use.
In July, Midjourney opened its public beta, accessible entirely through Discord. You joined a server, typed a command, and four image options appeared. The community exploded almost immediately. By the time Midjourney V3 launched in the same month, the Discord server had grown to one million users, overtaking the Fortnite and Minecraft servers as the largest on the platform.
In August, Stability AI released Stable Diffusion as fully open-source. This was the moment that changed the structural nature of the race. Not because it was the most beautiful output, but because anyone could download it, run it on their own computer, modify it, fine-tune it on their own images, and build on top of it for free. The genie was out of the bottle in a way that could not be put back.
By November, Midjourney V4 had arrived and the quality leap was significant enough to go viral. AI-generated images were showing up in news articles, social media feeds, and creative briefs. The conversation had started.
2023: The quality crosses a threshold
If 2022 was about proving the concept, 2023 was about AI image generation crossing a line that most people did not realise existed until it had already been crossed.
In April 2023, Midjourney V5 arrived. Not “impressive for AI.” Just impressive. Portrait photography that required a second look to determine whether a camera had been involved. Concept art that working professionals could use without embarrassment. Designers and marketers started incorporating it into real workflows, not as a novelty but as a genuine production tool.
Adobe entered in March 2023 with Firefly, and the positioning was deliberate. Trained only on licensed Adobe Stock images and public domain content, Firefly was the copyright-safe option for enterprises that could not afford legal risk. It reached one billion images generated in its first three months. The creative industry had a legitimate tool it could actually defend in front of a client or a legal team.
In October 2023, DALL-E 3 launched inside ChatGPT. The text rendering improved dramatically. For the first time, an AI image model could put legible words inside a generated image with reasonable reliability. That single capability unlocked an entire category of use cases that had previously been off the table.
Also in 2023, Getty Images sued Stability AI for copyright infringement. A group of artists filed a class action lawsuit against Midjourney, Stable Diffusion, and DeviantArt. The creative industry had gone from watching to litigating.
The aesthetic tool had become an industry threat.
2024: Video walks into the room
Everything above had been still images. Remarkable, disruptive, legally contested still images. Text-to-video had been around since 2022, when Meta’s Make-A-Video and Google’s Imagen Video showed it was theoretically possible. The results were blurry, three seconds long, and nowhere near usable. Runway Gen-2 in early 2023 was the first commercially available tool that produced something you could actually work with. But in February 2024, OpenAI showed the world something that changed the scale of the conversation entirely.
Enter Sora.
The demo videos were not “good for AI.” They were cinematic. A woman walking through neon-lit Tokyo rain, her reflection catching in wet pavement. A woolly mammoth charging across snowy grasslands with a camera pull that felt like a David Attenborough production. A slow drone shot moving through a fantasy cityscape. People who had spent careers on film sets watched the demos in awe.
Sora was not publicly available yet. OpenAI gave access to a small group of filmmakers and researchers. But the signal was clear. AI had learned to make video that looked real.
Meanwhile, a company called Runway had been building away in the background, producing the tool that professionals actually started using. Runway Gen-2 had shipped in 2023. Gen-3 Alpha arrived in 2024 and was the first public model to produce multi-second video with stable motion, coherent camera movement, and usable output. While Sora generated headlines, Runway became the working production standard.
From China, Kuaishou launched Kling. The motion quality on human subjects was striking, and the cost was significantly lower than Western equivalents. The pattern from the language model race was repeating itself. Chinese challengers arriving faster than the West had anticipated, matching capability at lower cost.
Pika launched its 2.0 version with a feature it called Pikaffects, stylised creative transformations that made it the tool of choice for social media creators who wanted something quick, expressive, and feed-native. It wasn’t trying to be production grade.
2024 into 2025: Hollywood fights back
The IP war that had started with still images escalated sharply when video entered the picture.
CAA, the talent agency representing some of the biggest names in Hollywood, called Sora a serious and harmful risk to their clients’ intellectual property. The Motion Picture Association demanded OpenAI take immediate and decisive action over copyright infringement. Writers, illustrators, voice actors, photographers, and musicians filed lawsuits across multiple jurisdictions. The EU AI Act kicked in with transparency and disclosure requirements for AI-generated content.
The creative industry had stopped being scattered and anxious. By this point it was organised, funded, and litigating on multiple fronts simultaneously.
OpenAI released Sora publicly in December 2024. The public version was more constrained than the demos suggested. Users found it difficult to access, expensive to run, and inconsistent in ways the curated demos had not revealed. The gap between the highlight reel and the actual product was significant.
Adobe continued adding AI capabilities to Photoshop, Illustrator, and Premiere, threading Firefly through the tools creative professionals already used rather than asking them to learn something new. For agencies and brand teams with existing Adobe workflows, this was the path of least resistance.
Google released Imagen 3 and began integrating video generation through what would become Veo. The Google approach was integration over standalone product, with AI capabilities threaded through Workspace, Gemini, and YouTube rather than a separate app to learn.
2025: The race goes wide
If 2024 had four or five serious players, 2025 had a dozen. The race went from a sprint to a field event.
Runway Gen-4 arrived in spring 2025 and became the professional production standard. Character consistency across scenes, precise camera controls, a reference image system that let you maintain visual continuity shot to shot. Ad agencies and film studios started building it into actual workflows, not just experimenting. If you worked in production in 2025, you knew Runway.
Google Veo 3 landed and changed the video conversation fundamentally. For the first time, a video generation model could produce native audio in the same pass as the footage. Ambient sound, dialogue, and music generated simultaneously rather than stitched together afterwards. The implications for content production were immediate. You could generate a scene with a person speaking and have their voice, the room tone, and background sound all arrive together.
Midjourney continued iterating through V7, with a strong emphasis on character consistency and a new feature called Consistent Characters that let you generate the same person across multiple scenes and angles. A capability that had been the exclusive domain of expensive production pipelines suddenly became available to anyone with a subscription.
Black Forest Labs released Flux, an open-source image model that stunned the community with its photorealism and prompt accuracy. Flux became the reference benchmark for what open-source image generation could achieve, and the community built an ecosystem of fine-tuned models around it rapidly.
Stable Diffusion 3.5 arrived as Stability AI’s most capable model, though the company’s financial troubles had begun to slow its competitive pace. The open-source ecosystem it had spawned remained vibrant even as the company itself struggled.
From China, ByteDance released image and video generation capabilities through its creative tools. Alibaba’s Wan 2.1 and 2.2 released open-source weights for a multimodal image and video model that held up against closed Western peers with few obvious weaknesses. Tencent’s HunyuanVideo had already released strong open-source video weights in December 2024. The Chinese open-source video models were arriving faster and more capable than anyone had forecasted.
Ideogram and Recraft carved out specific niches. Ideogram became known for graphic design and text-heavy image generation, posters, logos, and social graphics where accurate text rendering mattered. Recraft focused on clean vector-style illustration for designers who wanted something closer to professional graphic design output than photographic realism.
By this point, the AI image generation landscape had grown crowded enough that a new category emerged to make sense of it all. Aggregator platforms like Higgsfield, Fal.ai, Replicate, and Leonardo.ai gave users access to multiple models in one place, paying per generation rather than juggling five separate subscriptions. Higgsfield pulled together Runway, Kling, Luma, and image models like GPT Image and Recraft under one dashboard. Replicate let developers run open-source models on demand, paying per second of compute. Leonardo.ai built its own fine-tuned models on top of Stable Diffusion, aimed squarely at game artists and concept designers. When the market needs platforms just to manage access to all the platforms, that tells you something about how far things have come.
HeyGen and Synthesia dominated the talking-head video category. Not cinematic video, but AI avatars that could present in multiple languages with accurate lip-sync. The enterprise training video and corporate communications market found its tool.
2026: The market settles, mostly
Sora was shut down in March 2026. OpenAI announced the web and app experiences would be discontinued in April, the API in September. The economics had not worked. The IP battles had not helped. The most hyped AI product of 2024 became one of the more expensive product exits of 2026. Teams that had built workflows around it were told to migrate to Veo, Kling, Runway, or Seedance.
Midjourney reached V8.1. The company remained one of the most unusual success stories in the entire AI race, profitable since its second month, no venture capital funding, roughly 107 employees, $500 million in annual revenue, and valued at around $10 billion.
GPT Image 2 from OpenAI became the go-to for complex prompt-following image generation. If you needed an image that precisely matched a detailed description, including legible text, GPT Image 2 was the benchmark.
The market as of mid-2026 looks something like this. For image generation: Midjourney V8.1 for artistic and editorial output, GPT Image 2 for precise prompt execution, Adobe Firefly Image 3 for commercially safe enterprise use, Flux for open-source and developer workflows, Ideogram for text-heavy graphic design work, Recraft for clean illustration. For video: Google Veo 3.1 as the all-around quality benchmark, Runway Gen-4 as the professional production tool, Kling 3.0 for high-motion scenes and human subjects, Pika for social media content, HeyGen and Synthesia for talking-head and corporate video. Most professional teams in 2026 use two models, not one, a workhorse for volume work and a premium model for hero shots.
What this actually means
Here is the thing about the image and video race that makes it different from the language model race. The LLM story is largely about productivity, who can answer questions faster, write code more accurately, synthesise information more reliably. It affects how we work.
The image and video story cuts deeper than that. It goes after craft. It goes after the skills that creative professionals spent years developing. The photographer who spent a decade learning light. The illustrator who built a visual style that clients commissioned specifically. The motion graphics designer who understood how to make something feel cinematic. All of these skills are now being partially replicated by tools that cost a few cents per generation.
That does not mean the skills have no value. The best creative work still requires taste, judgement, direction, and a point of view that no prompt can substitute for. But the floor of what is acceptable has moved dramatically, and the economics of creative production have shifted in ways that are still working themselves out.
The lawsuits are not over. The IP questions are not resolved. The question of what happens to the next generation of designers and photographers and filmmakers who are learning their craft in an environment where AI can produce a decent version of almost anything is genuinely open.
I do not have a clean answer to that. But I think it is the more important question to sit with than “which model should I use this week.”
Here is the sum of it all. In under five years, AI image generation went from generating a blurry avocado on a chair to producing cinematic video with native audio. The tools went from research labs to Discord servers to professional production workflows. The lawsuits went from a few concerned artists to Hollywood studios and government regulators. And the cost of generating a five-second video clip dropped from $2.50 to less than 30 cents in under eighteen months. Whatever you think about where this is heading, the pace of what has already happened is worth sitting with for a moment.
Filter by country
| Company | Country | Model | Month | Year | What it could do / key leap forward |
|---|---|---|---|---|---|
| OpenAI | US | DALL-E v1.0 Image |
Jan | 2021 | OpenAI’s first public image model. Research-grade, not widely accessible. Famous for the “avocado chair” prompt. Proved a machine could generate an image from a text description. Blurry and strange by today’s standards, but the concept worked. |
| OpenAI | US | DALL-E 2 v2.0 Image |
Apr | 2022 | Dramatically sharper and more controllable than the original. First widely available consumer image generator from a major lab. Launched to a waitlist and proved public demand was real. Introduced inpainting — the ability to edit specific parts of an image with text. |
| Stability AI | UK | Stable Diffusion v1.4 Image |
Aug | 2022 | The model that changed the structural nature of the race. Fully open-source, free to download and run locally. Anyone with a decent GPU could generate images, fine-tune on their own data, or build products on top of it. Spawned an entire ecosystem of community models and tools. |
| Midjourney | US | Midjourney v1 to v3 Image |
Jul | 2022 | Open beta launched July 2022, accessible only through Discord. Outputs were raw, dreamlike and strange. Hands were wrong, faces melted. But the aesthetic stopped people in their tracks. Discord server hit one million users within months, overtaking Fortnite and Minecraft servers as the largest on the platform. |
| Meta | US | Make-A-Video v1.0 Video |
Sep | 2022 | One of the first text-to-video models from a major lab. Could generate short animated clips from text prompts. Results were blurry and only a few seconds long, but proved text-to-video was theoretically possible. Research release, not publicly available. |
| US | Imagen Video v1.0 Video |
Oct | 2022 | Google’s early text-to-video research model. Higher resolution than Make-A-Video but still limited to short clips. Research-only, never publicly released. Set the foundation for Veo several years later. | |
| Midjourney | US | Midjourney v4 to v5 Image |
Nov | 2022 | V4 went viral in November 2022 with a dramatic quality leap. V5 arrived in March 2023 and crossed a threshold — not “impressive for AI” but just impressive. Portrait photography requiring a second look to determine if a camera was involved. The tool creative professionals started actually using. |
| Adobe | US | Firefly v1.0 Image |
Mar | 2023 | The copyright-safe option. Trained only on licensed Adobe Stock images and public domain content. Reached one billion images generated in its first three months. Integrated across Photoshop, Illustrator, and Express, making it the enterprise default for teams that could not afford legal risk. |
| Runway | US | Gen-2 v2.0 Video |
Mar | 2023 | The first commercially available text-to-video tool that produced something actually usable. Short clips, imperfect physics, but coherent enough for creative use. Became the de facto standard for video professionals before Sora arrived and long after it departed. |
| OpenAI | US | DALL-E 3 v3.0 Image |
Oct | 2023 | Launched inside ChatGPT. Dramatically improved text rendering — for the first time, an AI image model could put legible words inside a generated image reliably. Unlocked an entire category of use cases including posters, infographics, and branded content that had previously been off the table. |
| Pika | US | Pika 1.0 v1.0 Video |
Nov | 2023 | Consumer-friendly video generation that made the category accessible to social media creators. Strong on stylised output — watercolour, stop-motion, cel-animation. The tool for fast iterations and feed-native content rather than production-grade work. |
| OpenAI | US | Sora v1.0 Video |
Feb | 2024 | The model that changed the conversation. Demo videos were cinematic — a woman walking through Tokyo rain, a woolly mammoth charging across snowy grasslands. People who had spent careers on film sets watched in awe. Not the first text-to-video model, but the first to produce genuinely cinematic output. Shut down March 2026 due to brutal economics and IP battles. |
| Black Forest Labs | US | Flux 1.0 v1.0 Image |
Aug | 2024 | Open-source image model that stunned the community with photorealism and prompt accuracy. Founded by former Stability AI researchers. Became the new reference benchmark for open-source image generation. The community rapidly built an ecosystem of fine-tuned models on top of it. |
| Luma Labs | US | Dream Machine v1.0 Video |
Jun | 2024 | Fast, accessible image-to-video tool known for cinematic motion on short clips. Popular with creators who needed speed over precision. Public release made it one of the first widely accessible video generation tools outside of Runway. |
| Kuaishou | CN | Kling v1.0 Video |
Jun | 2024 | China’s first serious challenger to Western video generation tools. Motion quality on human subjects was striking and the cost was significantly lower than Western equivalents. Could generate up to two minutes of video. Proved the East-West gap in video generation was closing faster than the press acknowledged. |
| Runway | US | Gen-3 Alpha v3.0 Video |
Jun | 2024 | The first public model to consistently produce multi-second video with stable motion and coherent camera movement. More stylistic control than Gen-2. While Sora generated headlines, Runway Gen-3 became the working production standard for ad agencies and studios. |
| US | Veo v1.0 Video |
May | 2024 | Google’s serious entry into video generation. Supported text, image, and video input. Matched Sora on quality benchmarks in independent testing. Deployed through Gemini and Google AI Studio rather than a standalone app. | |
| MiniMax | CN | Hailuo AI v1.0 Video |
Sep | 2024 | MiniMax’s video generation tool, known for creative and expressive motion on unusual prompts. Ranked near the top of the Artificial Analysis image-to-video leaderboard in mid-2026 alongside Seedance. Strong on dynamic scenes where other models produce stiff motion. |
| Tencent | CN | HunyuanVideo v1.0 Video |
Dec | 2024 | Open-source video model from Tencent that released strong weights on Hugging Face. Independent benchmarks placed it around Kling 1.5 on complex motion and camera control, ahead of Runway Gen-3 and Luma Dream Machine. Became a popular base model for the open-source community. |
| OpenAI | US | Sora public Video |
Dec | 2024 | Public release of Sora, eight months after the demo. The public version was more constrained than the demos suggested — difficult to access, expensive, and inconsistent. The gap between the highlight reel and the actual product was significant. Shut down in March 2026. |
| Midjourney | US | Midjourney v6 to v7 Image |
Jan | 2025 | V6 introduced dramatically improved photorealism and text rendering. V7 added Consistent Characters, allowing the same person to be generated across multiple scenes and angles. A capability previously exclusive to expensive production pipelines, now available on a subscription. |
| Runway | US | Gen-4 v4.0 Video |
Apr | 2025 | The professional production standard. Character consistency across scenes, precise camera controls, and a reference image system for visual continuity shot to shot. The tool ad agencies and film studios built into actual production workflows. Not the flashiest, but the most reliable for deliverables. |
| US | Veo 3 v3.0 Video |
May | 2025 | Changed the video conversation fundamentally. First model to generate native audio in the same pass as the video — ambient sound, dialogue, and music produced simultaneously rather than stitched together afterwards. A game-changer for content production workflows. | |
| OpenAI | US | GPT Image 2 v2.0 Image |
Apr | 2026 | The benchmark for precise prompt-following image generation. Launched April 21, 2026 with a built-in reasoning mode — it thinks before it draws. Near-perfect text rendering including non-Latin scripts, up to 2K resolution output. If you needed an image that matched a detailed description precisely, including legible text, GPT Image 2 was the reference. |
| Alibaba | CN | Wan 2.1 / 2.2 v2.2 Video |
Feb | 2025 | Open-source multimodal image and video model from Alibaba that held up against closed Western peers with few obvious weaknesses. Wan 2.2 introduced a mixture-of-experts diffusion backbone. Widely adopted by the open-source community as a free, capable alternative to paid closed models. |
| Stability AI | UK | Stable Diffusion 3.5 v3.5 Image |
Oct | 2024 | Stability AI’s most capable model, offering strong photorealism and improved prompt adherence over earlier versions. The open-source ecosystem built around Stable Diffusion remained vibrant even as the company faced financial difficulties. Full control for users who do not mind tinkering. |
| Black Forest Labs | US | Flux 2 v2.0 Image |
Nov | 2025 | Major leap from experimental image generation toward true production-grade visual creation. Available through managed APIs and open-weight checkpoints. Achieved state-of-the-art across realism, art, and anime styles with 17 billion parameters. The open-source community’s reference benchmark for image quality. |
| Ideogram | US | Ideogram v3 v3.0 Image |
Mar | 2025 | The specialist for text-heavy image generation. Posters, logos, social graphics, and anything where accurate text rendering inside an image matters. Carved out a distinct niche from photorealistic generators by focusing on graphic design use cases. |
| Recraft | US | Recraft V3 v3.0 Image |
Jan | 2025 | Clean vector-style illustration for designers who want something closer to professional graphic design output than photographic realism. Strong on brand consistency and style control. The designer’s choice when photorealism is not the goal. |
| HeyGen | US | HeyGen 3.0 v3.0 Video |
Mar | 2025 | The enterprise standard for talking-head AI video. Multilingual dubbing with accurate lip-sync, AI avatars that can present in dozens of languages. Became the tool of choice for corporate training, localisation, and marketing teams producing high-volume video content without production crews. |
| Synthesia | US | Synthesia 2.0 v2.0 Video |
Feb | 2025 | AI avatar video platform focused on enterprise learning and development. Create a training video in one language and localise it to 140 languages without reshooting. Adopted by Fortune 500 companies for internal communications and employee training at scale. |
| ByteDance | CN | Seedance v1.0 Video |
Jan | 2026 | ByteDance’s video generation model that topped the Artificial Analysis leaderboard in mid-2026 alongside Alibaba’s HappyHorse-1.0. Seedance 2.0 added Universal Reference for character and object consistency across scenes. Proof that ByteDance’s creative AI ambitions extend well beyond TikTok. |
| US | Veo 3.1 v3.1 Video |
Oct | 2025 | Better audio, more natural motion, stronger character preservation. Available via Gemini app, Vertex AI, and Google’s Flow filmmaking tool. Ranked third on the Artificial Analysis leaderboard in mid-2026. The all-around quality benchmark for video generation. | |
| Kuaishou | CN | Kling 3.0 v3.0 Video |
Feb | 2026 | Four entries in the top 10 of the Artificial Analysis leaderboard in mid-2026. Best-in-class for complex human motion when driven from a reference still. Lower cost than Western equivalents at comparable quality. The clearest sign the East-West video gap has closed entirely. |
| Adobe | US | Firefly Image 4 v4.0 Image |
May | 2026 | Adobe’s most capable image model. 24 billion assets generated across all Firefly versions by mid-2025. Integrated across 7 Creative Cloud apps reaching 32.5 million subscribers. The enterprise safe choice, commercially licensed training data, used by 72% of Fortune 500 design teams. |
| Lightricks | IL | LTX-Video v2.3 Video |
Mar | 2026 | Open-source video model from Israeli company Lightricks, optimised to run on consumer-grade NVIDIA GPUs. LTX-2.3 was one of the few frontier-quality models that did not require enterprise-level hardware. Made serious video generation accessible to independent creators and small teams. |
| Alibaba | CN | HappyHorse 1.0 v1.0 Video |
Apr | 2026 | Topped the Artificial Analysis image-to-video leaderboard in April 2026 alongside ByteDance’s Seedance 2.0, displacing all Western models from the top two slots. The most visible signal yet that Chinese video generation had overtaken Western labs at the frontier. |
| Midjourney | US | Midjourney v8.1 Image + Video |
Mar | 2026 | V8.1 reached with 500M annual revenue, 107 employees, no VC funding, valued at $10 billion. Added video generation via V1 Video model in June 2025. Now expanding into Midjourney Medical, an AI tool for medical imaging. The most improbable success story in the entire AI race. |
How are you navigating AI in your creative or marketing work? I would love to hear your perspective in the comments below.