AI Image & Video Generation Statistics 2026 — “Video Has Replaced Images” Is the Wrong Take
AI video is the fastest-moving creative frontier, but image generation remains the high-volume layer. Google says Nano Banana has produced more than 50 billion images and Gemini now generates 150 million+ images every day, while Flow’s video milestones climbed from tens of millions to hundreds of millions in months.
This report compiles the most decision-useful AI image generation statistics, AI video generation statistics and generative AI creator statistics for Midjourney, Sora, Veo and adjacent platforms—covering adoption, output volume, pricing, model capability, creator ROI, safety and the current 2026 competitive landscape.
Images generated with Google’s Nano Banana family by May 2026.
Images the Gemini app was generating daily by August 2026.
Images and videos created in Google Flow by February 2026.
Date OpenAI discontinued the Sora web and app experiences in 2026.
U.S. video creators in Adobe Express research who had used AI video generation or editing tools.
Current Midjourney monthly subscription range across four plans.
// Table of Contents
The “AI Video Replaced Images” Narrative Misses the Scale Difference
Video is where capability is changing fastest, but images still dominate volume. Google reported more than 50 billion Nano Banana images generated by May 2026 and 150 million+ Gemini images per day by August. By comparison, Flow’s October 2025 milestone was 275 million generated videos.
That is not a weakness for AI video. Video is structurally more expensive, takes longer to render, consumes more compute and has more failure modes across motion, continuity, audio and physics. The relevant comparison is value per generation, not raw output count.
Cumulative images generated with Nano Banana models by Google I/O 2026.
Daily image generations in the Gemini app as of August 2026.
Images generated with Nano Banana in Southeast Asia over the prior year.
ChatGPT users who created images in the first week of OpenAI’s native image launch.
Images those ChatGPT users generated during that first week.
Videos generated in Google Flow within roughly five months of launch.
Videos generated by Sora Android users in its first 24 hours.
Combined images and videos created in Flow by February 2026.
Do not compare “images generated” with “video-app users.” Image counts measure outputs, Sora download figures measure installs, Flow milestones measure creations, and Midjourney pricing/spec data measures capability. They answer different questions.
The second correction is more dramatic: Sora is no longer a current consumer competitor. OpenAI discontinued Sora’s web and app experiences on April 26, 2026. The Sora API remains scheduled for discontinuation on September 24, so the model still appears in developer pricing at the time of this report, but Sora should be treated as a 2025–2026 product-cycle case study rather than a live standalone app.
Scheduled discontinuation date for the Sora API.
Time OpenAI’s Sora lead said the iOS app took to reach one million downloads.
Appfigures estimate for Sora iOS downloads across its first seven days.
Sora U.S. weekly active users by the end of Q4 2025, per Sensor Tower.
Sora’s peak weekly U.S. mobile revenue in Sensor Tower’s Q4 analysis.
The current 2026 market therefore looks less like a three-way race and more like a layered stack: Midjourney remains a specialist visual creation platform, Veo is embedded across Google’s consumer, creator, Workspace and developer products, while OpenAI has shifted its consumer visual emphasis toward ChatGPT Images.
Generation Volume Is Compounding Faster Than Product Categories Can Stay Separate
The cleanest trend comes from Google Flow. It launched as an AI filmmaking tool, but by February 2026 Google had merged image generation, editing and animation into the same workspace—so later milestones combine modalities.
| Date | Milestone | Measurement | Source |
|---|---|---|---|
| Jul. 2025 | 40M+ | Veo 3 videos across Gemini + Flow in seven weeks | Google Labs · Flow / Veo 3 expansion · Jul. 2025 |
| Aug. 2025 | 100M | Flow videos since May launch | Google Labs · Flow 100M video milestone · Aug. 2025 |
| Oct. 2025 | 275M+ | Flow videos generated | Google Labs · Veo 3.1 / Flow update · Oct. 2025 |
| Feb. 2026 | 1.5B+ | Flow images + videos after workspace expansion | Google Labs · Flow creation milestone · Feb. 2026 |
| May 2026 | 50B+ | Nano Banana images generated across Google surfaces | Google I/O 2026 · Sundar Pichai keynote · May 2026 |
| Aug. 2026 | 150M+/day | Gemini image generation daily run rate | Google · Gemini 1B monthly users / creative usage · Aug. 2026 |
The milestones also show why “AI image generator” and “AI video generator” are becoming less useful product categories. Google Flow now handles images, video, editing and creative agents in one workspace. Midjourney creates images and turns them into video. ChatGPT Images combines generation and editing conversationally. The workflow is becoming multimodal even when the underlying models remain specialized.
Nano Banana 2 reached one billion images in nearly half the time required by Nano Banana 1.
Geographic availability Google cited for Flow after its expansion.
Google AI Pro expansion for Veo 3 access cited during the 2025 rollout.
Free Veo 3.1 generations available monthly in Google Vids to anyone with a Google account.
Veo video generations available to Google AI Ultra/Workspace AI Ultra accounts in Vids.
Google AI Ultra subscription price announced at I/O 2026.
The output curve also changes how teams should think about creative inventory. Image generation can support dozens or hundreds of variants for concepting, localization, product merchandising and testing because the marginal cost and wait time are low. Video remains more selective: every additional second introduces motion, continuity and audio decisions that can invalidate an otherwise strong generation. In practice, images increasingly operate as the exploration layer that feeds references, storyboards and keyframes into higher-cost video workflows.
Visual generation is moving from a destination product to a capability embedded everywhere: chat apps, design suites, work tools, APIs, social publishing and mobile creation. The winner may be the workflow with the lowest friction—not the model with the most spectacular demo.
Content Demand, Speed and Multi-Model Workflows Are Driving Adoption
Creative professionals are not adopting AI because demand disappeared. They are adopting it because demand increased faster than production capacity. Adobe’s March 2026 survey found content demand had at least doubled over two years, while campaign refresh cycles continued to compress.
Creative professionals who said AI lets them produce content faster.
Average time creative pros said AI saves them.
Average share of projects on which creative pros said they use AI.
Creative pros using more than one AI model on each asset.
Video creators who had used AI video generation or editing.
Creators saying creative AI helps them produce content faster.
Speed is only one driver. Creative teams are also using AI because marketing cycles are changing. Adobe found more than eight in ten marketers update campaigns weekly and more than a third update daily. That makes inexpensive variation, resizing, localization, thumbnail creation and iterative testing economically valuable even when the final hero asset still requires substantial human craft.
Creative pros saying generative AI has made their work better.
Creative pros saying it is easier to get the results they want from generative AI.
Creative pros choosing specific AI models for specific tasks.
Creative pros comparing outputs across models to select a result.
Creative pros concerned about maintaining brand continuity with AI.
Creative pros concerned about commercial safety.
Creative pros concerned about homogenized creative output.
Creative pros concerned AI efficiency will cause unsustainable content demands.
When creative capacity rises, demand often rises with it. AI can return 17 hours per week and still leave teams overloaded if output expectations expand faster than the saved time.
This is also why standalone model loyalty is weakening. When 91% of surveyed creative professionals use multiple models on an asset, the category behaves more like a creative supply chain than a winner-take-all SaaS market. Midjourney may win style exploration, Veo may win a controlled video shot, and another model may win typography or inpainting—all inside the same campaign.
Midjourney, Veo and Sora Now Represent Three Different Product Paths
The 2026 comparison is not apples-to-apples. Midjourney is a live specialist creation platform, Veo is a live model family distributed across multiple Google surfaces, and Sora’s consumer product has been discontinued. ChatGPT Images is now OpenAI’s live mainstream visual-creation surface.
| Platform | 2026 status | Core output | Key current capability | Source |
|---|---|---|---|---|
| Midjourney | Live | Images + image-to-video | V8.2 current default image model | Midjourney Docs · Model Versions · Aug. 2026 |
| Google Veo | Live | Video + native audio | Veo 3.1 family across Flow, Gemini, APIs and Workspace | Google DeepMind · Veo 3.1 model & evaluations · 2025–2026 |
| OpenAI Sora app/web | Discontinued | Video + audio | Consumer experience ended Apr. 26, 2026 | OpenAI Help · Sora discontinuation · Aug. 2026 |
| Sora API | Sunsetting | Video | Scheduled to end Sep. 24, 2026 | OpenAI Help · Sora discontinuation · Aug. 2026 |
| ChatGPT Images | Live | Images + editing | Images 2.0 available across ChatGPT tiers | OpenAI Help · ChatGPT Images FAQ · Aug. 2026 |
Midjourney’s current image product moved unusually quickly in 2026. V8.0 alpha arrived in March, V8.1 in April, V8.1 became the default in June, and V8.2 became default on July 24. V8.1 introduced native 2K HD generation and sharply reduced render time and GPU cost.
Date Midjourney V8.2 became the default version.
V8.1 standard-job speed improvement versus earlier versions.
Current team size stated on Midjourney’s official About page.
Midjourney FY2024 revenue reported by Forbes citing PitchBook; an older commercial baseline.
This distribution difference matters commercially. A specialist tool can win creators through quality and workflow depth even without owning a billion-user consumer surface, while a platform company can make generation ubiquitous by embedding it into products users already open every day. That means model comparisons should include acquisition cost and workflow switching cost alongside prompt adherence or visual fidelity. A technically stronger model can still lose the production decision if getting an asset into the team’s existing review, storage and publishing system creates too much friction.
Veo’s strategy is distribution breadth. The same model family appears in Flow, the Gemini app, Google Vids, AI Studio, the Gemini API and enterprise infrastructure. That gives Google both consumer reach and a developer route to production, while Midjourney’s differentiation remains a tightly integrated creative environment and aesthetic tooling.
Video Costs More Because Every Second Is a Compute Budget
The pricing data explains why image generation reaches much larger output volumes. Image subscriptions can support effectively unlimited relaxed generation; production video is metered by seconds, resolution, GPU minutes or monthly generation allowances.
| Video model | 720p | 1080p | 4K | Source |
|---|---|---|---|---|
| Veo 3.1 Standard + audio | $0.40/sec | $0.40/sec | $0.60/sec | Google AI for Developers · Veo API pricing · Aug. 2026 |
| Veo 3.1 Fast + audio | $0.10/sec | $0.12/sec | $0.30/sec | Google AI for Developers · Veo API pricing · Aug. 2026 |
| Veo 3.1 Lite + audio | $0.05/sec | $0.08/sec | Not supported | Google AI for Developers · Veo API pricing · Aug. 2026 |
| Sora 2 | $0.10/sec | — | — | OpenAI API · Image & Sora pricing · Aug. 2026 |
| Sora 2 Pro | $0.30/sec | $0.70/sec | — | OpenAI API · Image & Sora pricing · Aug. 2026 |
Sora pricing is shown because the API was still listed and available at the report date; OpenAI says it will be discontinued September 24, 2026.
Monthly Fast GPU allocation on Midjourney Basic.
Company-revenue threshold requiring Midjourney Pro or Mega.
At current API prices, an 8-second 720p clip costs about $0.40 on Veo 3.1 Lite, $0.80 on Veo Fast, and $3.20 on Veo Standard. That spread changes which model makes sense for storyboards, ad variations or final production shots.
Midjourney’s video economics use GPU minutes rather than per-second API billing. A default four-video SD batch costs eight GPU minutes; the comparable HD batch costs 26. That makes experimentation volume, resolution and batch size important cost controls before a creator starts extending clips.
The cost model should change how teams evaluate creative AI. A cheap first draft is valuable only if retries, upscale passes, extensions and manual cleanup stay inside the budget. The relevant procurement metric is therefore cost per approved asset or usable second—not list price per generation.
The Frontier Is Control: Audio, Continuity, Resolution and Editing
The 2026 frontier is no longer “can a model make a plausible clip?” The competition has moved to native audio, subject consistency, first/last frames, object insertion, reference images, scene extension, vertical output and lower-cost iteration.
MovieGenBench prompts used in Google’s Veo 3.1 text-to-video preference evaluation.
Image/text pairs used in Veo’s image-to-video evaluation.
Prompts used in Veo’s text-to-video-with-audio evaluation.
Examples used in a Veo reference-image control benchmark.
Examples used for one Ingredients-to-Video internal comparison.
Examples used for Scene Extension comparison.
Examples used for another first/last-frame or object-control evaluation.
Text-to-video prompts in the Veo 3.1 Lite vs. Fast evaluation.
Image-to-video prompts in the Veo 3.1 Lite evaluation.
Veo 3.1 Lite T2V overall win rate versus Veo 3.1 Fast.
Veo 3.1 Lite I2V overall win rate versus Veo 3.1 Fast.
Maximum Veo 3.1 Standard/Fast API output tier listed in current pricing.
Midjourney’s video model occupies a different niche: image-first motion. Clips start at five seconds, each extension adds four seconds, and creators can extend up to four times for a 21-second maximum. Output is 480p by default or 720p on qualifying plans, with low/high motion controls plus looping and end-frame options.
Maximum Midjourney video length after extensions.
Midjourney HD video resolution on Standard, Pro and Mega.
Default number of videos generated per video prompt.
Midjourney HD output dimensions for a 16:9 starting image.
Midjourney HD output dimensions for a square starting image.
Raw visual quality is converging. Production usefulness now depends on whether a model can preserve characters, follow references, maintain physics, synchronize audio, hit a required aspect ratio and survive multiple rounds of editing.
OpenAI’s Sora launch also pushed provenance into the product layer. Sora outputs used visible and invisible provenance signals plus C2PA metadata, while Google uses SynthID on Veo outputs. As photorealism improves, provenance is shifting from a policy footnote to a distribution requirement.
Industry-standard provenance metadata embedded in Sora videos under OpenAI’s deployment approach.
Maximum video duration highlighted for the original Sora research model in 2024.
Crash-free rate OpenAI reported for the Sora Android app.
Time OpenAI’s team took to build the production Sora Android app.
Codex usage consumed while building Sora for Android.
Creators Use AI Heavily—but Still Want the Final Say
Adobe’s June 2026 Creators’ Toolkit survey of more than 16,000 creators shows a mature pattern: AI is embedded in workflow, but human judgment remains the publishing gate. The strongest users are treating AI as leverage rather than authorship-by-default.
Creators using creative AI who say it accelerated business or follower growth.
Creators describing creative AI as integrated or essential to how they work.
Creators saying AI made them feel more confident, professional or serious about their work.
Creators saying AI makes them feel more secure about their future.
Creators saying AI-assisted content consistently performs better.
Creators who find it harder to stand out and blame sheer content volume.
Creators saying AI-generated content makes unique voices harder to distinguish.
Creators saying they can compete with larger teams/studios more effectively since using AI.
Creators saying AI-assisted work still reflects their unique voice.
Creators saying human judgment remains essential to creative taste.
Creators saying AI outputs typically need moderate or extensive editing before sharing.
Creators saying AI gives them more freedom to experiment before pitching.
Creators saying AI gives confidence to pursue more ambitious projects.
Creators saying the final creative decision should always remain human.
“Voice, taste and judgment remain what set great creators apart.”Mike Polner, VP & Head of Product Marketing for Creators, Adobe · Creators’ Toolkit Report 2026
Control requirements are equally concrete. Creators say they would give future creative agents more independence if they could review, edit or undo actions, see what the system is doing, and limit the data/tools it can access. Disclosure remains unresolved: audience expectations are rising faster than creator behavior.
Creators wanting review/edit/undo at any point before granting an agent more independence.
Creators wanting transparency into what an agent is doing and why.
Creators who would use agent-freed time to learn new creative skills.
Creators who would spend freed time on higher-level ideas and direction.
Creators saying audience expectations for AI disclosure are increasing or steady.
Creators believing audiences can tell when AI was meaningfully involved.
Creators saying copyright protection for AI-assisted work is important.
The underlying profession is not disappearing into a generic prompt role. Adobe’s separate June 2026 research finds AI skills increasingly written into existing creative job descriptions. The market is rewarding people who can combine craft judgment with model fluency, not simply people who can generate the most outputs.
Creative AI Is Global, but Distribution Still Depends on Product Surfaces
Country-level active-user counts for Midjourney, Veo and image/video generators are rarely disclosed on a comparable basis. The strongest public data instead shows availability, regional output and survey reach.
Southeast Asia
5B Nano Banana images generated over the prior year in Google’s 2026 regional report.
Veo 3 rollout
Google cited access for Pro users in 150+ countries during its 2025 expansion.
Sora iOS launch
Initially launched in the United States and Canada.
Sora Android launch
Launched across seven named markets in November 2025.
Adobe creators
2026 Creators’ Toolkit survey covered eight countries.
Adobe’s creator survey covered the U.S., U.K., France, Germany, South Korea, Japan, India and Australia. Its separate Future of Creative Work research surveyed creative professionals in the U.S., U.K. and Japan, making it useful for attitudes but not a global market-share estimate.
Creators in Adobe’s 2026 global Creators’ Toolkit survey.
U.S., U.K. and Japan coverage in Adobe’s June 2026 creative-attitudes research.
Creative professionals and other creatives in that Adobe attitudes study.
U.S. video creators surveyed by Adobe Express.
Adobe Express video-survey respondents who were solo creators.
Availability is not adoption. “140+ countries” tells you Flow can be used broadly; it does not tell you where generation volume or paid demand is highest. Avoid turning product-availability lists into fabricated country market shares.
The Best Creative-AI Workflows Are Multi-Model and Human-Edited
The strongest 2026 practitioner data points toward a hybrid workflow: use AI to accelerate drafts and production, switch models based on task strengths, then apply human taste, brand controls and editing before publishing.
Video-specific research shows similar economics. Seventy-one percent of surveyed video creators had already used AI video tools. More than half saved at least 30 minutes per video, with the largest teams much more likely to save two hours or more. The gain then showed up in publishing consistency, engagement and sponsorship outcomes.
AI-video users in Adobe Express research using AI video tools weekly.
AI-video users who had only tried the tools once or twice.
Creators saving more than 30 minutes per video with AI.
Creators saving more than four hours per video.
Teams of 4+ saving more than two hours per video.
Teams of 2–3 saving more than two hours per video.
Solo creators saving more than two hours per video.
Creators reporting higher audience watch time or completion rates.
Creators with brand deals reporting more sponsorships after AI adoption.
Creators reporting higher client satisfaction.
Creators reporting higher CPMs or sponsorship rates.
Video creators planning to increase AI-tool spending over the next year.
Creators planning to maintain current AI-tool budgets.
The data does not reward “one prompt, publish immediately.” The strongest pattern is faster drafting + model selection + human refinement. Speed matters, but judgment remains the conversion layer between generated output and publishable creative.
Adobe’s 2026 job-posting research reinforces that conclusion. AI skills are becoming an explicit requirement inside existing creative roles rather than simply replacing those roles. The most valuable operators increasingly look like creative directors of model portfolios: they know which system to use, what to keep, what to reject and how to maintain a consistent brand across generated variations.
Increase in U.S. creative-professional job postings over the six-month comparison window.
Share of creative job postings explicitly listing AI skills in the earlier wave.
Share explicitly listing AI skills by April 2026.
Mid-market/enterprise creative job postings explicitly requiring AI skills.
Solo/small-business creative postings explicitly requiring AI skills.
Mid-market/enterprise share of creative-professional hiring in the later wave.
Three Charts That Explain Visual AI in 2026
The most useful charts show video creation velocity, current generation economics, and the creator workflow pattern that turns raw model capability into usable output.
Google Flow video generation milestones
Video-only milestones before Flow expanded into a broader image + video workspace.
8-second 720p API video cost
Published per-second pricing × 8 seconds. Sora API is scheduled to sunset Sep. 24, 2026.
Creator workflow signals
Selected Adobe 2026 creator and creative-professional survey measures.
Together, the charts show the market’s operating logic: generation volume can compound rapidly, but costs widen sharply with video quality and resolution; meanwhile, the professional workflow still depends on human review. The economic winner is not the system that generates the most—it is the one that produces the most usable output per dollar and per hour of review.
10 Evidence-Based Moves for Creative Teams in 2026
The strongest visual-AI operating model is not “replace the creative stack.” It is to use cheap image iteration, selective video generation, multiple specialist models and explicit human approval—then measure whether faster production actually improves business outcomes.
Gemini generates 150M+ images daily while video APIs are metered per second. Use images for high-volume exploration and reserve video compute for shots that earn it. Evidence: S002 Evidence: S066
OpenAI ended Sora web/app access on April 26, 2026 and plans to sunset the API September 24. Update comparison pages and procurement matrices accordingly. Evidence: S004 Evidence: S015
91% of surveyed creative pros use more than one model per asset, and 50% choose different models for different tasks. Evidence: S036 Evidence: S043
57% of creators say AI output usually needs moderate/extensive editing and 85% say the final creative decision should remain human. Evidence: S130 Evidence: S133
An 8-second 720p clip can range from about $0.40 on Veo 3.1 Lite to $3.20 on Veo 3.1 Standard before retries. Evidence: S072 Evidence: S066
Midjourney V8.1 uses 0.8 GPU minutes for SD versus 1.3 minutes for HD; video HD batches also cost far more than SD. Evidence: S059 Evidence: S058
Creative pros report 17 hours saved per week on average, while 56% of video creators save more than 30 minutes per video. Evidence: S034 Evidence: S166
84% of creative pros worry about brand continuity and 64% worry about homogenized output. Build references, review criteria and reusable style systems. Evidence: S045 Evidence: S047
85% of creators say disclosure expectations are rising or steady, while Sora embeds provenance metadata into output. Evidence: S139 Evidence: S115
AI skills rose from 10% to 15% of creative job postings in Adobe’s comparison, reaching 18% at mid-market/enterprise companies. Evidence: S181 Evidence: S182 Evidence: S183
For measurement, separate experimentation from production. Track how many generations are created, how many survive human review, how many require substantial repair, and how many are ultimately published or shipped. The ratio between generated and approved assets is often more informative than generation speed alone. As model output becomes cheaper, review time, brand risk and rights management can become the dominant costs—even when the invoice for inference falls.
A useful 2026 stack can therefore be intentionally heterogeneous. Midjourney can remain the exploration and visual-style engine, Veo can handle high-control video with audio, ChatGPT Images or Nano Banana can handle conversational editing and high-volume image workflows, and traditional tools still own finishing, compositing, typography, color and brand governance. The important decision is not model loyalty—it is where each generation step creates enough value to justify its cost and review burden.
How This Report Was Built
Data is current as of August 26, 2026 and uses US English to match SEOScaleUp’s current statistics-page standard. We prioritized first-party product disclosures, model documentation, pricing pages, safety documentation and research from Google/Google DeepMind, OpenAI, Midjourney and Adobe. Secondary sources were used only where the primary company did not publish a comparable public number—notably Appfigures launch estimates via TechCrunch, Sensor Tower mobile-app data, and Forbes/PitchBook’s dated Midjourney revenue baseline.
Sources used: Google I/O 2026 keynote; Google Gemini 1B monthly-users update; Google Flow updates from July, August and October 2025 and February/May 2026; Google Veo 3.1 model page and Veo 3.1 Lite model card; Google Gemini API/Veo pricing; Google Vids April 2026 update; Google AI subscription updates; Google Nano Banana 2 and Southeast Asia reports; OpenAI native image-generation API launch; OpenAI ChatGPT Images 2.0 and Images FAQ; OpenAI API pricing; OpenAI Sora 2 launch, Sora safety documentation, Sora Android engineering case study and Sora discontinuation FAQ; TechCrunch/Appfigures Sora launch reporting; Sensor Tower Q4 2025 photo/video app data; Midjourney model-version, video and subscription documentation plus official company information; Forbes/PitchBook Midjourney company profile; Adobe Creators’ Toolkit Report 2026; Adobe/Advanis creative-professional survey; Adobe Express AI video creator survey; Adobe Future of Creative Work research and creative job-posting analysis.
Conflicting or non-comparable metrics are not averaged. Google’s 50B Nano Banana figure is cumulative images; Gemini’s 150M is a daily image run rate; Flow’s 1.5B combines images and videos; Sora’s app figures are downloads, weekly users or videos generated; Midjourney metrics are subscription, model and GPU-capacity statistics. Older 2024–2025 figures are labeled when they provide necessary trend or launch context. Vendor-run benchmark results are identified as vendor evaluations rather than neutral cross-market rankings.
