Doc (beta)
Guides
AI Model Selection Guide

AI Model Selection Guide

Last Updated: 2026-08-17

Overview

This guide helps you choose the right image model, video model, or processing tool for each task on WeShop AI.

Supported image models: GPT Image 2, Nano 2, Nano Pro, Seedream 5 Pro, Seedream 5.0 Lite, Midjourney, and Z-Image. Supported video models: Kling V3, Kling V3 Omni, MiniMax H3, Seedance 2.5, Seedance 2.0, and Seedance Mini.

For API fields and Agent parameters, see API Reference and Agent. If you have any questions, please email: hi@weshop.ai.

Usage Principles

Assess the Task Before Choosing a Model

Model selection should follow this order:

  1. Determine whether the task is pure text-to-image, reference-image generation, local editing, or multi-image consistency. Midjourney and Z-Image support text-to-image only; exclude both whenever image input or image-editing is required.
  2. Define the delivery tier: draft, internal review, live campaign, or premium commercial delivery.
  3. Confirm core constraints: subject consistency, Asian aesthetics, photographic quality, speed, cost, or batch scale.
  4. Choose the lowest-cost model that satisfies the core constraints.
  5. On generation failure, optimize inputs first, then upgrade the model; do not rely on repeated blind retries.

Default Recommendations

RequirementPrimary ModelFallback Model
Readable text (including Chinese and other languages)GPT Image 2 MediumNone; adjust tier only when Medium fails or the user explicitly requests it
Multilingual images (including Chinese), image translation, and localizationGPT Image 2 MediumNone; use text-free base images plus post-production layout when necessary
Artistic creation, inspiration exploration, and illustrationMidjourneyGPT Image 2
Photorealistic text-to-image or Chinese cultural elementsZ-ImageGPT Image 2
General high-quality generationGPT Image 2Nano 2
Subject or character consistencyGPT Image 2Nano Pro
Low-cost batch draftsGPT Image 2 (Low)Seedream 5.0 Lite (only when lighting requirements are high or Asian aesthetics apply)
Higher-quality fast iterationNano ProGPT Image 2
Commercial photography with high lighting requirementsSeedream 5 ProGPT Image 2
Asian subjects, fashion, and e-commerce photographySeedream 5 ProSeedream 5.0 Lite
Routine commercial images with high lighting requirements or Asian aestheticsSeedream 5.0 LiteGPT Image 2

Relative Price Order

Under the same pricing basis, this guide uses the following relative price order (high to low):

GPT Image 2 High
> Nano Pro
> Nano 2 = Seedream 5 Pro
> GPT Image 2 Medium = Seedream 5.0 Lite
> GPT Image 2 Low
Price TierModel or Quality Mode
P5: HighestGPT Image 2 High
P4Nano Pro
P3Nano 2, Seedream 5 Pro
P2GPT Image 2 Medium, Seedream 5.0 Lite
P1: LowestGPT Image 2 Low

The tiers above reflect relative invocation cost only; they are not equivalent to model quality ranking. Actual total cost is also affected by output dimensions, generation count, retry count, and platform billing rules.

Midjourney and Z-Image share the same per-image price. Their relative pricing against other P1–P5 models is not yet confirmed, so they are not included in the tiers above. Midjourney always outputs 4 images per call; when calculating per-invocation cost, bill for 4 images, not a single image.

Capability Declaration Rules

  • This guide describes recommended usage boundaries; it does not represent official hard limits of the underlying models.
  • GPT Image 2 supports Low, Medium, and High tiers. Other models and tools do not support the quality parameter.

Hard Routing Rules for Text-in-Image Tasks

  • Whenever the output image must include titles, poster copy, packaging text, signage, menus, UI text, or other readable characters—including Chinese—route uniformly to GPT Image 2 Medium.
  • Multilingual generation (including Chinese-inclusive combinations), image text translation, language replacement, and localized layout also route fixedly to GPT Image 2 Medium.
  • Z-Image is not a candidate for Chinese text requirements; consider Z-Image only for pure text-to-image tasks that explicitly require Chinese cultural elements.
  • Nano 2, Nano Pro, Seedream 5.0 Lite, and Seedream 5 Pro are not primary or fallback models for text-in-image generation.
  • GPT Image 2 tiers follow a progressive rule: Low for text layout validation; Medium as the default for routine text and translation deliverables; High only when Medium fails acceptance or the user explicitly requests it.
  • Text content must be given verbatim in the prompt, with explicit language, casing, placement, font style, and typographic hierarchy.
  • If text accuracy still does not meet delivery requirements, switch to “text-free base image + layout tool for text”; do not blindly switch among models that lack text input capability.

Hard Limits for Text-to-Image-Only Models

  • Midjourney and Z-Image perform text-to-image only; input must consist solely of text prompts and generation parameters.
  • When any reference image, product image, portrait, background, mask, first/last frame, or source image for editing is present, do not call Midjourney or Z-Image.
  • Model swap, outfit change, virtual try-on, multi-product compositing, subject consistency, inpainting, expand image, background removal, and image translation are not pure text-to-image tasks; handle them with the corresponding image models or tools.
  • Midjourney and Z-Image do not support quality; do not pass this field in requests.
  • Midjourney always outputs 4 images per call; the invocation layer must not request a different count, and the business layer must receive, bill, and filter all 4 results.
  • Unit pricing, relative pricing against other models, and speed for both models are not yet confirmed; Z-Image’s unified quality tier is also pending confirmation. Complete configuration before integration; do not infer values.

Seedream Series Usage Limits

  • Consider Seedream 5.0 Lite or Seedream 5 Pro only when the task has high lighting requirements or explicitly requires Asian aesthetics.
  • For general generation, subject consistency, routine product shots, and standard commercial images that do not meet either condition above, default to GPT Image 2 or the Nano series.
  • Seedream 5.0 Lite is for routine deliverables; Seedream 5 Pro is for the same class of needs with higher lighting, material, or final-delivery requirements.
  • Text-in-image rules take precedence: whenever readable text, multilingual output, or translation is required—including Chinese—use GPT Image 2 Medium.

GPT Image 2 Tier Hard Rules

GPT Image 2 must use progressive tier selection. Do not jump to High simply because the task description mentions “multiple images,” “subject consistency,” “commercial image,” or “final delivery.”

TierUsage Conditions
LowLowest-cost drafts, composition validation, text layout validation
MediumDefault tier; routine deliverables, subject consistency, multi-angle/multi-pose, standard commercial assets, multilingual and translation tasks
HighMedium output explicitly failed acceptance, or user explicitly requests premium final delivery, print-grade detail, etc.

Enforced sequence:

  1. Call Medium first for routine tasks; for batch tasks, generate 1 test image with Medium first.
  2. After the test image passes identity, outfit, composition, text, or product-structure acceptance, complete remaining outputs with Medium.
  3. Upgrade to High only when Medium results are actually substandard or the user explicitly requests High.
  4. High cannot compensate for missing reference assets. If outfit, subject angle, or product structure is not visible in references, add assets first.
  5. When upgrading to High, record the specific reason—for example, “Medium subject detail failed acceptance”—not merely “high quality required.”

Image Models

GPT Image 2

Basic Information

AttributeDetails
PositioningGeneral-purpose flagship model
Core strengthsVersatility, complex instruction following, subject and character consistency
Default qualitymedium
Output quality positioningHigh: high; Medium: medium-high; Low: basic
Recommended delivery tierInternal review, live campaigns, commercial delivery

Recommended Use Cases

ScenarioRationale
Subject consistency tasksSame subject with pose, outfit, or scene changes
Complex scene compositingMultiple subjects, multiple constraints, explicit spatial relationships
Product showcaseBalance product, environment, and marketing composition
Text-in-image generationPoster titles, packaging text, signage, and other readable characters
Multilingual and image translationTranslate or replace image text while preserving layout and visual hierarchy
Photorealistic contentNatural, credible overall imagery
General design needsWhen requirements are not yet biased toward a specific aesthetic
Multi-round editingIterating on the image from feedback

Tier and Alternative Boundaries

ScenarioAssessmentRecommended Choice
Draft seeking lowest per-call priceUse this model’s Low tier directlyGPT Image 2 Low
Top-tier Asian fashion photographySpecialized aesthetics and lighting can be further improvedSeedream 5 Pro
Quick composition validation onlyMedium or High not neededGPT Image 2 Low

For multilingual tasks, call and accept each target language separately; do not output multiple language versions in one request.

Nano 2

Basic Information

AttributeDetails
PositioningHigh-throughput fast generation model
Core strengthsSpeed, batch exploration; priced below Nano Pro, above GPT Image 2 Medium/Low
quality parameterNot supported; do not pass
Output quality positioningMedium
Recommended delivery tierDrafts, proof of concept, initial A/B test variants

Recommended Use Cases

  • Batch generation of visual directions and composition candidates.
  • Early A/B testing for ad creative.
  • Social content sketches, storyboards, and mood boards.
  • Tasks needing fast batch exploration with quality above the lowest-cost draft tier.
  • Filter prompts, then redo on a higher-quality model.

Not Recommended as Primary

ScenarioReasonRecommended Alternative
Stable subject identity requiredHigher consistency riskGPT Image 2
Fine product textureSmall details easily lostGPT Image 2; Seedream series when lighting requirements are high or Asian aesthetics apply
Print or premium client deliveryOutput detail may be insufficientGPT Image 2; Seedream 5 Pro when lighting requirements are high or Asian aesthetics apply
Complex multi-subject scenesInstruction load too highGPT Image 2
Readable text in imageNano 2 does not handle text tasksGPT Image 2 Medium

Nano Pro

Basic Information

AttributeDetails
PositioningQuality-enhanced fast generation model
Core strengthsBetter detail and instruction execution than Nano 2 while keeping fast iteration; priced just below GPT Image 2 High
quality parameterNot supported; do not pass
Output quality positioningHighest
Recommended delivery tierAdvanced drafts, internal review, lightweight live assets

Recommended Use Cases

ScenarioNotes
High-quality concept imagesWhen drafts should not be too rough
Medium-scale A/B testingQuality matters more than lowest cost
Social media assetsSimple scenes, short delivery cycles
Design reviewClearer material and lighting expression needed

Note

Readable text in images: do not call Nano Pro; route uniformly to GPT Image 2 Medium.

Boundary with Nano 2

CriterionNano 2Nano Pro
GoalExplore as many directions as possibleMature frames from fewer candidates
PriorityCost and speedBalance of quality and speed
Recommended output count4–8 directions2–4 directions
Best stageDivergenceConvergence and review

Seedream 5 Pro

Basic Information

AttributeDetails
PositioningSpecialized model for demanding lighting and Asian aesthetics
Core strengthsFine lighting control, material rendering, and Asian commercial aesthetics; same price tier as Nano 2
quality parameterNot supported; do not pass
Output quality positioningMedium-low
Recommended delivery tierPremium advertising, brand key visuals, print-grade delivery

Recommended Use Cases

  • Commercial photography with high requirements for light direction, tonal hierarchy, shadow transitions, rim light, or ambient light.
  • Automotive, jewelry, watch, fragrance, and luxury ads requiring studio lighting and texture sculpting.
  • Product shots where glass, metal, leather, and similar materials need fine lighting expression.
  • Asian subjects, Asian fashion, beauty, and localized e-commerce visuals.
  • Brand key visuals emphasizing both Asian aesthetics and professional studio lighting.

Not Recommended as Primary

ScenarioReasonRecommended Alternative
Fast large-scale divergencePositioned for final deliverables; throughput may not suit explorationNano 2
Routine internal review with high lighting or Asian aestheticsPro not required directlySeedream 5.0 Lite / Nano Pro
Complex multi-round instruction editingGeneral editing flexibility may not be optimalGPT Image 2
Strong abstract or experimental illustrationSpecialized photography strengths underusedGPT Image 2
Readable text in imageSeedream 5 Pro does not handle text tasksGPT Image 2 Medium
General tasks without high lighting or Asian aesthetic requirementsOutside Seedream series specializationGPT Image 2 / Nano series

Seedream 5.0 Lite

Basic Information

AttributeDetails
PositioningRoutine production tier for demanding lighting and Asian aesthetics
Core strengthsLighting control, Asian aesthetics, and commercial usability; same price tier as GPT Image 2 Medium
quality parameterNot supported; do not pass
Output quality positioningLow
Recommended delivery tierE-commerce, social ads, and routine brand assets with high lighting requirements or Asian aesthetics

Recommended Use Cases

ScenarioNotes
E-commerce product shots with high lighting requirementsPrecise control of light source, shadow, depth, and product texture
Asian fashion photographyLocalized subjects and aesthetic expression
Asian subject contentPrioritize localized aesthetics
Ad assets with high lighting requirementsRoutine batch production and commercial deliverables
Pro-tier previewConfirm lighting, composition, and Asian aesthetic direction first

Note

Readable text in images: do not call Seedream 5.0 Lite; route uniformly to GPT Image 2 Medium.

Note

If the task does not have high lighting requirements and does not emphasize Asian aesthetics, do not call Seedream 5.0 Lite; use GPT Image 2 or the Nano series instead.

Boundary with Seedream 5 Pro

CriterionSeedream 5.0 LiteSeedream 5 Pro
Primary goalDaily commercial productionTop-tier commercial delivery
Generation stageDirection confirmation and routine deliverablesFinal key visuals
Recommended batch sizeMediumSmall curated set
BudgetMedium-highHighest tier
Detail requirementsProduct pages, social adsLarge-format print, luxury, fine materials

Midjourney

Basic Information

AttributeDetails
Task typeText-to-image only
PositioningTop-tier artistic creation model
Core strengthsBest for artistic creation, inspiration, illustration, and style exploration
quality parameterNot supported; do not pass
Output quality positioningHigh (not a request parameter)
Output countAlways 4 images per call; cannot be changed
Per-image priceSame as Z-Image; relative pricing against other models pending confirmation

Recommended Use Cases

ScenarioNotes
Artistic creationEmphasis on visual expression, mood, form, and style language
Inspiration searchEarly concept exploration, mood boards, direction divergence
IllustrationClear artistic-style illustration work
Stylized visualsSurreal, experimental, or non-photorealistic expression

Prohibited Scenarios

ScenarioReasonAlternative
Any image input or reference-image generationText-to-image onlyGPT Image 2, Nano, or Seedream series per reference-image task
Outfit change, model swap, product compositingReference-image editing or multi-image compositingVirtual Try-On or reference-capable image models
Image translation and multilingual localizationRequires reading and modifying source imageGPT Image 2
Any readable text (including Chinese)Fixed text-task routingGPT Image 2 Medium
Photorealism as core goal without artistic intentNot this model’s primary routeZ-Image or GPT Image 2

Midjourney always generates 4 images. Even when the business needs only 1, receive all results and filter afterward.

Z-Image

Basic Information

AttributeDetails
Task typeText-to-image only
PositioningPhotorealistic text-to-image with Chinese cultural elements
Core strengthsStrong photorealism; better understanding and rendering of Chinese cultural symbols
Artistic capabilitySome artistic expression, but limited style diversity
quality parameterNot supported; do not pass
Output quality positioningPending confirmation
Per-image priceSame as Midjourney; relative pricing against other models pending confirmation

Recommended Use Cases

ScenarioNotes
Photorealistic text-to-imageSubjects, products, or scenes without reference images
Chinese cultural elementsBetter understanding of traditional architecture, festivals, attire, patterns, and cultural symbols

Prohibited Scenarios

ScenarioReasonAlternative
Any image input or reference-image generationText-to-image onlyGPT Image 2 or other reference-capable models
Image text translation, language replacement, or localizationRequires reading and editing source imageGPT Image 2
Any readable text, including Chinese and multilingualFixed text-task routingGPT Image 2 Medium
Rich artistic styles, top-tier art creation, or broad inspiration explorationLimited style diversity despite some artistic abilityMidjourney
Subject consistency, outfit change, multi-image compositingRequires reference-image capabilityGPT Image 2, Nano series, or Virtual Try-On

Video Generation Models

Video model pricing, quality tiers, supported duration, resolution, and concurrency limits are not yet confirmed. The following defines model boundaries and invocation structure only; concrete enums must be overridden by actual platform configuration.

Kling V3 / V3 Omni

Basic Information

AttributeKling V3Kling V3 Omni
ModelKling V3Kling V3 Omni
PositioningControllable video generationMulti-reference, multimodal video generation
Core inputsText, first frame, last frameText, multiple images, and other reference inputs supported by the platform
Primary scenariosFirst/last frame, single-image animation, controllable camera motionMulti-image consistency, complex references, video editing
Price / quality tierPending confirmationPending confirmation

Kling V3 Recommended Use Cases

ScenarioNotes
First/last frame videoExplicit control of start and end states
Single-image animationTurn product, portrait, or scene stills into motion
Controllable cameraClear requirements for subject motion and camera motion
Product showcaseRotation, push-in, pull-back, detail reveal
Medium-complexity motionBalance controllability and motion amplitude

Kling V3 Omni Recommended Use Cases

ScenarioNotes
Multi-image referenceReference character, scene, product, or style images simultaneously
Complex subject consistencyConstrain subject or product with multiple reference assets
Video referenceExtract motion, camera, or visual traits; capability per platform
Video-to-videoContinue generation or editing from existing video; capability per platform
Multimodal tasksInput combinations beyond standard first/last frame

Version Boundaries

CriterionKling V3Kling V3 Omni
Single first frameRecommendedUsable but may be overkill
Explicit last frameRecommendedPer platform support
Multi-image referenceDo not useRecommended
Video referenceDo not useRecommended
Simple controllable cameraRecommendedUsable
Complex reference combinationsInsufficient capabilityRecommended

Not Recommended as Primary

ScenarioReasonRecommended Alternative
Dance, sports, and large-amplitude motionDynamic range not primary strengthMiniMax H3
Audio-driven high-dynamic performanceNeeds motion and rhythm capabilityMiniMax H3
Low-cost concept validation onlyMay be overkillSeedance Mini; confirm pricing first
Long continuous narrativeSingle-shot duration limitedSplit shots and composite

MiniMax H3

Basic Information

AttributeDetails
PositioningHigh-dynamic motion video model
Core strengthsLarge-amplitude motion, motion continuity, dynamic camera
Price / quality tierPending confirmation

Recommended Use Cases

ScenarioNotes
Large-amplitude motionDance, sports, running, jumping, action performance
High-dynamic cameraFast movement, pronounced pose change, dynamic camera work
Motion continuityEmphasis on action trajectory and rhythmic flow
Audio-driven videoAudio combined with subject or video reference; support per platform
Complex pose changesMultiple consecutive action phases in one shot

Not Recommended as Primary

ScenarioReasonRecommended Alternative
Precise first/last frame controlCore need is start/end state, not dynamic rangeKling V3
Multi-image subject consistencyNeeds more reference constraintsKling V3 Omni
Static or subtle motionHigh-dynamic strengths unusedKling V3
Primarily artistic audiovisual expressionAesthetics and AV expression prioritizedSeedance 2.5

Required visual input combinations, duration, and format for audio-driven generation must follow the actual platform; do not assume pure-audio requests always work.

Seedance 2.5 / 2.0 / Mini

Basic Information

AttributeSeedance 2.5Seedance 2.0Seedance Mini
ModelSeedance 2.5Seedance 2.0Seedance Mini
PositioningAudiovisual expression and artistic deliverablesStandard creative videoLightweight concept validation
Primary scenariosAudio-visual sync, music video, artistic shortRoutine artistic expression, creative shortsDrafts, lightweight trials, fast exploration
Price / quality tierPending confirmationPending confirmationPending confirmation

Version Selection

RequirementPreferred VersionNotes
Audio-visual sync or music videoSeedance 2.5Audio and visual rhythm as core
Artistic final deliverableSeedance 2.5Creative and aesthetic expression prioritized
Routine artistic shortSeedance 2.0Does not need 2.5 enhancements
Concept validation and draftsSeedance MiniLightweight exploration; confirm cost on platform

Recommended Use Cases

  • Music videos, mood shorts, and artistic social content.
  • Non-photorealistic, emotional, or stylized visual expression.
  • Shorts centered on rhythm, color, transitions, and visual mood.
  • Seedance Mini for validating style, rhythm, and camera direction.

Not Recommended as Primary

ScenarioReasonRecommended Alternative
Precise product structure showcaseControllability over artistic expressionKling V3
Multi-image character consistencyNeeds stronger reference constraintsKling V3 Omni
Dance, sports, large-amplitude motionDynamic range prioritizedMiniMax H3
Precise first/last frameNeeds explicit start/end statesKling V3
Long complex narrativeSingle shot cannot guarantee continuityStoryboard generation then composite

Video Model Quick Decision

Does the task require large-amplitude motion or high-dynamic action?
├─ Yes → MiniMax H3
└─ No: Does it require multi-image, video, or complex references?
   ├─ Yes → Kling V3 Omni
   └─ No: Does it require precise first/last frame or controllable product showcase?
      ├─ Yes → Kling V3
      └─ No: Is audio-visual sync or artistic expression the core goal?
         ├─ Yes → Seedance 2.5
         └─ No: Is it only lightweight concept validation?
            ├─ Yes → Seedance Mini
            └─ No → Seedance 2.0

Video Routing Priority

  1. Precise first/last frame and product structure control take priority over artistic style → Kling V3.
  2. Multi-reference input takes priority over standard first/last frame → Kling V3 Omni.
  3. Large-amplitude motion takes priority over identity and first/last frame control → MiniMax H3.
  4. When audio-visual sync and artistic expression are prioritized → Seedance 2.5.
  5. Seedance 2.0 for routine artistic shorts; Mini for lightweight validation.
  6. When conflicting goals coexist, split shots or have the user confirm the highest priority first.

Image Processing Tools

Use processing tools when the task is not image or video generation. For API parameters, see Agent.

Tool Overview

ToolChoose WhenChoose Image Model Instead When
Remove BGNeed transparent PNG cutout onlyNeed new background, content edit, or generative fill
Expand ImageNeed aspect ratio or platform size change via cropNeed upscaling, outpainting, or content modification
Virtual Try-OnSingle product, at most 1 single-person model, at most 1 backgroundMultiple products, multiple people, multi-product outfit, or readable text in output

Virtual Try-On Routing Rules

Do not call Virtual Try-On when:

ScenarioAlternative
Two or more products at onceImage model workflow with multi-reference compositing
Multiple model images or multiple people in model imageImage models
Missing product imageAdd product image first
Readable text in outputGPT Image 2 Medium

When Virtual Try-On does not apply, route to image models per this guide: GPT Image 2 for text or multilingual tasks; Seedream series when lighting requirements are high or Asian aesthetics apply; otherwise GPT Image 2 or Nano series by consistency, quality, speed, and price.

Tool Selection Rules

Need to remove the original background?
├─ Yes → Remove BG
└─ No: Is it virtual try-on, model swap, or background swap?
   ├─ Yes: Is there exactly 1 product image, at most 1 single-person model image, and at most 1 background image?
   │  ├─ Yes → Virtual Try-On
   │  └─ No → Image model workflow
   └─ No: Need to change aspect ratio, platform size, or crop/composition?
      ├─ Yes → Expand Image
      └─ No → Image generation or editing workflow

Model Selection & Routing

Image Model Comparison Matrix

“Output quality positioning” in this table uses internal assessment criteria; it is not a request parameter and does not equal fit for a specific style. Only GPT Image 2 supports the quality parameter. Seedream enters the candidate set only for tasks with high lighting requirements or Asian aesthetics. Midjourney and Z-Image support text-to-image only; they share the same per-image price, and Midjourney always outputs 4 images per call.

Model / TierOutput Quality Positioning (Not a Parameter)SpeedPrice TierConsistencyPhotographic QualityRecommended Stage
GPT Image 2 HighHighMediumP5 (highest)HighHighFinal deliverables, complex high-quality tasks
Nano ProHighestFasterP4MediumMedium-highConvergence, internal review
Nano 2MediumFastP3Medium-lowMediumBatch divergence, drafts
Seedream 5 ProMedium-lowSlowerP3Medium-highVery highHigh-requirement delivery with demanding lighting or Asian aesthetics
GPT Image 2 MediumMedium-highMediumP2HighHighGeneral deliverables, routine complex tasks
Seedream 5.0 LiteLowMediumP2Medium-highHighRoutine production with demanding lighting or Asian aesthetics
GPT Image 2 LowBasicFasterP1 (lowest)Medium-highMedium-highLowest-cost validation and drafts
MidjourneyHigh; top-tier for artistic creationPending confirmationSame per-image price as Z-Image; tier pendingN/A for reference consistencyStrong and diverse artistic stylePure text-to-image art, inspiration, illustration; always 4 images
Z-ImagePending confirmation; strong photorealismPending confirmationSame per-image price as Midjourney; tier pendingN/A for reference consistencyStrong photorealism; limited artistic style diversityPure text-to-image photorealism and Chinese cultural elements

Image Model Quick Decision Tree

Does the task include any image input, reference image, or image-editing requirement?
├─ Yes → Exclude Midjourney and Z-Image
│  ├─ Image translation, multilingual, or readable text editing → GPT Image 2 Medium
│  └─ Other tasks → Choose GPT Image 2, Nano, or Seedream series by consistency, cost, lighting, and Asian aesthetics
└─ No (pure text-to-image)
   ├─ Artistic creation, inspiration, or illustration → Midjourney
   ├─ Any readable text or multilingual (including Chinese) → GPT Image 2 Medium
   ├─ Chinese cultural elements → Z-Image
   ├─ Photorealism as core without artistic emphasis → Z-Image
   └─ Other general tasks → GPT Image 2 Medium; then adjust by price, batch, and style needs

Recommended Image Production Pipelines

WorkflowRecommended Pipeline
General marketing assetsGPT Image 2 Low validation → Nano 2 batch divergence → GPT Image 2 Medium deliverables
E-commerce with demanding lighting or Asian aestheticsNano Pro composition → Seedream 5.0 Lite deliverables → Seedream 5 Pro curated hero images
Brand key visuals with demanding lighting or Asian aestheticsSeedream 5.0 Lite preview → Seedream 5 Pro delivery
Subject consistency seriesGPT Image 2 Medium: generate 1 test first → complete with Medium after acceptance; upgrade to High only if Medium fails
Pure text-to-image art concepts and illustrationMidjourney divergence → manual filter and prompt convergence
Pure text-to-image Chinese cultural photorealismZ-Image generation → cultural element acceptance; switch to GPT Image 2 Medium if readable text is also required

Model Fallback

When the primary model does not meet requirements, use the following fallback chains.

Image Model Fallback Chain

Original Model or TierFirst FallbackFinal Fallback
GPT Image 2 HighGPT Image 2 MediumGPT Image 2 Low
Nano ProNano 2GPT Image 2 Medium
Nano 2GPT Image 2 MediumGPT Image 2 Low
Seedream 5 ProSeedream 5.0 LiteGPT Image 2 Medium
Seedream 5.0 LiteGPT Image 2 MediumGPT Image 2 Low
MidjourneyNo automatic cross-model fallbackManual choice
Z-ImageGPT Image 2 MediumText-free base + post layout

Readable text, multilingual, translation, and image text editing tasks do not use the cross-model fallback chain above; default to GPT Image 2 Medium.

Video Model Fallback Chain

Original ModelFirst FallbackFinal Fallback
Kling V3 OmniKling V3Manual handling
Kling V3Seedance 2.0Seedance Mini
MiniMax H3Kling V3Seedance 2.0
Seedance 2.5Seedance 2.0Seedance Mini
Seedance 2.0Seedance MiniManual handling
Seedance MiniManual handling

Do not silently downgrade when fallback would break core capabilities such as multi-image reference, precise last frame, or large-amplitude motion.