AI Model Selection Guide
Last Updated: 2026-08-17
Overview
This guide helps you choose the right image model, video model, or processing tool for each task on WeShop AI.
Supported image models: GPT Image 2, Nano 2, Nano Pro, Seedream 5 Pro, Seedream 5.0 Lite, Midjourney, and Z-Image. Supported video models: Kling V3, Kling V3 Omni, MiniMax H3, Seedance 2.5, Seedance 2.0, and Seedance Mini.
For API fields and Agent parameters, see API Reference and Agent. If you have any questions, please email: hi@weshop.ai.
Usage Principles
Assess the Task Before Choosing a Model
Model selection should follow this order:
- Determine whether the task is pure text-to-image, reference-image generation, local editing, or multi-image consistency. Midjourney and Z-Image support text-to-image only; exclude both whenever image input or image-editing is required.
- Define the delivery tier: draft, internal review, live campaign, or premium commercial delivery.
- Confirm core constraints: subject consistency, Asian aesthetics, photographic quality, speed, cost, or batch scale.
- Choose the lowest-cost model that satisfies the core constraints.
- On generation failure, optimize inputs first, then upgrade the model; do not rely on repeated blind retries.
Default Recommendations
| Requirement | Primary Model | Fallback Model |
|---|---|---|
| Readable text (including Chinese and other languages) | GPT Image 2 Medium | None; adjust tier only when Medium fails or the user explicitly requests it |
| Multilingual images (including Chinese), image translation, and localization | GPT Image 2 Medium | None; use text-free base images plus post-production layout when necessary |
| Artistic creation, inspiration exploration, and illustration | Midjourney | GPT Image 2 |
| Photorealistic text-to-image or Chinese cultural elements | Z-Image | GPT Image 2 |
| General high-quality generation | GPT Image 2 | Nano 2 |
| Subject or character consistency | GPT Image 2 | Nano Pro |
| Low-cost batch drafts | GPT Image 2 (Low) | Seedream 5.0 Lite (only when lighting requirements are high or Asian aesthetics apply) |
| Higher-quality fast iteration | Nano Pro | GPT Image 2 |
| Commercial photography with high lighting requirements | Seedream 5 Pro | GPT Image 2 |
| Asian subjects, fashion, and e-commerce photography | Seedream 5 Pro | Seedream 5.0 Lite |
| Routine commercial images with high lighting requirements or Asian aesthetics | Seedream 5.0 Lite | GPT Image 2 |
Relative Price Order
Under the same pricing basis, this guide uses the following relative price order (high to low):
GPT Image 2 High
> Nano Pro
> Nano 2 = Seedream 5 Pro
> GPT Image 2 Medium = Seedream 5.0 Lite
> GPT Image 2 Low| Price Tier | Model or Quality Mode |
|---|---|
| P5: Highest | GPT Image 2 High |
| P4 | Nano Pro |
| P3 | Nano 2, Seedream 5 Pro |
| P2 | GPT Image 2 Medium, Seedream 5.0 Lite |
| P1: Lowest | GPT Image 2 Low |
The tiers above reflect relative invocation cost only; they are not equivalent to model quality ranking. Actual total cost is also affected by output dimensions, generation count, retry count, and platform billing rules.
Midjourney and Z-Image share the same per-image price. Their relative pricing against other P1–P5 models is not yet confirmed, so they are not included in the tiers above. Midjourney always outputs 4 images per call; when calculating per-invocation cost, bill for 4 images, not a single image.
Capability Declaration Rules
- This guide describes recommended usage boundaries; it does not represent official hard limits of the underlying models.
- GPT Image 2 supports Low, Medium, and High tiers. Other models and tools do not support the
qualityparameter.
Hard Routing Rules for Text-in-Image Tasks
- Whenever the output image must include titles, poster copy, packaging text, signage, menus, UI text, or other readable characters—including Chinese—route uniformly to GPT Image 2 Medium.
- Multilingual generation (including Chinese-inclusive combinations), image text translation, language replacement, and localized layout also route fixedly to GPT Image 2 Medium.
- Z-Image is not a candidate for Chinese text requirements; consider Z-Image only for pure text-to-image tasks that explicitly require Chinese cultural elements.
- Nano 2, Nano Pro, Seedream 5.0 Lite, and Seedream 5 Pro are not primary or fallback models for text-in-image generation.
- GPT Image 2 tiers follow a progressive rule: Low for text layout validation; Medium as the default for routine text and translation deliverables; High only when Medium fails acceptance or the user explicitly requests it.
- Text content must be given verbatim in the prompt, with explicit language, casing, placement, font style, and typographic hierarchy.
- If text accuracy still does not meet delivery requirements, switch to “text-free base image + layout tool for text”; do not blindly switch among models that lack text input capability.
Hard Limits for Text-to-Image-Only Models
- Midjourney and Z-Image perform text-to-image only; input must consist solely of text prompts and generation parameters.
- When any reference image, product image, portrait, background, mask, first/last frame, or source image for editing is present, do not call Midjourney or Z-Image.
- Model swap, outfit change, virtual try-on, multi-product compositing, subject consistency, inpainting, expand image, background removal, and image translation are not pure text-to-image tasks; handle them with the corresponding image models or tools.
- Midjourney and Z-Image do not support
quality; do not pass this field in requests. - Midjourney always outputs 4 images per call; the invocation layer must not request a different count, and the business layer must receive, bill, and filter all 4 results.
- Unit pricing, relative pricing against other models, and speed for both models are not yet confirmed; Z-Image’s unified quality tier is also pending confirmation. Complete configuration before integration; do not infer values.
Seedream Series Usage Limits
- Consider Seedream 5.0 Lite or Seedream 5 Pro only when the task has high lighting requirements or explicitly requires Asian aesthetics.
- For general generation, subject consistency, routine product shots, and standard commercial images that do not meet either condition above, default to GPT Image 2 or the Nano series.
- Seedream 5.0 Lite is for routine deliverables; Seedream 5 Pro is for the same class of needs with higher lighting, material, or final-delivery requirements.
- Text-in-image rules take precedence: whenever readable text, multilingual output, or translation is required—including Chinese—use GPT Image 2 Medium.
GPT Image 2 Tier Hard Rules
GPT Image 2 must use progressive tier selection. Do not jump to High simply because the task description mentions “multiple images,” “subject consistency,” “commercial image,” or “final delivery.”
| Tier | Usage Conditions |
|---|---|
| Low | Lowest-cost drafts, composition validation, text layout validation |
| Medium | Default tier; routine deliverables, subject consistency, multi-angle/multi-pose, standard commercial assets, multilingual and translation tasks |
| High | Medium output explicitly failed acceptance, or user explicitly requests premium final delivery, print-grade detail, etc. |
Enforced sequence:
- Call Medium first for routine tasks; for batch tasks, generate 1 test image with Medium first.
- After the test image passes identity, outfit, composition, text, or product-structure acceptance, complete remaining outputs with Medium.
- Upgrade to High only when Medium results are actually substandard or the user explicitly requests High.
- High cannot compensate for missing reference assets. If outfit, subject angle, or product structure is not visible in references, add assets first.
- When upgrading to High, record the specific reason—for example, “Medium subject detail failed acceptance”—not merely “high quality required.”
Image Models
GPT Image 2
Basic Information
| Attribute | Details |
|---|---|
| Positioning | General-purpose flagship model |
| Core strengths | Versatility, complex instruction following, subject and character consistency |
| Default quality | medium |
| Output quality positioning | High: high; Medium: medium-high; Low: basic |
| Recommended delivery tier | Internal review, live campaigns, commercial delivery |
Recommended Use Cases
| Scenario | Rationale |
|---|---|
| Subject consistency tasks | Same subject with pose, outfit, or scene changes |
| Complex scene compositing | Multiple subjects, multiple constraints, explicit spatial relationships |
| Product showcase | Balance product, environment, and marketing composition |
| Text-in-image generation | Poster titles, packaging text, signage, and other readable characters |
| Multilingual and image translation | Translate or replace image text while preserving layout and visual hierarchy |
| Photorealistic content | Natural, credible overall imagery |
| General design needs | When requirements are not yet biased toward a specific aesthetic |
| Multi-round editing | Iterating on the image from feedback |
Tier and Alternative Boundaries
| Scenario | Assessment | Recommended Choice |
|---|---|---|
| Draft seeking lowest per-call price | Use this model’s Low tier directly | GPT Image 2 Low |
| Top-tier Asian fashion photography | Specialized aesthetics and lighting can be further improved | Seedream 5 Pro |
| Quick composition validation only | Medium or High not needed | GPT Image 2 Low |
For multilingual tasks, call and accept each target language separately; do not output multiple language versions in one request.
Nano 2
Basic Information
| Attribute | Details |
|---|---|
| Positioning | High-throughput fast generation model |
| Core strengths | Speed, batch exploration; priced below Nano Pro, above GPT Image 2 Medium/Low |
quality parameter | Not supported; do not pass |
| Output quality positioning | Medium |
| Recommended delivery tier | Drafts, proof of concept, initial A/B test variants |
Recommended Use Cases
- Batch generation of visual directions and composition candidates.
- Early A/B testing for ad creative.
- Social content sketches, storyboards, and mood boards.
- Tasks needing fast batch exploration with quality above the lowest-cost draft tier.
- Filter prompts, then redo on a higher-quality model.
Not Recommended as Primary
| Scenario | Reason | Recommended Alternative |
|---|---|---|
| Stable subject identity required | Higher consistency risk | GPT Image 2 |
| Fine product texture | Small details easily lost | GPT Image 2; Seedream series when lighting requirements are high or Asian aesthetics apply |
| Print or premium client delivery | Output detail may be insufficient | GPT Image 2; Seedream 5 Pro when lighting requirements are high or Asian aesthetics apply |
| Complex multi-subject scenes | Instruction load too high | GPT Image 2 |
| Readable text in image | Nano 2 does not handle text tasks | GPT Image 2 Medium |
Nano Pro
Basic Information
| Attribute | Details |
|---|---|
| Positioning | Quality-enhanced fast generation model |
| Core strengths | Better detail and instruction execution than Nano 2 while keeping fast iteration; priced just below GPT Image 2 High |
quality parameter | Not supported; do not pass |
| Output quality positioning | Highest |
| Recommended delivery tier | Advanced drafts, internal review, lightweight live assets |
Recommended Use Cases
| Scenario | Notes |
|---|---|
| High-quality concept images | When drafts should not be too rough |
| Medium-scale A/B testing | Quality matters more than lowest cost |
| Social media assets | Simple scenes, short delivery cycles |
| Design review | Clearer material and lighting expression needed |
Note
Readable text in images: do not call Nano Pro; route uniformly to GPT Image 2 Medium.
Boundary with Nano 2
| Criterion | Nano 2 | Nano Pro |
|---|---|---|
| Goal | Explore as many directions as possible | Mature frames from fewer candidates |
| Priority | Cost and speed | Balance of quality and speed |
| Recommended output count | 4–8 directions | 2–4 directions |
| Best stage | Divergence | Convergence and review |
Seedream 5 Pro
Basic Information
| Attribute | Details |
|---|---|
| Positioning | Specialized model for demanding lighting and Asian aesthetics |
| Core strengths | Fine lighting control, material rendering, and Asian commercial aesthetics; same price tier as Nano 2 |
quality parameter | Not supported; do not pass |
| Output quality positioning | Medium-low |
| Recommended delivery tier | Premium advertising, brand key visuals, print-grade delivery |
Recommended Use Cases
- Commercial photography with high requirements for light direction, tonal hierarchy, shadow transitions, rim light, or ambient light.
- Automotive, jewelry, watch, fragrance, and luxury ads requiring studio lighting and texture sculpting.
- Product shots where glass, metal, leather, and similar materials need fine lighting expression.
- Asian subjects, Asian fashion, beauty, and localized e-commerce visuals.
- Brand key visuals emphasizing both Asian aesthetics and professional studio lighting.
Not Recommended as Primary
| Scenario | Reason | Recommended Alternative |
|---|---|---|
| Fast large-scale divergence | Positioned for final deliverables; throughput may not suit exploration | Nano 2 |
| Routine internal review with high lighting or Asian aesthetics | Pro not required directly | Seedream 5.0 Lite / Nano Pro |
| Complex multi-round instruction editing | General editing flexibility may not be optimal | GPT Image 2 |
| Strong abstract or experimental illustration | Specialized photography strengths underused | GPT Image 2 |
| Readable text in image | Seedream 5 Pro does not handle text tasks | GPT Image 2 Medium |
| General tasks without high lighting or Asian aesthetic requirements | Outside Seedream series specialization | GPT Image 2 / Nano series |
Seedream 5.0 Lite
Basic Information
| Attribute | Details |
|---|---|
| Positioning | Routine production tier for demanding lighting and Asian aesthetics |
| Core strengths | Lighting control, Asian aesthetics, and commercial usability; same price tier as GPT Image 2 Medium |
quality parameter | Not supported; do not pass |
| Output quality positioning | Low |
| Recommended delivery tier | E-commerce, social ads, and routine brand assets with high lighting requirements or Asian aesthetics |
Recommended Use Cases
| Scenario | Notes |
|---|---|
| E-commerce product shots with high lighting requirements | Precise control of light source, shadow, depth, and product texture |
| Asian fashion photography | Localized subjects and aesthetic expression |
| Asian subject content | Prioritize localized aesthetics |
| Ad assets with high lighting requirements | Routine batch production and commercial deliverables |
| Pro-tier preview | Confirm lighting, composition, and Asian aesthetic direction first |
Note
Readable text in images: do not call Seedream 5.0 Lite; route uniformly to GPT Image 2 Medium.
Note
If the task does not have high lighting requirements and does not emphasize Asian aesthetics, do not call Seedream 5.0 Lite; use GPT Image 2 or the Nano series instead.
Boundary with Seedream 5 Pro
| Criterion | Seedream 5.0 Lite | Seedream 5 Pro |
|---|---|---|
| Primary goal | Daily commercial production | Top-tier commercial delivery |
| Generation stage | Direction confirmation and routine deliverables | Final key visuals |
| Recommended batch size | Medium | Small curated set |
| Budget | Medium-high | Highest tier |
| Detail requirements | Product pages, social ads | Large-format print, luxury, fine materials |
Midjourney
Basic Information
| Attribute | Details |
|---|---|
| Task type | Text-to-image only |
| Positioning | Top-tier artistic creation model |
| Core strengths | Best for artistic creation, inspiration, illustration, and style exploration |
quality parameter | Not supported; do not pass |
| Output quality positioning | High (not a request parameter) |
| Output count | Always 4 images per call; cannot be changed |
| Per-image price | Same as Z-Image; relative pricing against other models pending confirmation |
Recommended Use Cases
| Scenario | Notes |
|---|---|
| Artistic creation | Emphasis on visual expression, mood, form, and style language |
| Inspiration search | Early concept exploration, mood boards, direction divergence |
| Illustration | Clear artistic-style illustration work |
| Stylized visuals | Surreal, experimental, or non-photorealistic expression |
Prohibited Scenarios
| Scenario | Reason | Alternative |
|---|---|---|
| Any image input or reference-image generation | Text-to-image only | GPT Image 2, Nano, or Seedream series per reference-image task |
| Outfit change, model swap, product compositing | Reference-image editing or multi-image compositing | Virtual Try-On or reference-capable image models |
| Image translation and multilingual localization | Requires reading and modifying source image | GPT Image 2 |
| Any readable text (including Chinese) | Fixed text-task routing | GPT Image 2 Medium |
| Photorealism as core goal without artistic intent | Not this model’s primary route | Z-Image or GPT Image 2 |
Midjourney always generates 4 images. Even when the business needs only 1, receive all results and filter afterward.
Z-Image
Basic Information
| Attribute | Details |
|---|---|
| Task type | Text-to-image only |
| Positioning | Photorealistic text-to-image with Chinese cultural elements |
| Core strengths | Strong photorealism; better understanding and rendering of Chinese cultural symbols |
| Artistic capability | Some artistic expression, but limited style diversity |
quality parameter | Not supported; do not pass |
| Output quality positioning | Pending confirmation |
| Per-image price | Same as Midjourney; relative pricing against other models pending confirmation |
Recommended Use Cases
| Scenario | Notes |
|---|---|
| Photorealistic text-to-image | Subjects, products, or scenes without reference images |
| Chinese cultural elements | Better understanding of traditional architecture, festivals, attire, patterns, and cultural symbols |
Prohibited Scenarios
| Scenario | Reason | Alternative |
|---|---|---|
| Any image input or reference-image generation | Text-to-image only | GPT Image 2 or other reference-capable models |
| Image text translation, language replacement, or localization | Requires reading and editing source image | GPT Image 2 |
| Any readable text, including Chinese and multilingual | Fixed text-task routing | GPT Image 2 Medium |
| Rich artistic styles, top-tier art creation, or broad inspiration exploration | Limited style diversity despite some artistic ability | Midjourney |
| Subject consistency, outfit change, multi-image compositing | Requires reference-image capability | GPT Image 2, Nano series, or Virtual Try-On |
Video Generation Models
Video model pricing, quality tiers, supported duration, resolution, and concurrency limits are not yet confirmed. The following defines model boundaries and invocation structure only; concrete enums must be overridden by actual platform configuration.
Kling V3 / V3 Omni
Basic Information
| Attribute | Kling V3 | Kling V3 Omni |
|---|---|---|
| Model | Kling V3 | Kling V3 Omni |
| Positioning | Controllable video generation | Multi-reference, multimodal video generation |
| Core inputs | Text, first frame, last frame | Text, multiple images, and other reference inputs supported by the platform |
| Primary scenarios | First/last frame, single-image animation, controllable camera motion | Multi-image consistency, complex references, video editing |
| Price / quality tier | Pending confirmation | Pending confirmation |
Kling V3 Recommended Use Cases
| Scenario | Notes |
|---|---|
| First/last frame video | Explicit control of start and end states |
| Single-image animation | Turn product, portrait, or scene stills into motion |
| Controllable camera | Clear requirements for subject motion and camera motion |
| Product showcase | Rotation, push-in, pull-back, detail reveal |
| Medium-complexity motion | Balance controllability and motion amplitude |
Kling V3 Omni Recommended Use Cases
| Scenario | Notes |
|---|---|
| Multi-image reference | Reference character, scene, product, or style images simultaneously |
| Complex subject consistency | Constrain subject or product with multiple reference assets |
| Video reference | Extract motion, camera, or visual traits; capability per platform |
| Video-to-video | Continue generation or editing from existing video; capability per platform |
| Multimodal tasks | Input combinations beyond standard first/last frame |
Version Boundaries
| Criterion | Kling V3 | Kling V3 Omni |
|---|---|---|
| Single first frame | Recommended | Usable but may be overkill |
| Explicit last frame | Recommended | Per platform support |
| Multi-image reference | Do not use | Recommended |
| Video reference | Do not use | Recommended |
| Simple controllable camera | Recommended | Usable |
| Complex reference combinations | Insufficient capability | Recommended |
Not Recommended as Primary
| Scenario | Reason | Recommended Alternative |
|---|---|---|
| Dance, sports, and large-amplitude motion | Dynamic range not primary strength | MiniMax H3 |
| Audio-driven high-dynamic performance | Needs motion and rhythm capability | MiniMax H3 |
| Low-cost concept validation only | May be overkill | Seedance Mini; confirm pricing first |
| Long continuous narrative | Single-shot duration limited | Split shots and composite |
MiniMax H3
Basic Information
| Attribute | Details |
|---|---|
| Positioning | High-dynamic motion video model |
| Core strengths | Large-amplitude motion, motion continuity, dynamic camera |
| Price / quality tier | Pending confirmation |
Recommended Use Cases
| Scenario | Notes |
|---|---|
| Large-amplitude motion | Dance, sports, running, jumping, action performance |
| High-dynamic camera | Fast movement, pronounced pose change, dynamic camera work |
| Motion continuity | Emphasis on action trajectory and rhythmic flow |
| Audio-driven video | Audio combined with subject or video reference; support per platform |
| Complex pose changes | Multiple consecutive action phases in one shot |
Not Recommended as Primary
| Scenario | Reason | Recommended Alternative |
|---|---|---|
| Precise first/last frame control | Core need is start/end state, not dynamic range | Kling V3 |
| Multi-image subject consistency | Needs more reference constraints | Kling V3 Omni |
| Static or subtle motion | High-dynamic strengths unused | Kling V3 |
| Primarily artistic audiovisual expression | Aesthetics and AV expression prioritized | Seedance 2.5 |
Required visual input combinations, duration, and format for audio-driven generation must follow the actual platform; do not assume pure-audio requests always work.
Seedance 2.5 / 2.0 / Mini
Basic Information
| Attribute | Seedance 2.5 | Seedance 2.0 | Seedance Mini |
|---|---|---|---|
| Model | Seedance 2.5 | Seedance 2.0 | Seedance Mini |
| Positioning | Audiovisual expression and artistic deliverables | Standard creative video | Lightweight concept validation |
| Primary scenarios | Audio-visual sync, music video, artistic short | Routine artistic expression, creative shorts | Drafts, lightweight trials, fast exploration |
| Price / quality tier | Pending confirmation | Pending confirmation | Pending confirmation |
Version Selection
| Requirement | Preferred Version | Notes |
|---|---|---|
| Audio-visual sync or music video | Seedance 2.5 | Audio and visual rhythm as core |
| Artistic final deliverable | Seedance 2.5 | Creative and aesthetic expression prioritized |
| Routine artistic short | Seedance 2.0 | Does not need 2.5 enhancements |
| Concept validation and drafts | Seedance Mini | Lightweight exploration; confirm cost on platform |
Recommended Use Cases
- Music videos, mood shorts, and artistic social content.
- Non-photorealistic, emotional, or stylized visual expression.
- Shorts centered on rhythm, color, transitions, and visual mood.
- Seedance Mini for validating style, rhythm, and camera direction.
Not Recommended as Primary
| Scenario | Reason | Recommended Alternative |
|---|---|---|
| Precise product structure showcase | Controllability over artistic expression | Kling V3 |
| Multi-image character consistency | Needs stronger reference constraints | Kling V3 Omni |
| Dance, sports, large-amplitude motion | Dynamic range prioritized | MiniMax H3 |
| Precise first/last frame | Needs explicit start/end states | Kling V3 |
| Long complex narrative | Single shot cannot guarantee continuity | Storyboard generation then composite |
Video Model Quick Decision
Does the task require large-amplitude motion or high-dynamic action?
├─ Yes → MiniMax H3
└─ No: Does it require multi-image, video, or complex references?
├─ Yes → Kling V3 Omni
└─ No: Does it require precise first/last frame or controllable product showcase?
├─ Yes → Kling V3
└─ No: Is audio-visual sync or artistic expression the core goal?
├─ Yes → Seedance 2.5
└─ No: Is it only lightweight concept validation?
├─ Yes → Seedance Mini
└─ No → Seedance 2.0Video Routing Priority
- Precise first/last frame and product structure control take priority over artistic style → Kling V3.
- Multi-reference input takes priority over standard first/last frame → Kling V3 Omni.
- Large-amplitude motion takes priority over identity and first/last frame control → MiniMax H3.
- When audio-visual sync and artistic expression are prioritized → Seedance 2.5.
- Seedance 2.0 for routine artistic shorts; Mini for lightweight validation.
- When conflicting goals coexist, split shots or have the user confirm the highest priority first.
Image Processing Tools
Use processing tools when the task is not image or video generation. For API parameters, see Agent.
Tool Overview
| Tool | Choose When | Choose Image Model Instead When |
|---|---|---|
| Remove BG | Need transparent PNG cutout only | Need new background, content edit, or generative fill |
| Expand Image | Need aspect ratio or platform size change via crop | Need upscaling, outpainting, or content modification |
| Virtual Try-On | Single product, at most 1 single-person model, at most 1 background | Multiple products, multiple people, multi-product outfit, or readable text in output |
Virtual Try-On Routing Rules
Do not call Virtual Try-On when:
| Scenario | Alternative |
|---|---|
| Two or more products at once | Image model workflow with multi-reference compositing |
| Multiple model images or multiple people in model image | Image models |
| Missing product image | Add product image first |
| Readable text in output | GPT Image 2 Medium |
When Virtual Try-On does not apply, route to image models per this guide: GPT Image 2 for text or multilingual tasks; Seedream series when lighting requirements are high or Asian aesthetics apply; otherwise GPT Image 2 or Nano series by consistency, quality, speed, and price.
Tool Selection Rules
Need to remove the original background?
├─ Yes → Remove BG
└─ No: Is it virtual try-on, model swap, or background swap?
├─ Yes: Is there exactly 1 product image, at most 1 single-person model image, and at most 1 background image?
│ ├─ Yes → Virtual Try-On
│ └─ No → Image model workflow
└─ No: Need to change aspect ratio, platform size, or crop/composition?
├─ Yes → Expand Image
└─ No → Image generation or editing workflowModel Selection & Routing
Image Model Comparison Matrix
“Output quality positioning” in this table uses internal assessment criteria; it is not a request parameter and does not equal fit for a specific style. Only GPT Image 2 supports the
qualityparameter. Seedream enters the candidate set only for tasks with high lighting requirements or Asian aesthetics. Midjourney and Z-Image support text-to-image only; they share the same per-image price, and Midjourney always outputs 4 images per call.
| Model / Tier | Output Quality Positioning (Not a Parameter) | Speed | Price Tier | Consistency | Photographic Quality | Recommended Stage |
|---|---|---|---|---|---|---|
| GPT Image 2 High | High | Medium | P5 (highest) | High | High | Final deliverables, complex high-quality tasks |
| Nano Pro | Highest | Faster | P4 | Medium | Medium-high | Convergence, internal review |
| Nano 2 | Medium | Fast | P3 | Medium-low | Medium | Batch divergence, drafts |
| Seedream 5 Pro | Medium-low | Slower | P3 | Medium-high | Very high | High-requirement delivery with demanding lighting or Asian aesthetics |
| GPT Image 2 Medium | Medium-high | Medium | P2 | High | High | General deliverables, routine complex tasks |
| Seedream 5.0 Lite | Low | Medium | P2 | Medium-high | High | Routine production with demanding lighting or Asian aesthetics |
| GPT Image 2 Low | Basic | Faster | P1 (lowest) | Medium-high | Medium-high | Lowest-cost validation and drafts |
| Midjourney | High; top-tier for artistic creation | Pending confirmation | Same per-image price as Z-Image; tier pending | N/A for reference consistency | Strong and diverse artistic style | Pure text-to-image art, inspiration, illustration; always 4 images |
| Z-Image | Pending confirmation; strong photorealism | Pending confirmation | Same per-image price as Midjourney; tier pending | N/A for reference consistency | Strong photorealism; limited artistic style diversity | Pure text-to-image photorealism and Chinese cultural elements |
Image Model Quick Decision Tree
Does the task include any image input, reference image, or image-editing requirement?
├─ Yes → Exclude Midjourney and Z-Image
│ ├─ Image translation, multilingual, or readable text editing → GPT Image 2 Medium
│ └─ Other tasks → Choose GPT Image 2, Nano, or Seedream series by consistency, cost, lighting, and Asian aesthetics
└─ No (pure text-to-image)
├─ Artistic creation, inspiration, or illustration → Midjourney
├─ Any readable text or multilingual (including Chinese) → GPT Image 2 Medium
├─ Chinese cultural elements → Z-Image
├─ Photorealism as core without artistic emphasis → Z-Image
└─ Other general tasks → GPT Image 2 Medium; then adjust by price, batch, and style needsRecommended Image Production Pipelines
| Workflow | Recommended Pipeline |
|---|---|
| General marketing assets | GPT Image 2 Low validation → Nano 2 batch divergence → GPT Image 2 Medium deliverables |
| E-commerce with demanding lighting or Asian aesthetics | Nano Pro composition → Seedream 5.0 Lite deliverables → Seedream 5 Pro curated hero images |
| Brand key visuals with demanding lighting or Asian aesthetics | Seedream 5.0 Lite preview → Seedream 5 Pro delivery |
| Subject consistency series | GPT Image 2 Medium: generate 1 test first → complete with Medium after acceptance; upgrade to High only if Medium fails |
| Pure text-to-image art concepts and illustration | Midjourney divergence → manual filter and prompt convergence |
| Pure text-to-image Chinese cultural photorealism | Z-Image generation → cultural element acceptance; switch to GPT Image 2 Medium if readable text is also required |
Model Fallback
When the primary model does not meet requirements, use the following fallback chains.
Image Model Fallback Chain
| Original Model or Tier | First Fallback | Final Fallback |
|---|---|---|
| GPT Image 2 High | GPT Image 2 Medium | GPT Image 2 Low |
| Nano Pro | Nano 2 | GPT Image 2 Medium |
| Nano 2 | GPT Image 2 Medium | GPT Image 2 Low |
| Seedream 5 Pro | Seedream 5.0 Lite | GPT Image 2 Medium |
| Seedream 5.0 Lite | GPT Image 2 Medium | GPT Image 2 Low |
| Midjourney | No automatic cross-model fallback | Manual choice |
| Z-Image | GPT Image 2 Medium | Text-free base + post layout |
Readable text, multilingual, translation, and image text editing tasks do not use the cross-model fallback chain above; default to GPT Image 2 Medium.
Video Model Fallback Chain
| Original Model | First Fallback | Final Fallback |
|---|---|---|
| Kling V3 Omni | Kling V3 | Manual handling |
| Kling V3 | Seedance 2.0 | Seedance Mini |
| MiniMax H3 | Kling V3 | Seedance 2.0 |
| Seedance 2.5 | Seedance 2.0 | Seedance Mini |
| Seedance 2.0 | Seedance Mini | Manual handling |
| Seedance Mini | Manual handling | — |
Do not silently downgrade when fallback would break core capabilities such as multi-image reference, precise last frame, or large-amplitude motion.