Qwen Image is an image generation and editing model from Alibaba’s Qwen team. Its original open-weight release combines visual generation, text rendering, and instruction-based editing in one model, with particular strength in Chinese and English typography.
In Phygital+, the Qwen Image node can start from a text prompt or use up to three images as references. Those references can define the subject, product, style, layout, or environment. The node also exposes LoRA support, multiple aspect ratios, seed and guidance controls, and camera presets.
This guide focuses on the Qwen Image implementation available in Phygital+. Newer Qwen Image 2.0 and 3.0 families now exist, so version-specific functions are separated instead of being presented as one feature set.
Qwen Image versions compared
| Version | Main strength | Best use |
|---|---|---|
| Qwen Image (2025) | 20B open-weight image generation, multilingual text, and editing | Phygital+ generation, reference editing, LoRA, and camera-control workflows |
| Qwen Image 2.0 | Unified generation and editing with a smaller, newer architecture | Current hosted workflows where this model ID is explicitly available |
| Qwen Image 3.0 | Complex layouts, fine text, reasoning, 1–3 references, and up to 2048×2048 | Newest Alibaba Cloud Qwen Image API workflows |
| Qwen Image 3.0 Pro | Higher-fidelity layout, typography, and photographic detail | Complex design work when the Pro endpoint is available |
What Qwen Image is good at
Text-to-image generation
The model handles detailed descriptions of subject, composition, style, lighting, and typography. It is useful for posters, illustrations, product concepts, editorial images, and structured visual assets.
Instruction-based editing
Connect an existing image and describe what should change. The clearest edit prompts specify the target region, the new content, and the elements that must remain untouched.
Multi-image composition
Up to three source images can contribute different parts of a result. One can define a person, another a product, and a third a location or style. The prompt should assign a clear job to every image.
Multilingual text
The original model was designed around complex text rendering, especially Chinese and English. Other scripts can work, but current official editing documentation guarantees Simplified Chinese and English, so every localized result requires proofreading.
Camera and LoRA control
The Phygital+ node exposes camera angle and distance controls as well as trained and preset LoRAs. These controls make it possible to explore composition and visual style without rewriting the entire prompt.
How to work with multiple references
Reference order has no universal meaning by itself. Define each role in the prompt:
- Image 1: preserve the person’s face and hairstyle.
- Image 2: use the jacket and accessories.
- Image 3: place the subject in this interior and match its lighting.
When an edit fails, remove references that describe the same element differently. A concise set of compatible inputs is easier for the model to reconcile.
Generation versus editing
Without a start image, the node generates from the prompt. With references, the task becomes editing or composition. In edit mode, describe the transformation and preservation rules instead of repeating every visible detail.
How to generate text inside images
- Put exact copy in quotation marks.
- Identify the language separately from the prompt language.
- Describe text position, size, and hierarchy.
- Limit the first generation to the essential headline and one supporting line.
- Request no additional words.
- Proofread the final result at full size.
Qwen Image can build posters and layouts around longer copy than many diffusion models, but small print, mixed scripts, unusual names, and dense paragraphs still require manual review.
How to prompt Qwen Image
Write the prompt as a production brief: output type, subject, composition, exact text, style, lighting, reference roles, and preservation constraints.
| Layer | What to describe | Example |
|---|---|---|
| Output | Poster, product image, edit, illustration | Vertical skincare advertisement |
| Subject | Main person or object | Amber serum bottle on pale stone |
| Composition | Placement, crop, and negative space | Bottle centered, copy in upper-left third |
| Text | Exact wording and language | “RESET YOUR GLOW” in English |
| Style | Medium, palette, texture, and lighting | Warm editorial photography, soft daylight |
| References | Role of each input | Image 1 product, Image 2 lighting |
| Constraints | Details that must stay fixed | Preserve cap, label, glass color, and proportions |
Qwen Image prompt examples
Bilingual poster
Create a vertical technology poster with a translucent blue sphere in a dark gallery. Headline in English: “NEW SIGNALS”. Supporting line in Simplified Chinese: “未来视觉实验”. Large white geometric typography, clean hierarchy, no extra words.
Skincare campaign
Create a premium skincare advertisement featuring the amber bottle from Image 1 on pale limestone. Use the warm window light from Image 2. Include “RESET YOUR GLOW” in refined black serif type. Preserve bottle shape, label, cap, and liquid color.
Product composite
Use the headphones from Image 1, the model from Image 2, and the concrete studio from Image 3. Create a horizontal campaign image with the model wearing the exact headphones. Match reflections and shadows. Preserve face, product geometry, and logo.
Packaging edit
Edit Image 1. Replace only the front label with a matte cream label reading “NORTH COAST” and “SEA SALT”. Keep the bottle, cap, glass, lighting, reflections, background, and camera angle unchanged.
Character outfit
Preserve the face and hairstyle from Image 1. Dress the character in the red coat from Image 2 and place them in the snowy street from Image 3. Full-body editorial photograph, natural winter light, consistent anatomy.
Infographic
Create a clean vertical infographic titled “A DAY OF ENERGY”. Show four labeled sections: “BREAKFAST”, “WORK”, “TRAINING”, and “REST”. Use simple icons, dark green text, cream background, clear reading order, no extra labels.
Restaurant menu
Design a one-page menu titled “AUTUMN TABLE”. Include three headings: “STARTERS”, “MAINS”, and “DESSERTS”. Warm white paper, deep burgundy serif typography, small line illustrations, generous spacing, readable text.
Camera variation
Create a three-quarter low-angle product photograph of the same silver sneaker from Image 1. Preserve all materials, stitching, sole shape, and logo. Dark studio, narrow blue rim light, subtle floor reflection.
Interior redesign
Edit the living room in Image 1. Replace the sofa with the green modular sofa from Image 2 and apply the warm wood palette from Image 3. Keep the architecture, windows, camera position, and daylight unchanged.
Book cover
Create a science-fiction book cover showing a small observatory under a huge pale moon. Title: “THE LAST TRANSMISSION”. Author: “ELI VOSS”. Minimal blue and ivory palette, sharp serif type, no additional text.
Food poster
Create a square food poster featuring a bowl of spicy noodles photographed from above. Add “MIDNIGHT NOODLES” in bold yellow type and “OPEN UNTIL 2 AM” below. Red table, strong flash, energetic editorial layout.
Illustration series
Using the visual style from Images 1 and 2, illustrate a cyclist crossing a bridge at dawn. Preserve the flat shapes, grain, limited coral-and-navy palette, and simplified perspective. Leave clean space for a headline.
Background replacement
Edit Image 1. Keep the person, pose, clothing, hair, and foreground shadows. Replace only the background with a modern train platform at blue hour. Match the new light direction and depth of field.
Logo concept
Create a clean wordmark for “FIELDNOTE”, a travel-journal brand. Rounded lowercase lettering with one simple compass detail. Black on white, flat vector-style presentation, no mockup, no slogan, no extra symbols.
Social carousel cover
Create a square social carousel cover titled “5 WAYS TO FIX A WEAK BRIEF”. Large black sans serif headline, one oversized red number 5, white background, editorial grid, strong mobile readability, no additional copy.
Building a Qwen Image workflow in Phygital+
Qwen Image can sit in the middle of a broader design pipeline. A text model can structure the brief and exact copy. Image nodes can create the subject, product, or environment references. Qwen combines or edits them, while branches test camera, layout, LoRA, and guidance settings. The approved output can continue into upscaling or video.
- Define the deliverable, copy, and preservation rules.
- Create or collect up to three references with distinct roles.
- Generate or edit in Qwen Image.
- Branch camera angles, compositions, or LoRA styles one variable at a time.
- Check text, identity, product details, and shadows.
- Upscale or animate the selected image.
The canvas preserves the relationship between every source and variation, which makes client revisions easier to reproduce.
Qwen Image pricing in Phygital+
Qwen Image uses Phygital credits. The current cost appears inside the node before generation and can depend on the active options and image count. The same subscription balance can be used for the surrounding text, image, video, and utility nodes.
Common Qwen Image problems
The edit changes too much
Name the exact target and list every important element that must remain unchanged.
References conflict
Assign one role to each image and remove duplicates that disagree about identity, product shape, or style.
Text is misspelled
Shorten the copy, use quotation marks, request no extra words, and proofread at full resolution.
The composite looks pasted together
Ask the model to match perspective, light direction, shadows, color temperature, and depth of field.
The character loses identity
Use the clearest face as the primary reference and avoid simultaneous changes to face, hair, pose, camera, and style.
Camera controls distort the subject
Make smaller angle changes and keep the subject or product constraints explicit.
Current limitations
Dense text, mixed-language layouts, tiny lettering, exact logos, complex hands, and aggressive viewpoint changes can still fail. LoRAs and references can also pull the image in competing directions. Treat generated trademarks, packaging copy, and regulated claims as drafts until they are reviewed.
FAQ
What is Qwen Image?
Qwen Image is Alibaba’s image generation and editing family. The original 2025 release is a 20B open-weight model focused on text rendering and instruction-based editing.
Which Qwen Image version is available in Phygital+?
The current model page documents the original Qwen Image implementation. Use the controls shown in the node and do not assume Qwen Image 3.0 API features are present.
How many reference images can I use?
The Phygital+ Qwen Image node supports up to three reference images.
Can Qwen Image edit an existing image?
Yes. Connect one or more images and describe the target change and the details that must remain fixed.
Is Qwen Image good at text?
Yes, especially for Chinese and English. Other scripts may work, but every final asset should be proofread.
Can I use a LoRA?
Yes. The Phygital+ node supports trained LoRAs and preset modifiers with adjustable strength.
What resolutions are available?
The current Phygital+ model page lists dimensions from 512 to 1664 pixels per side with multiple aspect-ratio presets.
What is the difference between Qwen Image and Qwen Image 3.0?
Qwen Image 3.0 is a newer hosted family with generation and editing, reasoning controls, up to three references, and output up to 2048×2048. The original model is a separate open-weight release.
How much does Qwen Image cost in Phygital+?
The model uses Phygital credits, and the current cost is displayed inside the node before generation.
Why use Qwen Image in Phygital+?
Phygital+ connects briefing, references, Qwen generation and editing, LoRA experiments, branching, upscaling, and video on one canvas.