ERNIE-Image Support in ComfyUI: Precise Text Rendering and Structured Image Generation

CN
2026-04-16 02:09:36

Today we announce that Baidu’s open-source ERNIE-Image framework, an Apache-2.0 licensed text-to-image model, joins the ComfyUI ecosystem. Powered by an 8-billion-parameter Diffusion Transformer architecture, this robust open-weights system creates more detailed outcomes from brief inputs with its integrated Prompt Enhancement capability.

Core Capabilities

  • Text precision: Generates complex English, Chinese, and multilingual layouts

  • Instruction compliance: Processes sophisticated directions, object relationships, and knowledge-rich descriptions

  • Structured visual creation: Generates posters, comics, multi-scene storyboards

  • Artistic flexibility: Produces images from photorealistic to cinematic aesthetics

  • Resource efficiency: Operates within 24GB VRAM at 8B parameters

  • Prompt enhancement: 3B auxiliary model enriches minimal inputs

Visual Demonstrations

Layout and Typography Applications

Subgraph Parameter Panel

Prompt: Educational comic book style infographic showing the coffee making process. The background has a light beige vintage paper texture. At the top center of the image is a bold brown title that reads 'The Coffee Making Process', with a smaller English subtitle 'How Coffee is Made' below. The main part consists of six step blocks connected by brown dotted arrows, arranged in two rows up and down, forming a Z-shaped visual guide.The first step in the upper left corner illustrates a coffee tree full of ripe red coffee cherries, with a cut coffee cherry next to it showing the beans inside, labeled 'Step 1: Harvesting Cherries' below the block.The second step in the upper middle shows wooden fermentation boxes filled with coffee beans, labeled 'Step 2: Pulping and Fermentation'.The third step in the upper right depicts coffee beans being dried under the sun on bamboo mats, labeled 'Step 3: Sun Drying'.The fourth step in the lower left shows a vintage metal roasting machine with coffee beans rolling inside and steam rising, labeled 'Step 4: Roasting'.The fifth step in the lower middle features a stone grinder pouring out smooth coffee grounds, labeled 'Step 5: Grinding'.The sixth step in the lower right shows a production line where coffee liquid is being poured into molds, with finished packaged coffee beside it, labeled 'Step 6: Brewing and Forming'.The four corners of the image are decorated with hand-drawn coffee leaves and coffee beans. The overall color palette consists of warm brown, caramel, cream, deep red and olive green, with delicate lines and clear layout.

Subgraph Parameter Panel

Prompt: A 6-panel comic page, 2 columns × 3 rows, black border between each panel.Panel 1 (top-left): wide shot — a young woman in a red coat stands at a rainy train station. Caption box: "She had waited three years for this moment."Panel 2 (top-right): close-up on her face — anxious eyes, rain on her cheek. No text.Panel 3 (mid-left): medium shot — a train arrives, doors slide open, steam rising. Sound effect text: "WHOOOOSH"Panel 4 (mid-right): her point-of-view — a man in a gray jacket steps out, back to camera.Panel 5 (bottom-left): extreme close-up — her hand trembling as she reaches forward.Panel 6 (bottom-right): wide shot — they face each other under one umbrella. Caption box: "Some arrivals change everything."Ink and watercolor style, cool blue-gray palette, expressive line art. Same character design consistent across all 6 panels.

Cinematic Expressions

Prompt: Split-screen conceptual poster with vertical split composition, left half representing the present and future featuring a solid stone analog clock dial with bright red hands and red scale markings, set against a deeply cracked, dry, parched earth texture in muted cool gray tones, embodying slow, static structural decay and temporal erosion, right half representing the past that is violently disintegrating, dissolving and exploding into chaotic dust, debris, flying stone shards and swirling cosmic nebula energy with a deep blood red and black color palette, the circular clock frame itself is half solid weathered stone and half crumbling into particles merging the two worlds, with the red clock hands fully positioned on the left present/future side of the dial, embodying the concept of time, past vs future, slow decay vs violent collapse, hyper-realistic 3D render with cinematic dramatic lighting, high contrast, sharp details on the left side, motion blur and dynamic particle effects on the right side, photorealistic textures of cracked earth and floating dust, atmospheric haze, futuristic graphic design with minimalist red vertical UI elements and custom technical text overlays in the corners, bold red stylized title text "ETERNAL DAWN" in the top right corner, small red technical metadata text "SPECDARY BOTHUNG" directly below the main title in the top right, dense lines of small red technical data blocks including "TIMESTAMP: 00:00:00", "COLLAPSE RATE: 99.8%", "TEMPORAL ANOMALY DETECTED", "PAST DIMENSION: DECAYING" in the bottom right corner with a red horizontal accent bar and footer text "FOR DECAY OF THE EONS" below, small red header text "TEMPORAL SHIFT PROTOCOL" with vertical red accent line running down the left edge in the top left, vertical red accent line with small red technical text annotations including "PRESENT: STABLE", "FUTURE: UNFOLDING" along it in the bottom left, red text "WARNING: TEMPORAL DISINTEGRATION" curving around the center clock dial perimeter, small red footer text "PAST FADES, FUTURE RISES" spanning the split at the bottom center, shot in the style of a high-end sci-fi movie poster, 8K ultra-high resolution, photorealistic, cinematic composition, dramatic depth, moody and intense atmosphere, detailed particle simulation, photorealistic material rendering.
Prompt: First-person vertical real-world urban night photo. Narrow, wet, busy city street with strong depth; towering buildings on both sides with fire escapes, pipes, AC units, warm window lights. Red vertical “HOTEL” neon on left, bright blue “BAR” sign and yellow “OPEN 24 HOURS” box on right. Distant huge digital billboard glowing cyan “NEON CITY”. Rough wet asphalt reflects neon lights; vintage yellow taxi with “TAXI” sign drives down center, red taillights on. Two pedestrians in dark coats with black umbrellas walk away on right sidewalk. Giant needle-like steel TV tower peeks through building gaps at end, red aircraft warning light glowing. Cinematic volumetric light, misty air, cool cyan-blue vs warm orange-red contrast, immersive and realistic.

Subgraph Parameter Panel
A charming, whimsical 3D illustration featuring a fluffy, plush-textured white duck sitting comfortably in a fuzzy coral-pink armchair, holding a bright red mug of steaming black coffee against a solid warm coral-red background; defined by its tactile, huggable fabric-like texture across all elements, a bold, minimalist warm color palette of creamy white, vibrant orange, and rich reds, gentle anthropomorphism that blends cuteness with relatable cozy relaxation, soft diffused lighting, and a subtle film grain that adds a nostalgic, handcrafted feel, creating a playful yet serene mood perfect for themes of calm downtime and morning routines.

Multi-frame Layouts

Prompt: Educational comic book infographic. Five vertical panels side by side, earthy color palette, hand-drawn illustration style. Top title: "NORTH AMERICAN NATIVE SPECIES". Style: comic line art, watercolor fills, white panel backgrounds, bold headers.Panel 1: Gray squirrel holding an acorn. Header: "EASTERN GRAY SQUIRREL" Fact: "Buries acorns to help forests grow." Arrow → tail: "Bushy tail"Panel 2: Robin with orange-red breast on a branch. Header: "AMERICAN ROBIN" Fact: "A sign of spring returning." Arrow → breast: "Orange-red breast"Panel 3: Red maple tree with autumn leaves. Header: "RED MAPLE TREE" Fact: "Shelters wildlife year-round." Arrow → leaves: "Red in autumn"Panel 4: Deer with white tail raised. Header: "WHITE-TAILED DEER" Fact: "White tail signals danger." Arrow → tail: "Alarm signal"Panel 5: Monarch butterfly on a flower. Header: "MONARCH BUTTERFLY" Fact: "Migrates 3,000 miles each year." Arrow → wings: "Warning colors"

Implementation Guide

  1. Update to ComfyUI v0.19.1 or higher

  2. Access 'Templates' and locate ERNIE-Image

  3. Select and activate the ERNIE-Image workflow

  4. Download required assets, configure prompts, and execute

Accessible Models

🤗 ERNIE-Image - Primary model delivering quality outputs in ~50 steps
🤗 ERNIE-Image-Turbo - Accelerated iteration generating in 8 steps

Begin your creative journey!