Learn how to run Ideogram 4 locally with SwarmUI and ComfyUI, download the required model bundle, use the ready Turbo, Balanced, and Highest Quality presets, and create accurate structured JSON prompts with Ultimate Image Captioner Pro.
The workflow covers image recreation, reliable text rendering, bounding-box editing, batch captioning, and training-dataset preparation.

Full Tutorial
Watch the complete step-by-step tutorial on YouTube:
Ideogram 4: The Ultimate JSON Prompting Masterclass
Tutorial Resources
Supported Models
Ultimate Image Captioner Pro supports the following models with robust torch.compile integration.
Qwen Vision Models
- Qwen3-VL 8B Instruct (default)
- Huihui Qwen3-VL 8B Instruct Abliterated
- Qwen3-VL 4B Instruct
- Qwen3-VL 2B Instruct
- Qwen3-VL 30B-A3B Instruct
- Qwen3.6 27B
- Huihui Qwen3.6 27B Abliterated
Joy Caption Models
- Joy Caption Beta 1
- Joy Caption Alpha 2
- Joy Caption Alpha 1
- Joy Caption Pre Alpha
Torch 2.13 Runtime
The application uses Torch 2.13 with the latest project-tested, precompiled supporting libraries.

Application Preview
Click the image to open the full-size screenshot.

The fully compiled captioning path delivers an 84% speed improvement in the demonstrated benchmark.

Installers
Installer workflows are available for Windows, RunPod, SimplePod, Massed Compute, and local Linux systems.

Video Chapters
Show all tutorial chapters
00:00:00 - Ideogram 4 overview: JSON prompting, SwarmUI presets, ComfyUI workflows, and model bundle
00:00:53 - Ultimate Image Captioner Pro for turning reference images into Ideogram JSON prompts
00:01:10 - Editing JSON elements, bounding boxes, wanted text fields, captions, and prompt layout
00:02:02 - Regeneration examples showing structure, objects, scene layout, and image text matching
00:03:18 - Captioner Pro feature tour: Qwen, JoyCaption, saved outputs, and JSON builder
00:04:30 - Dataset workflow: prompt presets, batch folder captioning, and automatic VRAM presets
00:05:13 - Tutorial roadmap: ComfyUI update, SwarmUI update, model download, app installation, and usage
00:05:40 - Updating ComfyUI by extracting the latest installer ZIP and overwriting old files
00:05:56 - Optional fresh ComfyUI virtual environment rebuild for outdated or broken installations
00:06:15 - Running the ComfyUI update script, Python choice, UV speed, and quantization support
00:07:05 - Installing recommended custom nodes bundle 100 for ComfyUI and SwarmUI compatibility
00:07:50 - Launching fresh ComfyUI and testing the Ideogram Turbo preset workflow
00:08:47 - Setting width, height, resolution, and matching the prompt aspect ratio
00:09:07 - Updating SwarmUI with the latest ZIP, overwrite method, and safe folder paths
00:09:48 - Automatic .NET SDK 10 installation and why SwarmUI needs the correct SDK version
00:10:51 - SwarmUI backend setup: ComfyUI backend, Triton, Sage Attention cautions, and extra arguments
00:11:44 - Downloading the Ideogram 4 core bundle with hash verification
00:12:28 - 16-connection parallel downloads, target folders, ComfyUI mode, and URL downloader
00:13:20 - Merging model parts and sharing SwarmUI models through extra_model_paths.yaml
00:13:51 - Setting the SwarmUI model root to reuse another model folder and avoid duplicates
00:14:12 - Updating SwarmUI presets with delete import, normal import, overwrite, and backup
00:14:58 - Refreshing presets and confirming Ideogram Turbo, Balanced, and Highest Quality
00:15:14 - First simple Ideogram prompt, false safety-filter block, and weak plain prompting
00:15:34 - Using Realism Engine Ideogram 5 LoRA to fix the blocked car prompt
00:15:57 - Why detailed JSON prompts are needed and downloading Captioner Pro
00:16:23 - Installing Captioner Pro with the Windows install/update app, virtual environment, and model downloads
00:16:34 - Windows requirements: Python, CUDA, cuDNN, C++ tools, FFmpeg, Git, and setup guide
00:17:03 - Cloud and Linux notes plus the Massed Compute interface, creator image, GPU, and coupon
00:17:34 - Captioner installer downloader: 16 connections, hash checks, and accurate setup
00:17:57 - Starting Ultimate Image Captioner Pro and saving custom user presets
00:18:14 - Loading the Bugatti reference image and generating official Ideogram JSON
00:18:39 - Prompt generation speed, copying the prompt, and understanding VRAM usage
00:19:09 - Subprocess mode to release all VRAM and RAM after each captioning run
00:19:54 - Reviewing generated JSON: high-level description, visible text, boxes, and details
00:20:21 - Pasting JSON into SwarmUI and matching the custom 5:3 aspect ratio
00:20:43 - Aspect-ratio calculator, side-length control, and high-resolution generation
00:21:36 - Comparing results with and without aspect-ratio metadata and avoiding false safety blocks
00:21:58 - Realism Engine LoRA strength, when to use it, and output comparison
00:22:34 - Choosing Turbo, Balanced, or Highest Quality and testing Turbo speed
00:22:54 - Ideogram 4 image-to-image, inpainting, image creativity, and image prompts
00:23:19 - Captioner Pro batch-folder processing: subfolders, overwrite, and append modes
00:23:35 - Post-processing captions with prefixes, suffixes, replacements, and sensitivity
00:24:07 - Final options, automatic quantization by GPU VRAM, support channels, and closing
Covered in the Tutorial
- Local Ideogram 4 installation
- SwarmUI and ComfyUI preset usage
- Automatic model downloads and hash verification
- Structured JSON prompt creation
- Bounding-box and visible-text editing
- Reference-image recreation
- Safety-filter troubleshooting and LoRA realism settings
- Folder-based batch captioning
- VRAM-friendly caption generation