jd-opensource/JoyAI-Image

JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.

What it solves

JoyAI-Image addresses the gap between image understanding and image generation. It provides a unified framework that allows a model to not only understand the spatial layout and content of an image but also generate new images or edit existing ones based on precise instructions, ensuring that changes preserve the scene structure and visual consistency.

How it works

The project combines an 8B Multimodal Large Language Model (MLLM) for understanding and a 16B Multimodal Diffusion Transformer (MMDiT) for generation. This creates a "closed-loop collaboration" where spatial understanding informs grounded generation and editing, while the generative capabilities (like changing viewpoints) provide new visual evidence to improve the model's spatial reasoning.

Who it’s for

This tool is designed for researchers and developers working in multimodal AI, specifically those needing high-fidelity image editing, precise spatial manipulation, and advanced text rendering within images.

Highlights

  • Unified Foundation: A single model family that handles understanding, text-to-image generation, and instruction-guided editing.
  • Spatial Intelligence: Supports precise object movement, object rotation to specific views, and camera control (yaw, pitch, and zoom).
  • Advanced Text Rendering: High performance in rendering long-form text, multilingual typography, and complex layouts like comics.
  • Multi-Image Editing: The "Plus" version supports cross-image composition and joint manipulation across multiple images.
  • Cofiguration Support: Native integration with Diffusers and ComfyUI.

Related

  • Project
  • Project
  • Project
  • Project
  • Project