viiika/Meissonic
[ICLR 2025] Official Implementation of Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
What it solves
Meissonic 은 고해상도 이미지를 효율적으로 생성하는 문제를 해결합니다. 소비자용 그래픽 카드(GPU)에서도 실행될 수 있을 만큼 가볍게 유지하면서 최대 1024x1024 픽셀까지 이미지를 합성할 수 있는 방법을 제공합니다.
How it works
Meissonic 은 비자기회귀 마스크 이미지 모델링 트랜스포머입니다. 픽셀이나 토큰을 하나씩 생성하는 전통적인 자기회귀 모델과 달리 마스킹 접근 방식을 사용해 이미지를 합성합니다. 텍스트‑투‑이미지 및 이미지‑투‑이미지 작업을 지원하며, FP8 양자화를 통해 메모리 사용량을 줄이고 성능을 유지하도록 최적화할 수 있습니다.
Who it’s for
이 프로젝트는 산업 규모의 컴퓨팅 자원이 필요 없이 자체 하드웨어에서 고품질·고해상도 이미지를 생성하고자 하는 AI 연구자, 개발자, 디지털 아티스트를 위한 것입니다.
Highlights
- High Resolution: Supports image generation up to 1024x1024.
- Consumer GPU Friendly: Optimized for accessibility on home hardware.
- Versatile Capabilities: Supports text-to-image, inpainting, and outpainting.
- Performance Optimization: Includes FP8 quantization support to lower VRAM requirements to approximately 8.7GB.