lidge-jun/ima2-gen

Local-first visual generation runtime and studio for people and coding agents, with reproducible image and video workflows across multiple providers.

What it solves

ima2-gen provides a local-first visual generation studio that simplifies the process of creating and iterating on images and videos across multiple AI providers. It eliminates the need to switch between different web interfaces by consolidating various APIs and OAuth flows into a single runtime, while adding professional iteration tools like node-based branching, a canvas for cleanup, and specialized workflows for coding agents.

How it works

The project acts as a visual generation runtime and studio. It connects to a wide array of providers including OpenAI (OAuth/API), Grok (OAuth/API), Gemini (via Antigravity CLI or direct API), NovelAI, and ComfyUI. It manages these connections through a core registry and provides both a CLI and a web-based UI for interaction.

Key operational modes include:

  • Classic Mode: Standard text-to-image/video generation with reference image support.
  • Node Mode: Allows users to branch a successful generation into multiple directions, creating a visual graph of iterations.
  • ** uma Canvas Mode**: A dedicated space for zooming, panning, annotating, and cleaning backgrounds (including a one-click GPT transparency tool).
  • Storyboard Mode: Maintains character and scene continuity across sequential frames for video and image production.

Who it’s for

  • Visual Creators: People who need a unified interface for multiple AI image and video generators.
  • AI Coding Agents: The project includes packaged "skills" (Markdown instructions) that allow agents to handle asset production, frontend design, and UI/UX discovery using the ima2 CLI.
  • Developers: Those who want a local-first tool with reproducible workflows and a local gallery for asset management.

Highlights

  • Multi-Provider Support: Integrated access to OpenAI, Grok, Gemini, NovelAI, and ComfyUI.
  • Advanced Iteration: Node-based branching and a Canvas for targeted cleanup and annotation.
  • Video Capabilities: Text-to-video and image-to-video generation via Grok, including keyframe extraction.
  • Agent-Ready: Specialized skill sets for frontend and UI/UX design tailored for AI coding agents.
  • Local-First: Local gallery with session-aware history and local prompt library imports.
  • Vectorization: Ability to trace raster art into real SVG paths.
  • SSE Multiplexing: Uses a single Server-Sent Events connection to prevent browser connection limits during concurrent jobs.

Related

  • Project
  • Project
  • Project
  • Project
  • Project