freestylefly/awesome-gpt-image-2

A Prompt-as-Code system and template library for GPT-Image2 that converts prose prompts into structured protocols for stable and controllable AI image generation.

yang0/handraw-style

A curated library of 216 numbered hand‑drawn illustration styles with ready‑made bilingual prompts, plus a simple web gallery. Users pick a style ID, add a subject, and get a prompt that works with text‑to‑image AIs, ensuring consistent visual tone without having to describe the style themselves.

abi/screenshot-to-code

screenshot‑to‑code is an open‑source app that converts UI screenshots, mock‑ups, or short screen recordings into ready‑to‑run front‑end code (HTML/CSS, React, Vue, etc.) using LLMs such as Gemini, GPT‑5.5, and Claude, with optional asset extraction and visual preview.

aldegad/sprite-gen

A Python CLI and agent skill that transforms a single base image into game-ready transparent sprite atlases or animation loops using AI generation pipelines.

perseval-BLR/NeuralScreen

NeuralScreen is a Windows utility that applies NVIDIA’s DLSS 5 neural up‑scaling (and optional frame generation) to the entire desktop or individual windows in real time, offering sharper visuals, screenshot/recording tools, and a configurable overlay UI.

s1dashu/ip-as-logo-skill

An Agent Skill that guides AI agents to generate simple, cute, and commercially viable IP mascot characters with consistent composition and style.

upscayl/upscayl

A free and open-source AI image upscaler that uses Real-ESRGAN and Vulkan to enlarge and enhance low-resolution images without losing quality.

wuyoscar/GPT-Image2-Skill

A prompt gallery, two reusable agent skills (image generation & prompt‑extraction), and a command‑line tool for OpenAI’s GPT‑Image 2/2.5 models.

yanliudesign/mono-color-skill

A design system and AI skill for creating professional one- or two-ink editorial prints, posters, and packaging with a focus on mechanical reproduction aesthetics.

danielgatis/rembg

A tool for removing image backgrounds that can be used as a CLI, Python library, or HTTP server, supporting various AI models and hardware acceleration.

Nutlope/logocreator

LogoCreator is an open‑source web app that generates brand‑ready logos using the FLUX‑2 image model via Together AI. Built with Next.js, TypeScript, Radix, and Tailwind, it runs without an account—just a Together AI key. Features include instant generation, FLUX‑1 edits, PNG/SVG export, optional rate‑limiting and auth, and a history dashboard. The repo provides clear setup steps and lists future tasks such as size selection, cost preview, reference‑logo uploads, and a showcase of redesigned famous logos.

basketikun/chatgpt2api

An OpenAI-compatible API proxy that reverse-engineers ChatGPT's web capabilities to provide image generation and editing services with built-in account pool management.

shitagaki-lab/see-through

A framework that decomposes single anime illustrations into up to 23 semantically distinct, inpainted layers exported as a PSD for 2.5D animation and rigging.

yjz211/vivid-figures-skill

Vivid Figures is an Agent‑Skill that equips AI assistants with 143 ready‑made chart templates and seven colour palettes. By feeding an assistant a data file (CSV, Excel, JSON) and a natural‑language request, the skill automatically selects a suitable plot, renders PNG/PDF outputs, and provides the full Python source for later tweaking. It follows the open Agent Skills spec, runs on Python 3.10+, and is free for personal, non‑commercial use.

Kizzuwatnaa/DLSS5-Autopilot

DLSS 5 Autopilot is a Windows utility that automatically adds NVIDIA’s DLSS 5 neural upscaling to games that never shipped it. It scans installed games, picks the appropriate community add‑on (native hook, OptiScaler, bridge, feeder, RTX Remix, etc.) based on the game’s graphics API and the user’s RTX 20‑50 GPU, downloads the needed DLLs at runtime, writes the configuration, and can cleanly uninstall everything. It supports many launchers, offers frame‑generation options, works on video/webcam streams, provides a CLI, and includes safety checks for anti‑cheat systems.

Comfy-Org/ComfyUI-Manager

ComfyUI‑Manager is an extension for the visual AI‑pipeline UI ComfyUI. It adds a “Manager” panel that lets users browse, install, update, enable/disable, and remove custom node packages and model files from a curated registry or local cache. Features include one‑click installs, snapshot‑based environment restore, workflow sharing to external sites, component copy‑&‑paste, a CLI tool for head‑less use, and extensive configuration (security level, network mode, pip/uv handling). It streamlines handling the large ecosystem of community‑made nodes that power Stable‑Diffusion image generation in ComfyUI.

T8RIN/ImageToolbox

An Android image editing tool that provides over 500 filters, AI-powered background removal, and batch processing for efficient photo manipulation.

ostris/ai-toolkit

Ostris AI Toolkit is an open‑source suite for fine‑tuning modern diffusion models (image, video, audio, edit) on consumer GPUs. It offers a web UI, an experimental manager that auto‑installs the correct PyTorch/Node.js/FFmpeg stack, and ready‑made YAML configs for LoRA/LoKr training. Installation works via simple scripts for Linux/macOS/Windows or manual virtual‑env steps. The toolkit supports dozens of recent checkpoints (FLUX‑1/2, SDXL, Qwen‑Image, etc.) and includes guides for cloud runs on RunPod and Modal.

SegFault42/HeliosGen

HeliosGen is an open‑source desktop app (macOS released, Windows/Linux pending) that provides a visual, node‑based canvas for building and running AI image and video generation pipelines. It works locally, sending requests to the kie.ai API (or optionally a Codex CLI) and stores all data on the user’s machine. Features include multi‑model support, reference images, parallel/sequential execution, and a real‑time history. The app is built with Tauri 2 (Rust) and a Next.js/React front‑end, uses SQLite for local persistence, and is MIT‑licensed.

facefusion/facefusion

FaceFusion is a high-performance face manipulation platform for swapping and editing faces in images and videos.

worldbench/DiffusionOPSD

DiffusionOPSD is an open‑source implementation of the “On‑Policy Self‑Distillation” algorithm for fine‑tuning diffusion image generators (Stable Diffusion 3.5‑M and Z‑Image‑Turbo). It repeatedly collects low‑noise states from a frozen behavior policy, builds explicit positive/negative targets using reward‑gradient steps, trains a new policy to match those targets, and refreshes the behavior policy via EMA. The repo provides scripts for installing dependencies, downloading reward‑model checkpoints, running quick smoke tests, and launching full‑scale distributed training for any of seven public evaluators (PickScore, CLIPScore, HPSv2.1/v3, Aesthetic, ImageReward, DeQA) or mixed‑reward objectives. Reported results show higher held‑out scores on 19/20 settings and 40‑63 % less GPU‑time versus prior methods. Licensed under Apache 2.0.

threerocks/hand-drawn-styles

A tool-agnostic library of 19 verified hand-drawn style prompt recipes that AI Agents can use to generate consistent, high-quality image prompts.

higgsfield-ai/cli

A cross‑platform command‑line interface for Higgsfield’s suite of generative AI models (image, video, 3D, audio) plus utilities for building React web apps, publishing browser games, and managing personal “Soul” avatars—all from the terminal.

LiamGvchi/gc-minimal-zine-poster

A Codex skill that transforms themes or photos into minimal, paper-textured editorial posters using a structured prompt-compilation system.

NVlabs/Sana

SANA is an open‑source, efficiency‑oriented diffusion codebase from NVIDIA for high‑resolution image and video generation. It includes multiple model families (SANA, SANA‑1.5, SANA‑Sprint, SANA‑Video/Video 2.0, SANA‑WM, SANA‑Streaming) and a reinforcement‑learning wrapper (Sol‑RL). Key tricks are linear attention, a 32× compression auto‑encoder (DC‑AE), decoder‑only text encoders, and sCM distillation, enabling 4K image generation and minute‑long 720p video generation with far less GPU cost than comparable models. The repo provides training/inference scripts, pretrained checkpoints on Hugging Face, Diffusers integration, ComfyUI nodes, and demos. Licensed Apache 2.0.

alisaitteke/photoshop-mcp

An MCP server that enables AI assistants to control Adobe Photoshop using plain-language commands for tasks like background removal and portrait retouching.

LunarXuan/image-prompt-reverse

A Codex skill that reverse-engineers reference images into high-fidelity positive and negative prompts for AI image-generation tools.

invoke-ai/InvokeAI

Invoke AI is an open‑source, locally‑run visual‑media generation platform. It provides a web‑based UI with a unified canvas, node‑based workflow editor, gallery management, and support for dozens of modern text‑to‑image models (Stable Diffusion, Flux, Qwen‑Image, etc.). Users install via the provided launcher, run a local server, and create or refine images through an intuitive browser interface. The project targets artists, designers, and developers who need a flexible, extensible AI image creation tool.

ciddwd/overlay-translator

Screen Translator is an Android app that captures the screen, runs OCR, translates the text (via on‑device models, LLM APIs, or traditional MT services), and overlays the result back onto the screen. It supports real‑time translation, batch image processing, word‑level dictionary lookup, TTS, and extensive customization, with both offline and cloud options.

Acly/krita-ai-diffusion

A Krita plugin that integrates generative AI diffusion models into the painting workflow, offering tools for inpainting, live painting, and precise structural control.

CookSleep/gpt_image_playground

A professional web UI for OpenAI's gpt-image-2.5 API that enables text-to-image generation, mask-based editing, and conversational AI agent image creation.

toyxyz/ComfyUI_toyxyz_test_nodes

A set of custom ComfyUI nodes for webcam/screen capture, region‑mask creation, OpenPose editing, video trimming/concatenation, and a MiniMax‑H3 prompt builder, enabling richer image‑and‑video generation workflows.

shanliuling/dsh-image-gen

An AI image creation suite for DeepSeek Harness that provides chat-based generation, a batch production studio, multi-model comparison, and local ComfyUI integration.

xororz/local-dream

An Android application that enables local Stable Diffusion image generation with Snapdragon NPU acceleration for improved performance.

kijai/ComfyUI-KJNodes

ComfyUI‑KJNodes is a plug‑in for the visual Stable Diffusion UI (ComfyUI) that adds utility nodes, cross‑sub‑graph Set/Get handling, and many keyboard/mouse shortcuts to make large generation graphs easier to build and maintain.

Hugo-Dz/spritefusion-pixel-snapper

A tool that snaps pixels to a perfect grid and quantizes colors to fix messy, inconsistent pixel art generated by AI.

mcmonkeyprojects/SwarmUI

SwarmUI is an open‑source, web‑based UI for running AI image, video, and audio generation models (Stable Diffusion, Flux, etc.). It offers a beginner‑friendly Generate tab, an advanced Comfy‑style workflow editor, multi‑GPU “swarm” support, and extensible plugins. Installers exist for Windows, Linux, macOS (Apple‑silicon), and Docker, with cloud‑GPU templates for Runpod and Vast.ai. The project is in an “Almost‑Release” beta (v0.9.8) and is MIT‑licensed, though it can auto‑install GPL/AGPL components.

Comfy-Org/ComfyUI_frontend

ComfyUI_frontend is the official browser/Electron UI for the ComfyUI AI workflow engine. It offers a node‑graph editor, mask editor, queue/history view, multilingual UI, and a JavaScript extension API for custom panels, commands, keybindings, and toast messages. Releases follow a 2‑week dev → 2‑week freeze cycle, with daily nightlies available via a launch flag.

jau123/MeiGen-AI-Design-MCP

MeiGen AI Design MCP is an MCP‑compatible server (npm package or remote HTTP endpoint) that lets AI assistants generate images, videos, and ecommerce‑focused assets (background removal, product‑detail batches, marketing posters, AI backgrounds, up‑scaling). It supports MeiGen Cloud, OpenAI‑compatible APIs, and local ComfyUI, provides 1,446 curated prompts, and offers a CLI and HTTP API for non‑MCP use.

liyue-aigc/xianxia-visual-director

A visual direction system for Codex/Agent skills that transforms simple ideas into structured, cinematic image prompts for monumental Eastern xianxia environments.

Akegarasu/lora-scripts

A GUI and script preset collection for training Stable Diffusion LoRA and Dreambooth models, providing an integrated environment for dataset tagging and training.

popopo-99/zy-cinematic-realism

A cinematic visual workflow for LLMs that transforms narrative ideas into professional, model-specific prompts for AI image generators to achieve authentic cinematic realism.

lbouaraba/comfyui-krea2edit

A ComfyUI node pack for instruction-based image editing using Krea 2, enabling high-fidelity identity preservation through dual latent and semantic conditioning.

llmsresearch/paperbanana

PaperBanana is an open‑source Python tool that turns method‑section text (or data files) into publication‑ready diagrams and statistical plots using a multi‑agent LLM/VLM pipeline, with support for batch runs, venue‑specific styling, and multiple AI providers.

goohai/Goohaitools-comfyui

A comprehensive toolkit of 70+ custom nodes for ComfyUI designed for professional image production, specializing in batch processing, automated ID photo creation, and advanced mask manipulation.

carolinaaafy/travel-memory-sticker-card

A Codex skill that transforms travel photos into collectible digital memory sticker cards.

jtydhr88/ComfyUI-See-through

A ComfyUI plugin that decomposes single anime illustrations into layered 2.5D models with depth ordering, specifically designed for Live2D workflows.

TheJoeFin/Text-Grab

Text Grab is a Windows‑only desktop app that captures any visible text (screenshots, PDFs, UI elements) using local OCR (WinAI, WinRT OCR, or Tesseract) and provides built‑in cleanup, spreadsheet editing, regex‑based extraction, reusable grab templates, bulk folder processing, and a Chrome/Edge extension. All processing stays on‑device, with optional NPU‑accelerated inference on Copilot+ PCs. Install via Microsoft Store, GitHub releases, or package managers; source can be built with Visual Studio or the .NET SDK.