cracker0dks/CaptchaSolver

CaptchaSolver is a JDownloader 2 plug‑in that uses a YOLO/Darknet neural‑network model (run via Node.js) to automatically solve 6‑digit and geometric captchas for a handful of file‑hosting sites. It provides pre‑compiled binaries for Windows/Linux (amd64) and build instructions for other platforms, letting users run JDownloader fully unattended on servers or NAS devices.

Nutlope/easyedit

EasyEdit is a prompt-based image editing tool powered by Flux 2 and Together AI that allows users to modify images with a single text prompt.

mmartial/ComfyUI-Nvidia-Docker

Docker images that run the ComfyUI Stable‑Diffusion web UI with NVIDIA GPU acceleration, offering GPU‑specific tags, UID/GID mapping, and built‑in ComfyUI‑Manager for easy updates.

SerialLain3170/adeleine

A deep-learning framework for automatic line art colorization that supports guidance via colored hints, text tags, or reference images.

Bing-su/adetailer

An extension for the Stable Diffusion web UI that automates masking and inpainting to refine details in generated images.

cocktailpeanut/fluxgym

Flux Gym provides a Gradio web UI that lets you train FLUX diffusion LoRAs on 12‑20 GB GPUs. It wraps Kohya‑ss training scripts, offering full script options, automatic model download, optional sample‑image generation, and one‑click publishing to Hugging Face. Install via Pinokio, Docker, or manual venv setup.

deepgenteam/deepgen

DeepGen 1.0 is a lightweight 5B parameter unified multimodal model that integrates image generation, editing, and reasoning into a single framework.

nihui/zimage-ncnn-vulkan

A portable ncnn-based implementation of the Z-Image generator that enables high-performance image generation, inpainting, and ControlNet support on any Vulkan-capable GPU.

Lornatang/SRGAN-PyTorch

A PyTorch reimplementation of SRGAN that uses generative adversarial networks to produce photo-realistic 4x upscaled images with recovered fine textures.

asagi4/comfyui-prompt-control

A ComfyUI extension that lets you control complex AI image generation workflows (like LoRA loading, prompt scheduling, and regional prompting) using simple text prompts instead of manually wiring dozens of nodes.

yanokusnir-ai/one-node-flux-2-klein

A ComfyUI custom node that wraps the full FLUX.2 [klein] workflow into a single, self-contained UI widget to eliminate complex node graphs.

worldbench/DanceOPD

DanceOPD is a research library that distills multiple diffusion‑based image‑generation capabilities (text‑to‑image, local edit, global edit, realism, etc.) into one model. It treats each teacher as a velocity field, queries the teacher on the student’s own rollout states, and trains with a simple velocity‑MSE loss. The repo supports SD‑3.5‑Medium and Z‑Image backends, includes a tiny smoke‑test, and provides scripts/configs for full training with LoRA adapters.

hustvl/ControlAR

ControlAR provides controllable image generation for autoregressive models, allowing spatial guidance via edges, depth, or masks with support for arbitrary resolutions.

Linfeng-Tang/SwinFusion

SwinFusion is a Swin Transformer-based framework for general image fusion that combines information from multiple source images across various domains, such as infrared, visible, and medical imaging.

wafer-bob/ASASR

ASASR is a faithful image super-resolution framework based on FLUX.1-dev that uses Adversarial Sobolev Alignment to reduce artifacts and improve structural fidelity in upscaled images.

bentoml/BentoDiffusion

A collection of example projects for self-hosting and deploying various Stable Diffusion family models using the BentoML framework.

shiimizu/ComfyUI_smZNodes

Custom ComfyUI nodes (CLIP Text Encode++ and Settings) that replicate AUTOMATIC1111's prompt parsing and embedding, enabling identical image generation between stable‑diffusion‑webui and ComfyUI.

Windsander/ADI-Stable-Diffusion

A C++ library and CLI tool that enables high-performance, Python-free inference for Stable Diffusion models using ONNXRuntime across multiple platforms.

mrhan1993/Fooocus-API

A FastAPI-powered REST API for Fooocus that allows developers to programmatically generate high-quality images without manual parameter tweaking.

melMass/comfy_mtb

A collection of custom nodes for ComfyUI that extends the image generation workflow with additional utility tools.

lobehub/sd-webui-lobe-theme

A modern, highly customizable interface framework for Stable Diffusion WebUI that improves the workflow and visual experience of AI image generation.

yejy53/RealGen

RealGen is a text-to-image generator that uses AI detectors as reward signals during reinforcement learning to produce highly photorealistic images with fewer artifacts.

cure-lab/MMA-Diffusion

A framework for evaluating the security of Text-to-Image models by using multi-modal adversarial attacks to bypass prompt filters and safety checkers to generate NSFW content.

ATH-MaaS/Ovis-Image

Ovis-Image is a 7B text-to-image model optimized for high-quality text rendering and layout precision, designed to run efficiently on accessible hardware.

nachifur/RDDM

A residual denoising diffusion model framework for high-quality image restoration and generation, supporting tasks like deraining and deblurring.

viiika/Meissonic

Meissonic is a non-autoregressive masked generative transformer for efficient high-resolution text-to-image synthesis designed to run on consumer GPUs.

nxnai/Voost

Voost is a unified Diffusion Transformer that enables high-quality bidirectional virtual try-on and try-off, allowing users to add or remove garments from images of people.

Ammmob/PixelSmile

PixelSmile is a fine-grained facial expression editing tool that allows users to change emotions in images of humans and anime characters while preserving their identity.

nekhtiari/image-similarity-measures

A Python package and CLI that computes eight standard image‑similarity metrics (RMSE, PSNR, SSIM, FSIM, ISSM, SRE, SAM, UIQ) for two images. Install via pip, optionally add speed‑ups, and use either the `image-similarity-measures` command or the `evaluation` function in code. Useful for quantitatively assessing image‑processing or AI models, especially in remote‑sensing contexts.

Stability-AI/stability-sdk

A client SDK for the Stability API that enables image generation, upscaling, and animation through a Python library and command-line tools.

mo-browser-apps/icons

A desktop app that uses AI to generate macOS app icons in the .icns format, allowing users to describe their desired design and iteratively refine the result.

HalfAI1102/anthropic-art

An agent skill that generates hand-drawn editorial illustrations in the Anthropic visual style, transforming abstract concepts into minimalist visual metaphors.

Nutlope/blinkshot

An open-source real-time AI image generator powered by Juggernaut Lightning Flux and Together AI.

Sygil-Dev/sygil-webui

A web-based user interface for Stable Diffusion that provides tools for image generation, upscaling, and prompt engineering with optimizations for low-VRAM GPUs.

ShuaixinHuang/image-multiple-angles-3d-camera

A 3D-controlled image generation tool that allows users to change the camera viewpoint of an uploaded image using an interactive 3D widget and the Qwen-Image-Edit-Plus model.

Lakonik/LakonLab

A high-performance codebase for experimenting with large diffusion models, implementing advanced flow-matching and distillation techniques for efficient few-step image generation.

Mukosame/Anime2Sketch

A sketch extractor that uses deep learning to convert illustrations, anime art, and manga into clean line art and sketches.

apple-aiml-research/ml-stable-diffusion

A toolkit for converting and running Stable Diffusion models on Apple Silicon using Core ML, featuring advanced weight compression for mobile deployment.

spectertoucanberth/thumbnail-maker-lab

Thumbnail Maker Lab is a professional design toolkit for creating thumbnails that combines vector graphics with AI-powered image editing and batch processing.

cathrynlavery/diagram-design

Diagram Design is an AI‑agent skill that generates a large library of brand‑aware, static HTML + SVG editorial diagrams (architecture, flowcharts, charts, maps, etc.) in three visual variants, with automatic colour/font onboarding and accessibility support. It integrates with Claude Code, Codex, Copilot, Factory Droid, Pi, Kiro and other Agent‑Skill platforms.

liujuntao123/smart-excalidraw-next

Smart Excalidraw is a Next.js web app that turns natural‑language descriptions into editable Excalidraw diagrams using an LLM (e.g., Claude sonnet‑4.5). It supports 20+ diagram types, auto‑routes arrows, and offers either a shared server‑side LLM (access‑password) or personal API‑key configuration. Run locally with `pnpm dev` or use the hosted demo.

ramjke/Translumo

Translumo is a Windows desktop app that captures a screen region, runs OCR (Windows OCR, Tesseract, EasyOCR), selects the best result with a small ML model, translates the text via DeepL/Google/Yandex/Papago, and overlays the translation back onto the screen. Designed for real‑time use in games, it supports many source/target languages, proxy rotation, and low‑latency operation. The project is open‑source (Apache 2.0) and built with .NET 8, Visual Studio 2022, and several OCR/vision libraries.

Picsart-AI-Research/MI-GAN

A lightweight GAN-based image inpainting model optimized for mobile devices, providing high-quality results with significantly fewer parameters and faster inference than SOTA models.

Physton/sd-webui-prompt-all-in-one

An extension for stable-diffusion-webui that enhances the prompt input experience with automatic translation, prompt organization tools, and ChatGPT integration.

TTPlanetPig/Comfyui_TTP_Toolset

A collection of ComfyUI nodes featuring Smart Tile 2.0, an object-aware tiled img2img workflow that uses semantic detection and variable-size tiles for high-detail upscaling.

leeguooooo/image-use

A zero-dependency Python CLI that enables image generation using existing ChatGPT or Gemini subscriptions, bypassing the need for a an OpenAI API key.

groundboxerrespect/Dlls5-auto

A self-hosted image enhancement service that applies photorealistic neural rendering aesthetics, such as cinematic lighting and bloom, to any image via a REST API.

HyperGAN/HyperGAN

A composable GAN framework built on PyTorch that makes it easy for developers and artists to train, share, and deploy generative adversarial networks.