op7418/guizang-yingzao-skill

An AI-powered design skill that transforms architectural and cultural photos into art-directed editorial posters with integrated Chinese typography and spatial awareness.

lixiaoxiao9888-create/manju-laoli-skill

An industrial-grade screenwriting and audiovisual directing system for AI agents to produce viral short dramas and animated shorts.

Blizaine/Maestro

A one-click AI video, image, and audio studio featuring a Director mode that uses LLMs to automatically plan and generate music videos and short films.

jtydhr88/ComfyTV

A canvas-based media workbench for ComfyUI that integrates AI generation with professional-grade editing tools for image, video, audio, music, and 3D assets.

ZSeven-W/openpencil

An open-source AI-native vector design tool that generates UI layouts from natural language prompts and exports them to production code.

nexmoe/VidBee

A free, open-source desktop app that downloads video and audio, transcribes them locally using ASR models, and allows AI-powered summarization and analysis of the transcripts.

wiltodelta/remove-ai-watermarks

A tool for removing visible labels, invisible pixel watermarks, and AI provenance metadata from AI-generated images and videos.

aaronyi97/image-story-video-wizard

A step-by-step AI video production workflow for Codex and WorkBuddy that guides users from topic selection to final rendering of story-based videos.

mnemosyne-oss/mnemosyne

Mnemosyne is a local-first, SQLite-backed memory layer for AI agents. It gives assistants persistent memory across sessions via a three-tier architecture (working memory, episodic memory, temporal knowledge graph), with hybrid semantic/keyword search, heavy vector compression, a Python SDK, CLI, and built-in MCP server. It's privacy-focused: no telemetry, local by default, with optional client-side-encrypted sync.

Open-LLM-VTuber/Open-LLM-VTuber

An open-source, voice-interactive AI companion featuring a Live2D avatar and visual perception, capable of running entirely offline for private, real-time conversations.

NomaDamas/CozyClay

A browser-based 3D staging studio for blocking scenes and sequencing AI-generated motion, featuring direct AI control via the Model Context Protocol.

TencentARC/Pixal3D

Pixal3D is a high-fidelity 3D asset generator that creates detailed geometry and PBR textures from a single image or multiple views using pixel-aligned back-projection.

tigerowo/infinite-canvas

An open-source creative workbench that integrates an infinite canvas with AI image, video, and audio generation tools for iterative visual design.

lightningpixel/modly

A local, open-source desktop application that turns photos into 3D meshes using AI models running on the user's GPU.

techjarves/Uncensored-Local-Studio

A zero-configuration, offline AI studio that integrates Stable Diffusion, LLMs, Whisper, and Kokoro TTS into a single private desktop interface with hardware acceleration.

kgoedecke/doop

An open-source multiplayer design canvas that allows humans and AI agents to collaborate live by streaming HTML designs into shared frames.

liustack/modlens

A vision plugin that gives text-only AI models the ability to see by transcribing pasted images into structured text using external vision engines.

waooAI/waoowaoo

An AI-powered tool that automatically transforms novel text into short dramas or comic videos by generating storyboards, consistent characters, and AI voiceovers.

jianchang512/pyvideotrans

An open-source video translation tool that automates speech recognition, subtitle translation, and AI dubbing with support for voice cloning and multi-role synthesis.

AutoArk/TinyEngram

An open research project exploring Engram-based memory injection for LLMs and Stable Diffusion to enable parameter-efficient knowledge updates without catastrophic forgetting.

jingyaogong/minimind-v

A minimal implementation of a Vision Language Model (VLM) that allows users to train a 65M parameter multimodal model on a single consumer GPU in just 2 hours.

WEIFENG2333/VideoCaptioner

An all-in-one video subtitle tool that uses LLMs for speech recognition, subtitle optimization, and translation to automate the video captioning process.

penecho/penecho

PenEcho is an open‑source AI‑enhanced infinite canvas that lets you write, sketch, and attach documents, then ask LLMs to generate explanations, formulas, plots, or interactive widgets directly on the canvas. It supports multiple model back‑ends (Kimi, Claude, OpenAI, DeepSeek, etc.), offers a multi‑step “PenEcho Agent” for reading files and web research, and includes an optional cloud service for cross‑device sync and public sharing. The app runs locally (Node.js) with a secure six‑digit LAN guard, and the UI is available as a downloadable desktop binary or via npm.

gastownhall/gastown

Gas Town is an open‑source Go‑based workspace manager that lets you run many AI coding agents (Claude, Copilot, Codex, Gemini, etc.) on multiple Git projects while persisting every step in a Git‑backed ledger. A central AI coordinator (the *Mayor*) creates *convoys* of work items (*beads*), hands them to worker agents (*Polecats*), and a set of watchdog services (Witness, Deacon, Dogs) keep the system healthy. Completed work is merged via a Bors‑style queue (*Refinery*), and an optional federation layer (*Wasteland*) lets different installations share tasks. Installable via Homebrew, `go install`, or Docker, it provides a CLI (`gt`) for creating rigs, crews, convoys, and for monitoring progress.

Darkatse/TauriTavern

TauriTavern is a cross‑platform native client (Windows, macOS, Linux, Android, iOS) that wraps the SillyTavern AI‑chat UI inside a Rust‑based Tauri application. It replaces the original Node.js backend with Rust, offering one‑click installation, local‑only data storage, LAN/cloud sync, a built‑in extension manager, and an evolving AI‑agent framework. Distributed under AGPL‑3.0, it can be installed via package managers or direct downloads, and the source follows a clean‑architecture Cargo workspace.

PurpleDoubleD/locally-uncensored

A plug-and-play local AI studio for Windows and Linux that integrates uncensored chat, image and video generation, and a coding agent into a single, no-cloud desktop application.

NVIDIA/cosmos

An open platform of omnimodal world models and tools for building Physical AI, enabling the joint processing and generation of text, vision, audio, and action sequences for robotics and autonomous systems.

koharu-rs/koharu

An ML-powered manga translator written in Rust that automates text detection, OCR, translation, and generative inpainting for a local-first workflow.

OpenBMB/ChatDev

ChatDev 2.0 (DevAll) is an open‑source, zero‑code platform that lets you design and run multi‑LLM‑agent workflows via a visual web console or a tiny Python SDK. Define agents and their connections in YAML, launch from the browser or code, and use built‑in templates for data viz, 3‑D generation, game dev, research, etc. The backend is FastAPI, the frontend Vue 3, and the project ships Docker/Makefile support, extensible Python tools, and research‑grade orchestrators (Puppeteer, MacNet).

oil-oil/oil-motion

An Agent-driven animation skill that converts AI-generated continuous motion into interactive web elements responsive to scrolling, mouse movement, and touch.

jjyaoao/HelloAgents

HelloAgents is a Python library that provides a production‑grade framework for building multi‑agent AI applications. It wraps OpenAI‑compatible, Anthropic, and Gemini LLMs, offers built‑in tools (file I/O, task delegation, todo tracking), and includes engineering features such as context management, session persistence, circuit‑breaker, optimistic locking, observability, streaming SSE, and async lifecycle. Install via `pip install hello-agents`, configure a `.env` with your API credentials, and start agents like `ReActAgent` with a registered tool set.

Open-Less/openless

An open-source voice-input tool for macOS and Windows that uses AI to polish spoken text and insert it directly at the cursor, featuring a specialized mode for composing AI prompts.

zhayujie/CowAgent

CowAgent is an open‑source, multi‑modal AI assistant platform. It receives messages from many chat channels, uses a chosen LLM to plan and reason, accesses a rich set of system tools and user‑defined “skills”, and maintains a three‑tier memory plus an auto‑curated knowledge base. Installable via a one‑line script or Docker, it supports dozens of LLM providers and runs on a personal computer or server.

BasedHardware/omi

An open-source AI memory system that captures screen and conversation data across wearables, desktop, and mobile to provide real-time transcription, summaries, and a searchable AI chat.

flyteorg/flyte

Flyte 2 is an open‑source Python framework for defining, orchestrating and serving ML pipelines and AI agents. It lets you write ordinary Python functions as tasks, compose them into workflows, and run them locally or on a future Kubernetes‑native backend. The project includes a CLI, a developer‑focused TUI, and first‑class support for FastAPI model serving.

StarTrail-org/PixelRAG

PixelRAG is an open‑source toolkit that renders web pages, PDFs or images into screenshot tiles, embeds those tiles with a fine‑tuned visual‑language model (Qwen3‑VL‑Embedding‑2B), builds a vector index (FAISS or Qdrant), and provides a fast API for visual‑plus‑text search. It ships a CLI (`pixelshot`), an orchestrator (`pixelrag`), a hosted 8.28 M Wikipedia visual index, and a Claude Code plugin that lets Claude “see” screenshots.

MirroS-Lab/Code-as-World

Code-as-World is a framework that represents the physical world as executable code to enable AI models to perform quantitative physical reasoning by discovering world representations through iterative simulation.

nicobailon/pi-web-access

Pi Web Access is an npm package that equips the Pi AI‑agent with zero‑config web search, multi‑provider routing, GitHub repo cloning, YouTube/local‑video understanding, PDF conversion, caching, and a structured source‑checking tool—all exposed via a small TypeScript API (`web_search`, `fetch_content`, etc.). Install with `pi install npm:pi-web-access`; it works out‑of‑the‑box using Exa MCP and can be extended by adding API keys to `~/.pi/web-search.json`.

ddcat-ai/open-ai-canvas

An open-source AI film and short-drama creation workbench that integrates an infinite canvas with multimodal generation tools to turn text briefs into cinematic assets.

antirez/h3.c

A native Metal inference engine for MiniMax-H3 on Apple Silicon, enabling local text-to-video and text-to-audio generation with advanced memory and performance optimizations.

Anionex/agent-vision-toolkit

A toolkit that gives text-only LLM agents visual capabilities through task-aware vision tools and seamless proxy integrations for image Q&A and GUI automation.

linyqh/NarratoAI

An automated AI-powered tool for movie and TV commentary videos that handles scriptwriting, video editing, voiceovers, and subtitles in one workflow.

Forget-C/Jellyfish

An end-to-end AI production workspace for short dramas that manages the full pipeline from script breakdown and asset consistency to video generation.

umlx5h/LLPlayer

LLPlayer is a Windows‑only C# media player that adds AI‑driven subtitle generation (Whisper ASR), real‑time translation (via cloud or local LLMs), OCR for bitmap subtitles, dual‑subtitle display, word lookup, and online video support. It targets language learners and is open‑source under GPL‑3.0.

llm-as-a-verifier/llm-as-a-verifier

llm-as-a-verifier is a Python framework that turns LLMs into fine‑grained, probabilistic verifiers for AI agent outputs. It scores candidate trajectories, selects the best‑of‑N via an efficient Probabilistic Pivot Tournament, tracks progress step‑by‑step, and supports multimodal (image) inputs. The library achieves state‑of‑the‑art results on several agentic benchmarks and includes scripts for reproducible evaluation, a cache‑optimised token usage tracker, and a Claude Code plugin (TurboAgent) for seamless integration.

xinnan-tech/xiaozhi-esp32-server

A backend server for the xiaozhi-esp32 project that enables ESP32 devices to become AI assistants with voice, vision, and smart home control capabilities.

llmsresearch/paperbanana

An agentic framework that automates the generation of publication-quality academic diagrams and statistical plots from text descriptions and data.

KangLiao929/Puffin

A series of unified multimodal models for 3D world modeling that enable camera-centric spatial intelligence and 3D world generation using native physics, geometry, and appearance states.