nexu-io/open-design

An open-source, local-first design tool that enables coding agents to generate and render high-fidelity prototypes, decks, and motion graphics based on a shared design system.

SegFault42/HeliosGen

A free and open-source visual workflow builder for chaining AI image and video generation models on an infinite node-based canvas.

Comfy-Org/ComfyUI

A modular, node-based AI engine for visual professionals to create complex workflows for generating images, video, audio, and 3D models with precise parameter control.

livekit/agents

A framework for building real-time, programmable multi-modal AI agents that can see, hear, and speak, featuring integrated job scheduling and telephony support.

ahujasid/camera-to-blender

A tool that converts photos of real-world objects into 3D models using AI and automatically imports them into Blender.

ahujasid/blender-mcp

A bridge connecting Blender to LLMs via the Model Context Protocol, enabling prompt-assisted 3D modeling, scene creation, and asset manipulation.

dramaclaw/dramaclaw

A source-available AI drama production pipeline that transforms scripts into finished films using a dual-mode workflow of a node-based canvas and a structured production pipeline.

chatfire-AI/huobao-drama

An AI-powered automated production platform that streamlines the creation of short dramas, from script generation and character design to final video synthesis.

Merserk/dlss5-visual-enhancer

A Windows application that applies NVIDIA DLSS 5 Neural Rendering and Frame Generation to images and videos via a local Gradio interface.

yejy53/Editable-Design

A coding-agent-driven framework that turns prompts into editable, structured visual designs using HTML and semantic layers instead of flattened images.

Anil-matcha/Open-Generative-AI

An unrestricted, open-source AI studio for generating images, videos, and audio using 400+ models without content filters or subscription fees.

Tencent/WeMM-Embedding

A family of universal multimodal embedding models that provide unified representations for text, images, videos, and visual documents to enable high-performance cross-modal retrieval.

NoizAI/HelixWorld

HelixWorld 1.0 is a forthcoming real‑time model that jointly generates video and spatial audio as a user moves a virtual camera. Trained on first‑person video/audio and game‑engine captures, it predicts the next frame and sound field causally, then distills the pipeline for interactive speeds. Code, weights, and a technical report are slated for release soon; the project will be Apache‑2.0 licensed.

OpenSenseNova/SenseNova-U1

SenseNova‑U is an open‑source 8 B‑parameter unified multimodal model family (U1 and U1.5) that natively handles text‑to‑image, image editing, VQA, and interleaved generation. Built on the NEO‑unify architecture, it removes separate visual encoders/VAEs, delivering high‑quality 4K generation, strong infographic rendering, and efficient inference (including LoRA and community‑provided 8‑bit GGUF quantizations). The repository contains training code, inference scripts, benchmarks, and deployment guides, all under Apache 2.0.

ningzimu/codex-ppt-skill

An AI agent skill that converts articles, reports, and papers into professional image-based PowerPoint presentations by automating outline planning and visual style generation.

darkzOGx/youtube-automation-agent

An open-source AI agent system that automates the entire YouTube channel lifecycle, from trend research and scriptwriting to video production, publishing, and analytics-driven optimization.

HBAI-Ltd/Toonflow-app

An AI-powered workbench for short drama production that automates the workflow from scriptwriting and storyboarding to final video generation.

basketikun/infinite-canvas

An open-source, node-based workbench for AI image creation that integrates generation, editing, and prompt libraries on an infinite canvas.

pipecat-ai/pipecat

An open-source Python framework for building real-time voice and multimodal AI agents with support for multi-agent orchestration and low-latency transports.

wide-trace/open-higgsfield

An open-source, self-hosted studio for generating images and videos using 40 different AI models via a user-provided platform key.

jaredrhod/fullstack-agent

A framework that assembles a multimodal AI assistant featuring persistent text-based memory, voice interaction, visualizers, and webcam-based gesture controls.

PurpleDoubleD/locally-uncensored

A plug-and-play local AI studio for Windows and Linux that integrates uncensored chat, image and video generation, and a coding agent into a single, no-cloud desktop application.

zenstory-ai/drama-skills

An AI-powered short drama creation workflow that transforms ideas or novels into scripts, storyboards, and production prompts via a set of modular agent skills.

moeru-ai/airi

An open-source framework for creating interactive AI virtual characters and VTubers capable of chatting, playing games, and interacting across multiple platforms.

ooolabdev/ooosplat

A desktop application that converts local videos into 3D Gaussian Splatting models by automating the pipeline of frame extraction, camera reconstruction, and training.

alchaincyf/huashu-design

A design-automation skill for AI agents that transforms text prompts into professional interactive prototypes, motion graphics, and editable presentation decks.

yihong0618/bilingual_book_maker

bilingual_book_maker is a Python CLI that translates public‑domain books (EPUB, PDF, TXT, MD, SRT) into bilingual versions using a variety of AI translation services (OpenAI, Claude, Gemini, DeepL, Google, Qwen‑MT, etc.). It parses the source, sends passages to the chosen model, and outputs a new bilingual EPUB (or paired TXT/SRT). Features include multi‑model selection, custom provider configs, plan mode for selective translation, tag control, context‑aware prompting, parallel processing, and resume support.

vllm-project/vllm-omni

A framework that extends vLLM to provide fast and efficient serving for omni-modality models, supporting text, image, audio, video, and robotic action generation.

TencentCloud/Octop

Octop is an open‑source, self‑hosted AI assistant platform. It runs a single FastAPI‑based process that serves a React dashboard, a CLI, and bots for Feishu, DingTalk, QQ, Discord, WeCom, etc. Users can create multiple agents with distinct MBTI‑style personalities, each with its own persistent memory and configurable LLM provider (OpenAI‑compatible, DashScope, Ollama, …). All data stays locally under `~/.octop/`, and the system includes guardrails, JWT isolation, browser/terminal automation, and a bidirectional ACP protocol for IDE integration. Installation is via a one‑line script, PyPI, or Docker; the project is MIT‑licensed and actively maintained.

localai-org/kimodo.cpp

A GGML/C++ implementation of NVIDIA's Kimodo model that generates human motion sequences (SMPL-X22) from text prompts on CPU or Vulkan.

AutoArk/TinyEngram

An open research project exploring Engram-based memory injection for LLMs and Stable Diffusion to enable parameter-efficient knowledge updates without catastrophic forgetting.

flatkey-ai/flatkey-cli

A command-line interface for unified multimodal AI generation, allowing media teams and AI agents to generate images, videos, audio, and text through a single API key and credit balance.

penecho/penecho

PenEcho is an open‑source AI‑enhanced infinite canvas that lets you write, sketch, and attach documents, then ask LLMs to generate explanations, formulas, plots, or interactive widgets directly on the canvas. It supports multiple model back‑ends (Kimi, Claude, OpenAI, DeepSeek, etc.), offers a multi‑step “PenEcho Agent” for reading files and web research, and includes an optional cloud service for cross‑device sync and public sharing. The app runs locally (Node.js) with a secure six‑digit LAN guard, and the UI is available as a downloadable desktop binary or via npm.

AutoArk/EVA-OS

An AIOS for real-time multimodal applications and smart hardware, providing an integrated development experience from AI-native coding to on-device inference.

lipku/LiveTalking

A real-time interactive streaming engine for digital humans that synchronizes audio and video for lifelike AI-driven conversations.

PhiloLabs/fable51-worlds

A system that uses AI agent swarms to generate verifiable, walkable 3D worlds from text, images, or video, rendered as code-based Three.js applications.

jangles-byte/Pythia

A local, swarm-intelligence global forecasting system that fuses over 40 real-time data feeds into a 3D intelligence globe to predict world events.

op7418/guizang-yingzao-skill

An AI-powered design skill that transforms architectural and cultural photos into art-directed editorial posters with integrated Chinese typography and spatial awareness.

bytedance/UI-TARS-desktop

A multimodal AI agent stack that enables natural language control of desktop applications and web browsers through visual recognition and precise GUI interaction.

Core-Mate/OpenGUI

A mobile GUI agent framework for Android that enables AI agents to understand and operate app interfaces on real devices for automated workflows.

Devin-AXIS/deepseek-design

A visual design system for DeepSeek Harness that enables AI to generate and edit real, editable design files for websites, presentations, and videos.

HKUDS/RAG-Anything

An all-in-one multimodal RAG framework that enables seamless processing and querying of documents containing interleaved text, images, tables, and mathematical equations.

Open-LLM-VTuber/Open-LLM-VTuber

An open-source, voice-interactive AI companion featuring a Live2D avatar and visual perception, capable of running entirely offline for private, real-time conversations.

jaredrhod/barehands

A webcam-based hand-tracking interface that lets you manipulate notes, images, and 3D models spatially, designed to be wired into an AI agent as a visual and interactive body.

ningzimu/image-to-editable-ppt-skill

A multi-agent AI skill that converts images, PDFs, and image-based PPTs into editable PowerPoint presentations by reconstructing text, shapes, and visual assets.

NomaDamas/CozyClay

A browser-based 3D staging studio for blocking scenes and sequencing AI-generated motion, featuring direct AI control via the Model Context Protocol.

duolahypercho/codex-router

Codex Router is a community tool that lets the Codex desktop/CLI app use many external LLM providers (Anthropic, Kimi, DeepSeek, xAI/Grok, Gemini, Copilot, Ollama Cloud, etc.). It installs a background service, an optional Electron control‑center (tray/menu‑bar app with a macOS widget), and CLI helpers that securely store API keys or OAuth tokens. After a guided setup you can pick a “routed model” in Codex, and the router forwards requests to the chosen provider, optionally adding a Perplexity search side‑car. Installation is via a one‑click script (macOS/Linux/Windows) or a Homebrew formula (CLI‑only).

OpenMOSS/MOSS-VL

MOSS‑VL is an open‑weight 11 B‑parameter family of video‑language models that can understand and converse about live video streams in real time. It uses a cross‑attention architecture with timestamped frames and a 3‑D rotary embedding (XRoPE) to achieve low‑latency, interruptible dialogue, proactive silence, and dynamic correction. Three variants are released (Realtime, Instruct, Base), with quantized checkpoints for single‑GPU use, integration with FlashAttention‑3, SGLang, ms‑swift, and LoRA fine‑tuning. The project includes benchmarks showing state‑of‑the‑art streaming performance and provides scripts for both realtime and offline inference.