ganbo-gab/open-storyboard-canvas

Open Storyboard Canvas is an open‑source, cross‑platform desktop app (Tauri + React + Rust) that provides a visual node canvas for AI‑generated images, videos, and 3‑D storyboard scenes. It includes a chat‑based Canvas Agent (beta) that can create, edit, and track generation tasks, supports multiple providers (OpenAI‑compatible, Claude, Dreamina CLI), and offers director‑studio tools, prompt libraries, bulk import, and detailed logs. Licensed MIT, built on the Storyboard‑Copilot upstream project.

xszyou/Fay

A comprehensive digital human framework that integrates LLMs, ASR, and TTS to power interactive virtual avatars for apps, websites, and embedded devices.

lycohana/BiliSum

BiliSum is a video-to-knowledge tool that converts Bilibili, YouTube, and local videos into transcriptions, visual notes, and a local RAG-powered knowledge base.

death34018-hue/AionsHome

A locally-hosted AI companion system integrating multi-modal chat, smart camera surveillance, and RAG-based memory to create a proactive personal assistant.

google-labs-code/stitch-sdk

A TypeScript SDK for generating and editing UI screens from text prompts, providing HTML and screenshots for rapid prototyping and AI agent integration.

omicverse/omicverse

OmicVerse is a Python library that unifies bulk RNA‑seq, single‑cell, and spatial transcriptomics analysis. It bundles dozens of established bio‑informatics tools, standardises on the anndata/mudata data structures, and adds AI‑assisted agents (J.A.R.V.I.S.) and an MCP server so you can run pipelines via natural‑language commands or LLM integration. Optional extensions provide synthetic‑biology modelling and GPU‑accelerated molecular‑dynamics workflows.

xandergos/terrain-diffusion

A learned successor to Perlin noise that uses diffusion models to generate infinite, deterministic, and randomly-accessible real-time terrain and climate data.

oil-oil/wolfcha

Wolfcha is a web game that lets a single human play Werewolf against a full table of AI‑controlled personalities. The AIs remember past dialogue, act with intent, and generate spoken lines (with optional voice) in real time, recreating the chaotic social deduction experience without needing other human players.

PunithVT/ai-avatar-system

A production-ready platform for building photorealistic AI avatars with real-time lip-sync, zero-shot voice cloning, and multi-LLM support.

LC044/TrailSnap

An AI-powered self-hosted photo album that automatically organizes travel memories through ticket OCR, face recognition, and geographic footprint mapping.

ashuoAI/SHUO-Canvas

SHUO Canvas is a desktop, node‑based canvas that lets creators combine text, images, video and audio with AI generation (via OpenAI‑compatible, RunningHub, ComfyUI, etc.) to build scripts, storyboards, character‑swap videos, 360° panoramas, voice‑overs and 3‑D scenes—all within one visual workflow.

MashiroSaber03/Saber-Translator

An AI-powered manga translation and management suite that automates text detection, OCR, translation, and image inpainting, featuring a RAG-based story analysis engine.

jipraks/yt-short-clipper

A Windows desktop app that uses LLMs to identify highlights from YouTube videos and automatically converts them into captioned, 9:16 portrait clips.

vanloctech/youwee

An open-source GUI for yt-dlp that provides video downloading from 1,800+ sites with integrated AI video summaries, natural language editing, and AI-powered subtitle generation.

openstory-so/openstory

An AI-powered video production platform that transforms scripts into styled video sequences with consistent characters and visual styles.

Softeria/ms-365-mcp-server

A Node.js server that exposes Microsoft 365 Graph API as MCP tools for LLMs, supporting personal and org accounts, multi‑account use, permission filtering, and an experimental token‑saving TOON output format.

google-deepmind/alphagenome_research

AlphaGenome is a unified DNA sequence model that predicts the effects of regulatory genetic variants on gene expression, splicing, and chromatin features at single base-pair resolution.

facebookresearch/eb_jepa

A lightweight library for Energy-Based Joint-Embedding Predictive Architectures used to learn representations for prediction and planning across images and video.

huangjunsen0406/py-xiaozhi

py‑xiaozhi is a Python‑based, async‑first framework for building multimodal AI assistants that can listen, speak, see, and control hardware. It offers real‑time Opus voice streaming, offline wake‑word detection, vision‑language integration, a JSON‑RPC tool server, and secure WebSocket/MQTT communication. Runs on Windows/macOS/Linux desktops and on edge boards (Raspberry Pi, Jetson Nano). Provides GUI (PySide6/QML), CLI, and GPIO modes, and a plugin system for easy extension.

NVlabs/FastGen

A PyTorch-based framework for accelerating diffusion models using various distillation and acceleration techniques for image and video generation.

ogkalu2/comic-translate

An automated comic translator that uses LLMs, OCR, and inpainting to translate comics and manga from multiple languages while preserving the original visual layout.

microsoft/skills-for-copilot-studio

Skills for Copilot Studio is a plugin (for Claude Code, GitHub Copilot CLI, and VS Code) that lets you create, edit, sync, test, and get design advice on Microsoft Copilot Studio “standard” agents using YAML files—all from your terminal or editor.

NVIDIA/TransformerEngine

NVIDIA Transformer Engine is a GPU‑focused library that provides FP8/MXFP8/NVFP4 mixed‑precision support, fused kernels and a simple autocast API for PyTorch and JAX, enabling faster, lower‑memory training and inference of large language and MoE models.

Alos21750/UAV-Downloader

UAV Downloader is a Windows‑focused video downloader for JableTV, MissAV, SupJav and Hanime1 that can automatically generate Japanese, English and Traditional Chinese subtitles using local speech‑recognition (ReazonSpeech K2 v2 or Whisper) and translation models (FuguMT, OPUS‑MT). It offers three operation modes – an interactive GUI (UAV Browser), an unattended scheduler (UAV Watcher), and a headless Docker/CLI version – with proxy support, parallel downloads, and optional LLM‑based translation. The project is released under Apache 2.0 and runs offline by default.

tianjiangqiji/nova-image-studio

A self-hosted AI image and video generation workbench that allows users to bring their own models and extend functionality via a plugin system.

wassermanproductions/motion-previs-studio

A desktop app for filmmakers to extract pose, depth, and camera movement from reference videos to create control bundles for AI video generation pipelines.

FlowElement-xinliuyuansu/m_flow

M-flow is a memory system for AI agents that retrieves information by retrieving information through scoring evidence paths in a knowledge graph, operating like a cognitive memory system rather than relying solely on similarity search.

visualbruno/ComfyUI-Trellis2

ComfyUI‑Trellis2 is a custom‑node pack that integrates Microsoft’s TRELLIS.2 3‑D diffusion model into the visual programming UI ComfyUI. It provides dozens of nodes for generating, refining, texturing, and rendering meshes from single or multi‑view images (or video). Installation requires Windows, Python 3.11, PyTorch 2.7/2.8 with CUDA, and the Facebook DINOv3 checkpoint. The repo ships pre‑built wheels for native extensions and includes example workflows. It’s an active, genuine AI/ML project aimed at 3‑D generation.

Asad-Ismail/MoneyPrinterTurbo-Extended

An automated video generation tool that creates scripts, synthesizes voiceovers with local voice cloning, and matches relevant video clips with synchronized word-level subtitles.

opentokenz/mcpx

MCPX is a Go‑based local server that implements the MCP protocol, enabling LLM agents (ChatGPT, Claude, etc.) to securely read, edit, execute, and manage files, commands, plans, and artefacts in registered local workspaces. It provides persistent Remote Sessions, fine‑grained security policies, audit logs stored in SQLite, and extensibility via Skills and custom tools.

jangles-byte/Pythia

A local, swarm-intelligence global forecasting system that fuses over 40 real-time data feeds into a 3D intelligence globe to predict world events.

NVIDIA-BioNeMo/Proteina-Complexa

A generative model for atomistic protein binder design that unifies generative modeling and hallucination to create binders for proteins and small-molecule ligands.

aurekaresearch/OpenDDE

An all-atom biomolecular foundation model that uses co-folding for structure prediction, design, and optimization in drug discovery.

OpenBMB/MiniCPM-o-Demo

A demo system for MiniCPM-o 4.5 that enables real-time, full-duplex omnimodal interaction, allowing the AI to see, listen, and speak simultaneously.

mynaparrot/plugNmeet-server

A scalable, open-source web conferencing system built on LiveKit that provides self-hosted video meetings with AI-powered transcription and summaries.

bytedance/SALMONN

A suite of advanced multi-modal LLMs that provide large language models with generic hearing abilities and integrated audio-visual perception.

mcp-router/mcp-router

MCP Router is a discontinued, open‑source desktop app (Windows/macOS) for locally managing Model Context Protocol (MCP) servers. It lets users group servers into projects and workspaces, toggle individual tools, view request logs and basic analytics, and integrate with popular AI front‑ends. All data stays on the user’s machine; the last released version is v0.6.4 (end‑of‑support Sept 2026).

NVIDIA-NeMo/Megatron-Bridge

NeMo Megatron Bridge is a PyTorch‑native library that connects Hugging Face model checkpoints with NVIDIA’s Megatron‑Core training engine. It offers bidirectional checkpoint conversion, verification, and high‑throughput training (tensor/pipeline parallelism, FP8/BF16/FP4). The repo ships ready‑made recipes for dozens of LLM, VLM, and diffusion models (e.g., Llama 3.2, Nemotron‑3 Ultra, K‑EXAONE‑2) and supports SFT/LoRA fine‑tuning. Install via the official NeMo Docker container or from source, and use the provided quick‑start scripts to convert, train, or export models.

tl2012tl/TE_MAN

TE MAN is a ComfyUI plugin that adds an infinite‑canvas workflow, asset library, audio/video/image nodes, 3‑D director, batch prompt tools, and an AI‑assistant chat panel, enabling both local model inference and remote API usage for image, video, and multimodal generation.

nagadomi/nunif

A collection of PyTorch-based tools for image super-resolution, 2D-to-3D video conversion for VR, and image quality assessment for dataset filtering.

ysr666/dsh-vision-router

A vision routing plugin for DeepSeek Harness that enables agents to interact with image pixels via tool calls and provides a built-in keyless free vision fallback.

tracel-ai/cubecl

CubeCL is a Rust procedural‑macro‑based language that lets you write a single compute kernel in Rust and JIT‑compile it to CUDA, HIP, Metal, WebGPU (WGSL/SPIR‑V) or CPU SIMD. It models parallelism with four axes (vector, plane, cube dimension, cube count), supports compile‑time specialization, automatic vectorization, and runtime autotuning, and forms the low‑level foundation for the Rust deep‑learning framework Burn and the kernel library cubek.

nicolasvonluetzow/GaussianGPT

GaussianGPT is a transformer-based model that generates 3D Gaussian scenes via next-token prediction, supporting controllable 3D scene generation, completion, and outpainting.

MiniMax-AI/MiniMax-MCP

An official Model Context Protocol (MCP) server that enables AI clients like Claude and Cursor to generate speech, clone voices, and create images and videos using MiniMax APIs.

InternLM/Intern-S1

A series of multimodal foundation models optimized for scientific intelligence and long-horizon agents, capable of reasoning across complex scientific domains like chemistry and biology.

ZeroTang05/cyber-doctor

A multimodal AI health assistant that provides disease diagnosis and medical record analysis using RAG, knowledge graphs, and voice interaction.

ACEsuit/mace

MACE is an open‑source PyTorch library for equivariant graph‑neural‑network interatomic potentials. It provides fast training/evaluation, multi‑GPU and Apple‑Silicon support, optional CUDA acceleration, and a catalog of pre‑trained foundation models for materials and organic chemistry. Install via `pip install mace-torch`, train with the `mace_run_train` CLI (or YAML configs), and evaluate with `mace_eval_configs`. Documentation, Colab tutorials, and Weights & Biases integration are included. The project is stable on the v0.3 line while a major v1.0 rewrite is in progress.

danilo-znamerovszkij/draw-your-font

An open-source tool that turns photos of handwriting into installable TTF/WOFF fonts using local processing and optional AI vision for character labeling.