ganbo-gab/open-storyboard-canvas

Open Storyboard Canvas is an open‑source, cross‑platform desktop app (Tauri + React + Rust) that provides a visual node canvas for AI‑generated images, videos, and 3‑D storyboard scenes. It includes a chat‑based Canvas Agent (beta) that can create, edit, and track generation tasks, supports multiple providers (OpenAI‑compatible, Claude, Dreamina CLI), and offers director‑studio tools, prompt libraries, bulk import, and detailed logs. Licensed MIT, built on the Storyboard‑Copilot upstream project.

Mr-funny/hbg-classical-poem-silk-video

An Agent Skill that converts Chinese classical poems into vertical, Chinese‑painting‑style animated videos (1080 × 1920) using AI image generation, a Gemini‑based image‑to‑video Docker container, brush‑stroke subtitles, and automated QA.

Mobile-Artificial-Intelligence/maid

Maid is an open‑source Android app (React Native) that lets you run GGUF LLMs locally via llama.cpp and connect to many remote LLM APIs. It offers one‑tap model downloads, chat management, adjustable generation settings, optional cloud sync, and a companion TTS app (Maise). Install from GitHub releases or Google Play, or build the APK yourself.

raysonmeng/agent-bridge

AgentBridge is a local tool that connects Claude Code and OpenAI Codex so they can exchange messages, split tasks, and review each other's code automatically. It runs a Bun‑based daemon and a short CLI (`abg`/`agentbridge`) that launches both agents, handles turn coordination, quota hand‑off, and filters noisy output. Designed for developers who want two LLM coding assistants to collaborate without manual copy‑pasting.

ZiYang-xie/WorldGen

WorldGen is a tool that generates immersive 3D scenes from text prompts or images in seconds, supporting both Gaussian Splatting and mesh outputs for VR, games, and simulations.

bcurts/agentchattr

agentchattr is a cross‑platform local chat server that lets you and multiple AI coding agents converse in real‑time. By @‑mentioning an agent in the web UI, the server injects a prompt into the agent’s terminal (via MCP), the agent reads the chat and replies automatically, enabling hands‑free coordination, multi‑agent workflows, jobs, rules, sessions, and a rich UI with channels, pins, decision cards, scheduling, voice typing, and more.

sorker/ai-shotlive

AI ShotLive Director is an open‑source, React + Express (or Electron) platform that turns novels or story outlines into short videos or motion comics. It uses a keyframe‑driven workflow (first/last frames → interpolation) and lets you plug in any text, image, video or audio model (OpenAI, Anthropic, Google, AntSK, etc.). Features include novel‑to‑script conversion, asset generation (character look‑books, scene concepts), a grid storyboard editor, multi‑track video editing with AI subtitles/TTS, and full backup/restore. Deployable via Docker, local SQLite, or as a desktop app, and licensed under CC BY‑NC‑SA 4.0.

ChrisChen667788/wind-comic

A multi-agent AI studio that transforms a single sentence into a finished short-form drama, automating everything from script and character design to voiceover and final MP4 export.

NVIDIA-AI-Blueprints/video-search-and-summarization

A reference architecture for building GPU-accelerated vision agents that use natural language to search, summarize, and reason over live or recorded video streams.

saddam213/AmuseAI

A local multimodal AI application for generating and editing images, video, audio, and text using various open-source pipelines and GPU backends.

aaronyi97/image-story-video-wizard

A step-by-step AI video production workflow for Codex and WorkBuddy that guides users from topic selection to final rendering of story-based videos.

Evianis/travel-photo-abstraction

A Codex skill that analyzes photographs to create sparse editorial abstractions by mapping visual evidence like shapes and spatial rhythm to minimal abstract marks.

lidge-jun/ima2-gen

A local-first visual generation studio and runtime that provides a unified interface for image and video generation across multiple AI providers with advanced iteration tools.

scraed/LanPaint

A training-free universal inpainting sampler for diffusion models that uses a "Think Mode" to improve quality across images, video, and audio.

facebookresearch/fairchem

fairchem is Meta’s open‑source library of machine‑learned interatomic potentials (the UMA models) for chemistry and materials science. It provides a pip‑installable `fairchem-core` package, ASE‑compatible calculators, multi‑GPU inference, and pretrained models for molecules, polymers, inorganic crystals, catalysts, MOFs, etc., enabling fast, quantum‑accurate simulations.

songlab-cal/gpn

A suite of genomic language models designed to predict the functional effects of genome-wide variants using both aligned and unaligned genomic sequences.

corvo007/MioSub

An AI-powered subtitle editor that automates transcription, translation, and millisecond-precise timeline alignment using Gemini and Whisper.

anliyuan/Ultralight-Digital-Human

A lightweight digital human model designed for real-time talking-head generation on mobile devices using personalized training from short videos.

NVIDIA/cudnn-frontend

NVIDIA’s cuDNN Frontend is an open‑source, header‑only C++ API plus Python bindings that simplify building high‑performance deep‑learning kernels (attention, GEMM, fused ops) on Hopper and Blackwell GPUs. It ships ready‑made, JIT‑compiled kernels for Flash‑Attention, Mixture‑of‑Experts GEMM fusions, sparse/linear attention, and more, all accessible via a unified Graph API and PyTorch‑compatible ops.

omnichar/OmniChar

An open-source studio for creating consistent AI characters using a portable .char format that works across multiple image and video generation models.

kunchenguid/autopreso

A hands-free presentation tool that uses an AI agent to draw and rearrange an Excalidraw whiteboard in real time based on the speaker's voice.

safetensors/safetensors

safetensors is a Rust‑backed Python library that defines a tiny, JSON‑headered binary format for storing tensors safely (no pickle code execution) while enabling zero‑copy, lazy loading and support for modern dtypes. It’s used for distributing large model weights on Hugging Face and can be installed via `pip install safetensors`.

zhizinan1997/jimeng-free-api-all

An OpenAI-compatible API service for Jimeng AI that enables programmatic image and video generation with multi-account rotation and a management dashboard.

kyleskom/NBA-Machine-Learning-Sports-Betting

A Python project that scrapes NBA stats and sportsbook odds, builds matchup features, trains XGBoost and TensorFlow models to predict game winners and over/under totals, and outputs expected value and Kelly‑criterion stake sizes. Includes scripts for data collection, model training, a command‑line predictor, and a simple Flask web UI.

NVIDIA/cosmos-framework

An end-to-end framework for training and serving omnimodal world models, including the Cosmos 3 family, designed to unify language, vision, audio, and action for Physical AI.

xmarre/ComfyUI-Spectrum-MiniMax-H3

A ComfyUI implementation of Spectrum that accelerates MiniMax H3 audio-video generation by forecasting hidden states to skip expensive transformer evaluations.

openxla/xprof

XProf is an open‑source, scalable profiler (with a TensorBoard plugin) for modern machine‑learning workloads. It visualises step‑time breakdowns, operation timelines, memory usage and HLO graphs, supports distributed processing, and can be installed via `pip install xprof`.

EvolvingLMMs-Lab/LLaVA-OneVision-2

A fully open 8B multimodal model that unifies image, long-form video, and spatial understanding using a codec-aligned vision encoder for efficient long-video reasoning.

facebookresearch/meshflow

MeshFlow is an efficient 3D mesh generation system that uses a MeshVAE and a flow-matching Diffusion Transformer to create artistic meshes in about one second.

AlayaLab/WildWorld

A large-scale action-conditioned world modeling dataset with explicit state annotations from a AAA ARPG, designed for training generative game environments.

dineshsoudagar/local-llms-on-android

An Android application that enables fully offline, private LLM chat with support for voice, image, and camera input using ONNX and LiteRT backends.

jedzqer/manga-translator-android

An Android app for manga translation that uses local bubble detection and OCR combined with LLM-based translation to overlay translated text on original images.

OneMoreGres/ScreenTranslator

A screen capture and OCR-based tool that translates text appearing on the screen using online translation services.

GalTransl/GalTransl

GalTransl is a desktop GUI that automates extracting, translating (via GPT‑4/Claude/DeepSeek/etc.), and reinjecting Japanese galgame scripts into Chinese, with GPT‑driven dictionaries, caching, and subtitle support.

Wing900/ManimCat

An AI-powered workspace for creating mathematical animations and static visuals using natural language, powered by Manim and matplotlib.

mll-lab-nu/VAGEN

VAGEN is a reinforcement learning framework for training multi-turn VLM agents that improves performance by reinforcing their internal world model through explicit visual state reasoning.

tmustier/pi-for-excel

Pi for Excel is an open‑source Office add‑in that embeds a conversational AI agent inside Microsoft Excel. The agent can read, write, format, and analyse spreadsheets using any LLM you configure, offering 16 built‑in spreadsheet tools, session management, automatic checkpoints for one‑click rollback, and extensible sandboxed extensions.

google-deepmind/tips

TIPS is a series of foundational image-text encoders designed for general-purpose computer vision and multimodal AI, featuring enhanced spatial awareness and patch-text alignment.

xikhar/persona

A cross-platform desktop application that provides a real-time, expressive 3D character avatar that reacts to voice output from AI assistants.

vinhhien112/img2obj

A Codex plugin that transforms reference images into animation-ready procedural Three.js objects by generating the necessary construction code through a structured analysis and review workflow.

RareSense/Nova3D

Nova3D is a code-native 3D generation system that creates assets as executable Blender Python scripts, enabling the production of structured, editable, and animatable 3D models with named parts.

OrRon/EpicInfographics

A skill for AI agents that enables the creation of professional, studio-quality infographics and animations by replacing generic templates with scene-based design and mathematical accuracy.

gradio-app/trackio

Trackio is a free, lightweight experiment‑tracking library (compatible with wandb) that stores logs locally in SQLite, offers a fast non‑blocking API, a Gradio‑style dashboard, CLI/SQL query tools, alerts, and optional free hosting on Hugging Face Spaces or self‑hosting.

nodetool-ai/nodetool

An agent-first creative workspace for producing images, video, audio, and text using a node-based graph system and integrated AI agents.

ATH-MaaS/Ovis

Ovis is a Multimodal Large Language Model (MLLM) architecture that structurally aligns visual and textual embeddings to enable high-resolution perception and reflective reasoning.

morsoli/aimangastudio

An AI-powered comic creation tool that provides an end-to-end pipeline for script generation, storyboarding, and character style control.

jeremychone/rust-genai

genai is a Rust crate that provides a single, ergonomic async API for chatting with over 200 LLM models from 26+ providers (OpenAI, Anthropic, Gemini, Ollama, Bedrock, etc.). It selects the correct native protocol automatically, supports custom OpenAI‑compatible endpoints, streaming, tool calling, image analysis, and offers extensive configuration via ChatOptions. The library is actively maintained (v0.6 released, v0.7 beta upcoming) and includes many examples and docs.

X-GenGroup/Flow-Factory

A unified framework for online RL and offline fine-tuning of diffusion and flow-matching models across image, video, and audio-video modalities.