EvoLinkAI/GPT-Image-2-Seedance-2.5-Workflow

A curated collection of workflows and prompt templates for combining GPT Image 2 and Seedance to create high-quality AI videos from storyboards.

zhouwei713/seedance-prompt

A prompting framework and skill for AI video generators that replaces generic descriptors with detailed simulations of real-world recording equipment and human imperfections to eliminate the "AI look."

geekjourneyx/hyperframes-motion-director

An Agent Skill that transforms text content into structured motion-video productions, featuring a professional two-phase workflow for planning and asset generation.

nvidia-cosmos/cosmos-predict2.5

NVIDIA Cosmos‑Predict 2.5 is an open‑source diffusion model that turns text, images or video into future‑state video predictions. It supports 2 B/14 B checkpoints, robot‑ and autonomous‑vehicle‑specific variants, LoRA fine‑tuning, distillation, and integrates with Hugging Face Diffusers. Designed for physical AI (robotics, AV simulation, video analytics), it offers detailed docs, a cookbook of post‑training recipes, and is licensed under Apache 2.0 (code) and NVIDIA Open Model License (weights).

microsoft/mlvc

A neural video codec that provides real-time performance and high compression efficiency across Apple, Intel, and Qualcomm NPUs.

madebyollin/taehv

TAEHV is a tiny auto‑encoder that speeds up and memory‑optimizes video VAE decoding for many diffusion models (Hunyuan, MiniMax‑H3, Wan‑Video, CogVideoX, Open‑Sora, LTX‑2, etc.). It reduces decoding of a 61‑frame 512×320 clip from ~2‑3 s / 6‑9 GB to ~0.5 s / <0.5 GB, with a small quality trade‑off. The repo ships model checkpoints, Python API (`TAEHV`, `StreamingTAEHV`), example notebooks, and integration notes for ComfyUI, stable‑diffusion.cpp, SDNext, and Diffusers.

PKU-YuanGroup/Helios

Helios is a 14B parameter real-time long video generation model capable of producing minute-scale, high-quality videos at up to 19.5 FPS on a single H100 GPU.

microsoft/DCVC

An ultra-fast neural video codec that uses a chunk-based coding framework to achieve high compression efficiency and rapid encoding/decoding speeds.

wuwukaka/ComfyUI-WanAnimatePlus

An extension for ComfyUI's WanVideo pipeline that enables seamless video connection and multi-reference image injection for improved consistency in AI video generation.

google-deepmind/physics-IQ-benchmark

A high-quality benchmark dataset and workflow for evaluating the physical understanding and realism of generative video models using real-world footage.

EvoLinkAI/awesome-seedance-2.5-guide

An official guide and showcase for Seedance 2.5, providing detailed prompts and API examples for high-end AI video generation.

NVlabs/AnyFlow

AnyFlow is an any-step video diffusion framework that allows a single model to adapt to arbitrary inference budgets for high-quality video generation.

facebookresearch/mae_st

A PyTorch implementation of Masked Autoencoders for spatiotemporal learning, enabling models to learn video representations by reconstructing masked video frames.

MCG-NJU/VideoChat3

VideoChat3 is a 4B generalist video MLLM designed for efficient video understanding, covering everything from subtle motion and long-form reasoning to live stream proactive responses.

Eyeline-Labs/Vista4D

Vista4D is a video reshooting framework that uses 4D point clouds to synthesize dynamic scenes from source videos using novel camera trajectories and viewpoints.

microsoft/World-R1

World-R1 is a reinforcement learning framework that improves 3D geometric consistency in text-to-video generation by aligning generated videos with 3D constraints without changing the base model architecture.

SandAI-org/MAGI-1

MAGI-1 is an autoregressive video generation model that produces high-fidelity videos chunk-by-chunk to ensure temporal consistency and physical accuracy.

linzzzzzz/openclip

An automated AI pipeline that extracts engaging highlights from long videos and livestreams, providing transcription, AI analysis, and automatic clip generation with subtitles and covers.

marcelo-earth/generative-manim

Generative Manim is an open‑source tool that uses LLMs (GPT‑4o, Claude, Gemini, etc.) to turn plain‑English prompts into runnable Manim animation code and rendered videos. It provides a web demo, a REST API, and a desktop app (Animo) for local rendering, plus a training pipeline and benchmark suite for fine‑tuning open‑weight models to generate Manim code.

chenfengxu714/StreamDiffusionV2

An interactive diffusion pipeline for real-time streaming video generation that optimizes throughput and latency across single and multi-GPU setups.

louisedesadeleer/clipify

A Claude Code skill that automatically turns long-form talking-head videos into short, social-ready clips with automated reframing and subtitles.

yaoyao-jpg/PhiZero

PhiZero is a world model that predicts future physical dynamics by reasoning in a discrete Physical Language before rendering the results as video.

Kosinkadink/ComfyUI-AnimateDiff-Evolved

An advanced AnimateDiff integration for ComfyUI that enables the creation of high-quality AI animations with infinite length and extensive control over motion and sampling.

WeChatCV/Stand-In

A lightweight, plug-and-play framework for identity-preserving video generation that maintains high subject consistency by training only 1% of additional parameters.

NevermindNilas/TheAnimeScripter

An AI-powered video enhancement toolkit specialized for anime that provides upscaling, motion interpolation, and restoration in a single processing pass.

Lixsp11/sekai-codebase

Sekai is a high-quality egocentric video dataset featuring over 5,000 hours of annotated first-person and drone footage for world exploration and generation.

AMAP-ML/DreamX-World

DreamX-World is a general-purpose interactive world model that generates high-fidelity, controllable simulations allowing users to explore and transform environments via event prompts.

zju3dv/street_crafter

StreetCrafter is a framework for street view synthesis that uses controllable video diffusion models and 3D Gaussian Splatting to generate realistic videos from novel trajectories.

Vchitect/Latte

A Latent Diffusion Transformer for high-quality video generation, supporting both text-to-video and text-to-image tasks.

Orkas-AI/Orkas-VideoStudio

A toolkit that enables AI coding agents to compose, edit, and generate videos via a readable plan file, providing a deterministic pipeline for automated video production.

thu-ml/DiT-Extrapolation

A plug-and-play framework for Diffusion Transformers that enables the generation of longer videos and higher-resolution images through positional embedding extrapolation.

barefootford/buttercut

An AI video editing agent that generates XML cuts for professional software like Final Cut Pro, Premiere, and DaVinci Resolve to automate the initial assembly of footage.

alibaba/Tora

Tora is a trajectory-oriented Diffusion Transformer for video generation that allows precise control of object motion using textual, visual, and trajectory-based guidance.

Kevin-thu/StoryMem

StoryMem is a multi-shot video generation framework that uses a memory-conditioned diffusion model to create long, coherent narrative videos with consistent characters.

wernerturing/multi-delogo

An application for removing logos and watermarks from videos, featuring automatic detection and the ability to handle moving logos.

ossrs/oryx

Oryx is an all-in-one open-source video solution for creating live streaming and WebRTC services with integrated AI capabilities for transcription and dubbing.

Agions/splicr

A Rust-based AI video narration engine that uses a multi-agent system to automate scene analysis, scriptwriting, voice cloning, and timeline alignment for CapCut exports.

IgorShadurin/app.yumcut.com

An open-source AI short video generator that automates the creation of scripts, voiceovers, visuals, and captions for TikTok, YouTube Shorts, and Instagram Reels.

gulucaptain/Camera-Transformer-1

CT-1 is a Vision-Language-Camera model that estimates precise camera trajectories from images and text prompts to enable spatially aware, controllable video generation.

showlab/Kiwi-Edit

Kiwi‑Edit is an open‑source video‑editing system that takes natural‑language instructions (and optional reference images) to modify whole videos. It fuses a multi‑modal LLM encoder with a video Diffusion Transformer, offering style transfer, object add/replace/remove, and background changes. The repo provides full training scripts (three‑stage curriculum with Qwen2.5‑VL‑3B + Wan2.2‑TI2V‑5B), pretrained Diffusers checkpoints, and evaluation pipelines on OpenVE‑Bench and RefVIE‑Bench.

Anil-matcha/Seedance-2-API

A Python wrapper for ByteDance's Seedance AI video generator, enabling high-fidelity text-to-video and image-to-video generation with a focus on realistic human faces and character consistency.

simchowitzlabpublic/nano-world-model

A minimalist framework for training video world models based on diffusion-forcing, used for future video prediction, 3D reconstruction, and robotic planning.

microsoft/LatentSpatialMemory

A framework for video world models that stores 3D scene content as latent tokens to improve generation speed and reduce memory overhead by avoiding repeated RGB rendering.

TencentARC/VerseCrafter

VerseCrafter is a controllable video world model that uses 4D geometric control and a GeoAdapter to provide precise, interpretable control over camera and multi-object motion in realistic videos.

nvidia-cosmos/cosmos-predict1

A suite of world foundation models for future state prediction that generates visual simulations from text or video prompts to support Physical AI development.

Tencent-Hunyuan/HY-WorldPlay

HY-World 1.5 is a real-time interactive world modeling framework that uses streaming video diffusion to generate geometrically consistent 3D environments based on user input.

Tencent-Hunyuan/HunyuanVideo-1.5

A lightweight 8.3B parameter video generation model that enables high-quality text-to-video and image-to-video synthesis on consumer-grade GPUs.

JingyunLiang/VRT

VRT is a PyTorch implementation of the Video Restoration Transformer, a transformer‑based model that handles video super‑resolution, deblurring, denoising, frame interpolation and space‑time SR. It uses Temporal Mutual Self‑Attention and parallel warping to capture long‑range temporal dependencies, delivering state‑of‑the‑art PSNR gains on nine benchmarks. The repo includes pretrained weights, test scripts, a Colab demo, and instructions for training on standard video datasets.