EvoLinkAI/GPT-Image-2-Seedance-2.5-Workflow
A curated collection of workflows and prompt templates for combining GPT Image 2 and Seedance to create high-quality AI videos from storyboards.
zhouwei713/seedance-prompt
A prompting framework and skill for AI video generators that replaces generic descriptors with detailed simulations of real-world recording equipment and human imperfections to eliminate the "AI look."
geekjourneyx/hyperframes-motion-director
An Agent Skill that transforms text content into structured motion-video productions, featuring a professional two-phase workflow for planning and asset generation.
nvidia-cosmos/cosmos-predict2.5
NVIDIA Cosmos‑Predict 2.5 is an open‑source diffusion model that turns text, images or video into future‑state video predictions. It supports 2 B/14 B checkpoints, robot‑ and autonomous‑vehicle‑specific variants, LoRA fine‑tuning, distillation, and integrates with Hugging Face Diffusers. Designed for physical AI (robotics, AV simulation, video analytics), it offers detailed docs, a cookbook of post‑training recipes, and is licensed under Apache 2.0 (code) and NVIDIA Open Model License (weights).
microsoft/mlvc
A neural video codec that provides real-time performance and high compression efficiency across Apple, Intel, and Qualcomm NPUs.
madebyollin/taehv
TAEHV is a tiny auto‑encoder that speeds up and memory‑optimizes video VAE decoding for many diffusion models (Hunyuan, MiniMax‑H3, Wan‑Video, CogVideoX, Open‑Sora, LTX‑2, etc.). It reduces decoding of a 61‑frame 512×320 clip from ~2‑3 s / 6‑9 GB to ~0.5 s / <0.5 GB, with a small quality trade‑off. The repo ships model checkpoints, Python API (`TAEHV`, `StreamingTAEHV`), example notebooks, and integration notes for ComfyUI, stable‑diffusion.cpp, SDNext, and Diffusers.
PKU-YuanGroup/Helios
Helios is a 14B parameter real-time long video generation model capable of producing minute-scale, high-quality videos at up to 19.5 FPS on a single H100 GPU.
microsoft/DCVC
An ultra-fast neural video codec that uses a chunk-based coding framework to achieve high compression efficiency and rapid encoding/decoding speeds.
wuwukaka/ComfyUI-WanAnimatePlus
An extension for ComfyUI's WanVideo pipeline that enables seamless video connection and multi-reference image injection for improved consistency in AI video generation.
google-deepmind/physics-IQ-benchmark
A high-quality benchmark dataset and workflow for evaluating the physical understanding and realism of generative video models using real-world footage.
EvoLinkAI/awesome-seedance-2.5-guide
An official guide and showcase for Seedance 2.5, providing detailed prompts and API examples for high-end AI video generation.
NVlabs/AnyFlow
AnyFlow is an any-step video diffusion framework that allows a single model to adapt to arbitrary inference budgets for high-quality video generation.
facebookresearch/mae_st
A PyTorch implementation of Masked Autoencoders for spatiotemporal learning, enabling models to learn video representations by reconstructing masked video frames.
MCG-NJU/VideoChat3
VideoChat3 is a 4B generalist video MLLM designed for efficient video understanding, covering everything from subtle motion and long-form reasoning to live stream proactive responses.
Eyeline-Labs/Vista4D
Vista4D is a video reshooting framework that uses 4D point clouds to synthesize dynamic scenes from source videos using novel camera trajectories and viewpoints.
microsoft/World-R1
World-R1 is a reinforcement learning framework that improves 3D geometric consistency in text-to-video generation by aligning generated videos with 3D constraints without changing the base model architecture.
SandAI-org/MAGI-1
MAGI-1 is an autoregressive video generation model that produces high-fidelity videos chunk-by-chunk to ensure temporal consistency and physical accuracy.
linzzzzzz/openclip
An automated AI pipeline that extracts engaging highlights from long videos and livestreams, providing transcription, AI analysis, and automatic clip generation with subtitles and covers.
marcelo-earth/generative-manim
Generative Manim is an open‑source tool that uses LLMs (GPT‑4o, Claude, Gemini, etc.) to turn plain‑English prompts into runnable Manim animation code and rendered videos. It provides a web demo, a REST API, and a desktop app (Animo) for local rendering, plus a training pipeline and benchmark suite for fine‑tuning open‑weight models to generate Manim code.
chenfengxu714/StreamDiffusionV2
An interactive diffusion pipeline for real-time streaming video generation that optimizes throughput and latency across single and multi-GPU setups.
louisedesadeleer/clipify
A Claude Code skill that automatically turns long-form talking-head videos into short, social-ready clips with automated reframing and subtitles.
yaoyao-jpg/PhiZero
PhiZero is a world model that predicts future physical dynamics by reasoning in a discrete Physical Language before rendering the results as video.
Kosinkadink/ComfyUI-AnimateDiff-Evolved
An advanced AnimateDiff integration for ComfyUI that enables the creation of high-quality AI animations with infinite length and extensive control over motion and sampling.
WeChatCV/Stand-In
A lightweight, plug-and-play framework for identity-preserving video generation that maintains high subject consistency by training only 1% of additional parameters.
NevermindNilas/TheAnimeScripter
An AI-powered video enhancement toolkit specialized for anime that provides upscaling, motion interpolation, and restoration in a single processing pass.
Lixsp11/sekai-codebase
Sekai is a high-quality egocentric video dataset featuring over 5,000 hours of annotated first-person and drone footage for world exploration and generation.
AMAP-ML/DreamX-World
DreamX-World is a general-purpose interactive world model that generates high-fidelity, controllable simulations allowing users to explore and transform environments via event prompts.
zju3dv/street_crafter
StreetCrafter is a framework for street view synthesis that uses controllable video diffusion models and 3D Gaussian Splatting to generate realistic videos from novel trajectories.
Vchitect/Latte
A Latent Diffusion Transformer for high-quality video generation, supporting both text-to-video and text-to-image tasks.
Orkas-AI/Orkas-VideoStudio
A toolkit that enables AI coding agents to compose, edit, and generate videos via a readable plan file, providing a deterministic pipeline for automated video production.
thu-ml/DiT-Extrapolation
A plug-and-play framework for Diffusion Transformers that enables the generation of longer videos and higher-resolution images through positional embedding extrapolation.
barefootford/buttercut
An AI video editing agent that generates XML cuts for professional software like Final Cut Pro, Premiere, and DaVinci Resolve to automate the initial assembly of footage.
alibaba/Tora
Tora is a trajectory-oriented Diffusion Transformer for video generation that allows precise control of object motion using textual, visual, and trajectory-based guidance.
Kevin-thu/StoryMem
StoryMem is a multi-shot video generation framework that uses a memory-conditioned diffusion model to create long, coherent narrative videos with consistent characters.
wernerturing/multi-delogo
An application for removing logos and watermarks from videos, featuring automatic detection and the ability to handle moving logos.
ossrs/oryx
Oryx is an all-in-one open-source video solution for creating live streaming and WebRTC services with integrated AI capabilities for transcription and dubbing.
Agions/splicr
A Rust-based AI video narration engine that uses a multi-agent system to automate scene analysis, scriptwriting, voice cloning, and timeline alignment for CapCut exports.
IgorShadurin/app.yumcut.com
An open-source AI short video generator that automates the creation of scripts, voiceovers, visuals, and captions for TikTok, YouTube Shorts, and Instagram Reels.
gulucaptain/Camera-Transformer-1
CT-1 is a Vision-Language-Camera model that estimates precise camera trajectories from images and text prompts to enable spatially aware, controllable video generation.
showlab/Kiwi-Edit
Kiwi‑Edit is an open‑source video‑editing system that takes natural‑language instructions (and optional reference images) to modify whole videos. It fuses a multi‑modal LLM encoder with a video Diffusion Transformer, offering style transfer, object add/replace/remove, and background changes. The repo provides full training scripts (three‑stage curriculum with Qwen2.5‑VL‑3B + Wan2.2‑TI2V‑5B), pretrained Diffusers checkpoints, and evaluation pipelines on OpenVE‑Bench and RefVIE‑Bench.
Anil-matcha/Seedance-2-API
A Python wrapper for ByteDance's Seedance AI video generator, enabling high-fidelity text-to-video and image-to-video generation with a focus on realistic human faces and character consistency.
simchowitzlabpublic/nano-world-model
A minimalist framework for training video world models based on diffusion-forcing, used for future video prediction, 3D reconstruction, and robotic planning.
microsoft/LatentSpatialMemory
A framework for video world models that stores 3D scene content as latent tokens to improve generation speed and reduce memory overhead by avoiding repeated RGB rendering.
TencentARC/VerseCrafter
VerseCrafter is a controllable video world model that uses 4D geometric control and a GeoAdapter to provide precise, interpretable control over camera and multi-object motion in realistic videos.
nvidia-cosmos/cosmos-predict1
A suite of world foundation models for future state prediction that generates visual simulations from text or video prompts to support Physical AI development.
Tencent-Hunyuan/HY-WorldPlay
HY-World 1.5 is a real-time interactive world modeling framework that uses streaming video diffusion to generate geometrically consistent 3D environments based on user input.
Tencent-Hunyuan/HunyuanVideo-1.5
A lightweight 8.3B parameter video generation model that enables high-quality text-to-video and image-to-video synthesis on consumer-grade GPUs.
JingyunLiang/VRT
VRT is a PyTorch implementation of the Video Restoration Transformer, a transformer‑based model that handles video super‑resolution, deblurring, denoising, frame interpolation and space‑time SR. It uses Temporal Mutual Self‑Attention and parallel warping to capture long‑range temporal dependencies, delivering state‑of‑the‑art PSNR gains on nine benchmarks. The repo includes pretrained weights, test scripts, a Colab demo, and instructions for training on standard video datasets.