MaximeVandegar/Papers-in-100-Lines-of-Code
A curated set of ~64 ultra‑compact (≈100‑line) implementations of influential machine‑learning papers, spanning from early deep‑learning milestones to the latest neural rendering and diffusion models. Designed as a learning and prototyping resource, the repo provides minimal, runnable code for each paper under an MIT license.
leigest519/ScreenCoder
ScreenCoder is a multi-agent system that transforms UI screenshots and design mockups into production-ready HTML/CSS code using visual understanding and layout planning.
microsoft/bioemu
A generative model that samples the equilibrium distribution of protein monomer structures from amino acid sequences, featuring a steering system to ensure physical plausibility.
OpenSenseNova/SenseNova-Vision
SenseNova-Vision is a unified multimodal model that treats diverse computer vision tasks as a generation problem, enabling a single model to perform detection, segmentation, and geometric prediction via text and image outputs.
dc-ai-projects/DC-Gen
An acceleration framework for diffusion models that transfers pre-trained models into a deeply compressed latent space to achieve up to 53.8x faster inference without losing quality.
LYiHub/pub-local-jarvis
A local multimodal desktop assistant for Windows that perceives screen and system audio to provide real-time gaming companionship, course notes, and contextual AI interaction.
IamCreateAI/NeoVerse
NeoVerse is a 4D world model that reconstructs 3D scenes from monocular videos to enable high-quality video generation along novel camera trajectories.
Wakals/CoVT
CoVT is a framework that enables Vision-Language Models to reason using continuous visual tokens, improving their spatial reasoning and geometric awareness by grounding semantic thoughts in perceptual cues.
Visionary-Laboratory/holi-spatial
A data curation pipeline that transforms video streams into annotated 3D spatial intelligence datasets, including 3D Gaussian Splatting geometry and spatial QA pairs.
Alibaba-NLP/VRAG
A multi-turn multimodal RAG framework that uses agentic reinforcement learning to enable VLMs to reason and retrieve information across text, images, and videos.
PiSugar/whisplay-ai-chatbot
A pocket-sized AI chatbot device built on Raspberry Pi that enables voice-to-voice interaction, image generation, and local AI acceleration.
bakrianoo/mazinger
An end-to-end video dubbing pipeline that automates transcription, translation, voice cloning, and audio assembly into a single workflow.
facebookresearch/seamless_communication
A family of multimodal AI models for high-quality, expressive, and real-time translation across nearly 100 languages, supporting speech and text modalities.
octimot/StoryToolkitAI
An AI-powered film editing assistant that transcribes, indexes, and searches video footage, allowing editors to use LLMs to automatically create story selections for export to editing software.
MMMU-Benchmark/MMMU
A massive multi-discipline benchmark for evaluating multimodal AI models on college-level subject knowledge and complex reasoning across 30 different subjects.
GPTomics/bioSkills
bioSkills is a library of ready‑made code templates (“skills”) that teach LLM coding agents how to perform bioinformatics analyses—from sequence I/O to single‑cell, variant calling, structural biology, chemoinformatics, and workflow management. Installable via thin scripts for Claude Code, Codex, Gemini/Antigravity, OpenCode, and OpenClaw, the skills provide best‑practice command‑line flags, example scripts and usage guides, improving the correctness of AI‑generated bio‑informatics code.
Jane-xiaoer/paper-collage-ad-codex
A comprehensive workflow skill for OpenAI Codex that automates the production of paper-collage style advertisements, from creative scripting to final MP4 rendering.
xeronsh/openOii
An AI comic-video generation project that uses LangGraph to orchestrate a multi-agent pipeline from story ideas to final video synthesis.
shang-zhu/violin
An open-source video translation tool that transcribes, translates, and dubs videos into 33 languages with native-sounding AI voices.
livekit-examples/python-agents-examples
A collection of runnable Python examples and full-stack applications for building voice, video, and telephony AI agents using the LiveKit Agents framework.
Voine/ChatWaifu_Mobile
An Android application that integrates ChatGPT, local VITS voice synthesis, and Live2D animation to create an immersive, character-based AI companion.
microsoft/psi
Platform for Situated Intelligence (\psi) is a Microsoft‑open‑source .NET framework for building real‑time multimodal AI applications—robots, mixed‑reality assistants, smart‑space systems—by providing a high‑performance streaming infrastructure, visualization tools (PsiStudio), and a library of sensor/AI components. It runs on Windows and Linux, is distributed via MIT‑licensed NuGet packages, and includes tutorials, samples, and an active community.
ChaitanyaEswarRajeshJakki/gemini-youtube-automation
An autonomous AI bot that uses Gemini 2.5 Flash to write, produce, and upload daily educational YouTube videos and Shorts without human intervention.
hassancs91/claude-youtube-editor
An open-source pipeline that automates YouTube video editing by turning raw talking-head recordings into polished videos with AI-generated visuals, audio cleaning, and automated uploads.
suki0dayo/AI_film_studio
A node-based visual editor that streamlines AI movie production by automating the workflow between LLMs for prompting and ComfyUI for image and video generation.
Anionex/dsh-vision-toolkit
A vision toolkit for DeepSeek Harness that gives text-only models visual capabilities for tasks like UI restoration, long-screenshot OCR, and GUI automation.
mintdotgg/mint-playground
A collection of open-source Three.js experiences demonstrating interactive 3D showrooms, games, and data visualizations built with the Mint asset pipeline.
NVIDIAGameWorks/kaolin
A PyTorch library from NVIDIA providing GPU-optimized modules for 3D deep learning, including differentiable rendering, physics simulation, and support for 3D Gaussian splats.
ParisNeo/lollms-webui
A unified local web interface for accessing hundreds of LLMs and multimodal AI models for text, image, video, and music generation.
GoogleCloudPlatform/vertex-ai-creative-studio
A web application for exploring Google Cloud's generative media models, providing a unified interface for image, video, music, and speech generation and creative workflows.
Stonesjtu/pytorch_memlab
pytorch_memlab is a Python library that profiles CUDA memory usage line‑by‑line, lists live tensors, and can temporarily move GPU data to CPU. It integrates with IPython/Jupyter via `%mlrun` magics and helps debug out‑of‑memory errors in PyTorch models.
pnnx/pnnx
pnnx is an open‑source tool that optimises PyTorch models and exports them to a lightweight, dependency‑free format (PNNX) and to ncnn/ONNX‑zero files, enabling fast inference on edge devices.
ljquan/opentu
Opentu is a canvas-based AI application platform that integrates multi-model generation and tool management into a single workspace for continuous AI task execution.
OpenSQZ/MiniCPM-V-CookBook
A comprehensive cookbook of deployment and fine-tuning recipes for the MiniCPM series of multimodal models, supporting text, vision, and audio capabilities across various hardware.
Cjbuilds/Codex-Orchestration
Codex Orchestration is a Codex plug‑in that lets you attach other LLMs (Claude Fable 5, Opus 5, OpenRouter models, etc.) to a task and assign them specific roles—Planner, Advisor, Designer, Executor. Codex stays the top‑level orchestrator, iterating plans, reviewing them, optionally designing UX artefacts, and finally executing code. The plug‑in is installed via the Codex marketplace, configured with skill prompts inside a Codex chat, and supports role‑specific effort levels. It aims to improve plan quality, enable parallel work, and reduce premium‑model usage. MIT‑licensed.
dromara/wgai
WGAI is a Spring‑Boot + Vue web platform that bundles vision (OpenCV, YOLO, OCR, face ID), audio (speech‑to‑text, ChatGPT‑style dialogue), digital‑human avatars, and AGV navigation into an offline‑first, industrial‑grade AI management system. It offers online data annotation, model training, real‑time video/audio analysis, and system monitoring, with a UI for configuration and a public demo instance. The code is AGPL‑3.0 and support is mainly via a paid Chinese community.
visualbruno/ComfyUI-Hunyuan3d-2-1
A ComfyUI wrapper for Tencent's Hunyuan3D-2.1 model that enables the generation of textured 3D models within a visual node-based workflow.
zai-org/GLM-V
A series of vision-language models (VLMs) that enhance multimodal reasoning, supporting native function calling, visual grounding, and complex document understanding.
PaddlePaddle/ERNIE
A family of large-scale multimodal MoE models and development toolkits providing state-of-the-art performance in text and vision-language understanding.
cambrian-mllm/cambrian-s
A suite of multimodal LLMs and datasets designed to enhance spatial reasoning and "spatial supersensing" in video understanding.
qualcomm/aimet
AIMET is Qualcomm’s open‑source toolkit for quantizing and compressing PyTorch/ONNX models, enabling fast, low‑memory inference on edge devices with minimal accuracy loss.
SamurAIGPT/AI-Influencer-Generator
An open-source pipeline that uses Stable Diffusion, gTTS, and SadTalker to create and animate consistent virtual AI influencers for social media.
volcengine/rtc-aigc-demo
A demo project showcasing an interactive AIGC pipeline that integrates RTC, ASR, LLM, and TTS for real-time voice-based AI interactions.
ShayneP/local-voice-ai
A private, low-latency voice assistant that runs locally on your hardware using a stack of speech recognition, LLM, and speech generation models.
sympozium-ai/sympozium
Sympozium is an open‑source, Kubernetes‑native platform that coordinates multi‑agent AI systems. It provides CRD‑based policies, shared SQLite memory, sidecar‑isolated skills, and a visual workflow canvas, letting teams orchestrate agents at scale while keeping governance and observability built‑in.
nv-tlabs/Gamma-World
Gamma-World is a generative multi-agent world model that simulates shared environments for multiple independently controllable agents, supporting real-time interactive video generation at 24 FPS.
kandinskylab/kandinsky-5
A family of diffusion models for high-quality video and image generation from text or images, featuring Pro and Lite versions with strong English and Russian language support.
OliBomby/Mapperatorinator
A multi-model AI framework that generates fully featured osu! beatmaps from audio spectrograms and provides an AI-driven modding tool to detect mapping inconsistencies.