AaronFeng753/Waifu2x-Extension-GUI
A graphical user interface for AI-powered image and video upscaling and frame interpolation, supporting a wide range of hardware and neural network models.
JuneYaooo/gpt-image2-ppt-skills
gpt-image2-ppt-skills is a Python skill for multimodal AI agents that uses OpenAI’s gpt‑image‑2 model to generate high‑quality 16:9 slide images and package them into a .pptx file. It supports a visual‑first mode and an optional editable mode that reconstructs slides into native PowerPoint objects. Users can create new decks, clone existing .pptx templates, or edit specific slide elements via natural‑language prompts. Installation is done by asking a supported agent to install the skill or by running a provided script; an OpenAI API key and a local PowerPoint/Keynote/LibreOffice renderer are required. The project is Apache‑2.0 licensed.
kleinlee/DH_live
A lightweight 2D digital human engine that enables real-time animation and dialogue on mobile browsers and low-power devices without requiring a GPU.
open-compass/VLMEvalKit
An open-source evaluation toolkit for large vision-language models that enables one-command evaluation across 70+ benchmarks and 200+ models.
SeemSeam/claude_codex_bridge
CCB (Claude‑Codex‑Bridge) is a cross‑platform terminal UI that lets you run many LLM‑backed CLI agents (Codex, Claude, Gemini, etc.) together in visible panes, define collaboration graphs, share a project‑wide memory file, and even control the workspace from an Android app. Install via npm, configure with a built‑in UI, and manage providers, roles, and rich‑media terminals from the command line.
JimLiu/baocut
An agent skill that enables AI coding agents to control BaoCut for automated video transcription, subtitle translation, and timeline editing via natural language.
PhiloLabs/fable51-worlds
A system that uses AI agent swarms to generate walkable 3D worlds from text, images, or video, rendering them as browser-based Three.js applications.
QwenLM/Qwen3.8-Flash-Next
A multimodal MoE model that serves as an architectural preview for Qwen4, offering high efficiency in coding and office tasks with significantly reduced training and inference costs.
amitshekhariitbhu/build-your-own-x-machine-learning
A collection of pure‑Python, from‑scratch implementations of classic ML algorithms, deep‑learning models, and applied projects (recommendation systems, CV apps, NLP, etc.). Designed as a hands‑on learning resource for anyone who wants to understand the inner workings of machine‑learning techniques.
DrEAmSs59/CS2-insight-agent
A one-stop creation suite for CS2 that automates demo analysis, highlight recording via OBS, and video editing with AI-powered commentary.
redai-studio/Relax
Relax is an open‑source, Ray‑Serve‑based reinforcement‑learning framework for multimodal large language models. It decouples rollout and training via a TransferQueue, supports fully asynchronous or hybrid execution across GPU clusters, and provides a rich set of on‑policy algorithms (PPO, GRPO, M2‑PO, RLOO, REINFORCE++, etc.) plus plug‑in rewards. The system works with Megatron‑LM training back‑ends and SGLang inference, handling text, vision, video, and audio in a single pipeline. Official Docker images, bilingual docs, and production‑grade ops (health manager, metrics, elastic rollout scaling) make it ready for research and large‑scale production fine‑tuning of models such as Qwen‑3‑Omni.
VAST-AI-Research/TripoSplat
TripoSplat is a tool that converts a single 2D image into high-quality 3D Gaussian Splats, enabling fast 3D asset creation for games and AR/VR.
apple-aiml-research/ml-mobileclip
MobileCLIP (and MobileCLIP2) are compact, fast image‑text models for zero‑shot classification and captioning, optimized for mobile devices. The repo offers pretrained checkpoints, training/evaluation scripts (built on OpenCLIP), an iOS demo, and CoCa caption models, all under MIT / Apple research licenses.
xr843/insect-world
An interactive 3D insect encyclopedia featuring 63 species generated entirely through parametric TypeScript code rather than external 3D assets.
morettt/my-neuro
A comprehensive workbench for creating personalized, human-like AI companions with customizable voice, personality, and Live2D visuals.
nexu-io/codex-slides
An open-source AI slide studio for Codex coding agents that transforms prompts, repos, or files into professional presentations through a steerable, live-editing workflow.
ArcInstitute/state
A framework for predicting cellular responses to perturbations and generating cell embeddings to model biological state transitions.
HisMax/RedInk
An AI-powered tool that generates complete multi-page social media posts, including outlines, copywriting, and visually consistent images, from a single sentence.
GetStream/Vision-Agents
A framework for building low-latency, multi-modal AI agents that can watch, listen, and understand video in real time using a combination of CV models and LLMs.
TEN-framework/ten-framework
An open-source framework for building real-time multimodal conversational AI agents with low-latency audio, text, and visual capabilities.
SakuraMathcraft/LaTeXSnipper
A desktop application that converts screenshots, images, and PDFs into editable LaTeX formulas and text using OCR.
microsoft/mcp-gateway
MCP Gateway is an open‑source Kubernetes‑native reverse proxy and management layer for Model Context Protocol (MCP) servers. It provides a REST control‑plane for creating, updating, and monitoring MCP adapters and tools, session‑aware routing, Azure Entra ID RBAC, a built‑in React portal, and optional LLM‑driven agents. Deploy locally with Docker/K8s or with a one‑click Azure template.
KangLiao929/Puffin
A series of unified multimodal models for 3D world modeling that enable camera-centric spatial intelligence and 3D world generation using native physics, geometry, and appearance states.
nv-tlabs/lyra
A series of open generative 3D world models from NVIDIA that create 3D and 4D scenes from a single image or video.
dmMaze/BallonsTranslator
A deep learning-powered comic translation tool that automates text detection, OCR, inpainting, and translation for manga and comics.
google-deepmind/gemma
A JAX library for deploying and fine-tuning the Gemma family of open-weights large language models, supporting multi-modal conversations and multiple hardware backends.
EvolvingLMMs-Lab/lmms-eval
LMMs‑Eval is an open‑source Python toolkit that unifies >100 multimodal benchmark tasks (image, video, audio) for evaluating large language models with vision/audio capabilities. It offers deterministic, statistically‑sound results, supports 30+ model families via vLLM, SGLang, or OpenAI‑compatible APIs, and includes a web UI, an HTTP evaluation server, and extensible model/task interfaces.
Lynpoint/CyberVerse
An open-source framework for building real-time digital-human agents that combine voice interaction, persona memory, and video animation from a single photo.
duixcom/Duix-Mobile
An open-source SDK for deploying real-time interactive AI avatars on mobile and embedded devices with low latency and on-device execution.
momori777/Artemis
A fully local, uncensored AI companion system that integrates LLMs, TTS, image generation, and Live2D animations to create private, character-driven virtual partners.
chatfire-AI/huobao-canvas
Huobao Canvas is an open‑source, node‑based visual editor that lets you chain text, image, and video AI models from 11 providers on an infinite canvas. It offers a lightweight Node server for persistence and async run queues, Docker and Electron deployment options, and a single‑key setup via Huobao. Ideal for building multimodal creative pipelines without writing code.
MirroS-Lab/Code-as-World
Code-as-World is a framework that represents the physical world as executable code to enable AI models to perform quantitative physical reasoning by discovering world representations through iterative simulation.
IBM/AssetOpsBench
AssetOpsBench is an open‑source benchmark/framework for building, orchestrating, and evaluating LLM‑based AI agents that operate on industrial asset‑management data (sensors, work orders, failure modes, etc.). It provides domain‑specific MCP tool servers, several ReAct‑style agent back‑ends, 141+ realistic scenarios, and a multi‑dimensional evaluation pipeline used in KDD/AAAI/NeurIPS competitions.
Kiri-Innovation/3dgs-render-blender-addon
A Blender add-on for importing, editing, animating, and rendering 3D Gaussian Splats, integrating high-fidelity 3D captures into a standard 3D production workflow.
Runfusion/Fusion
Fusion is an open‑source AI‑powered software factory that turns plain‑language task descriptions into fully reviewed, merge‑ready code. It coordinates a fleet of LLM‑driven agents, runs each task in its own Git worktree, and presents the whole process on a visual dashboard with configurable workflows and human oversight.
ArtificialAnalysis/Stirrup
Stirrup is a lightweight Python framework for building LLM‑driven agents. It supplies ready‑made tools (code execution, web search, file I/O, multimodal handling) and a flexible `Tool`/`ToolProvider` system, while automatically managing context limits and session lifecycles. Install via `pip install stirrup` (optional extras for Docker, browsers, etc.), create a client (OpenRouter, LiteLLM, or any OpenAI‑compatible API), instantiate an `Agent`, and run a session that lets the model invoke tools to solve tasks. Custom tools and providers are easy to add, making Stirrup suitable for rapid prototyping, domain‑specific assistants, and research on tool‑using agents.
zyddnys/manga-image-translator
An AI-powered image translator that detects, removes, and replaces text in manga and comics, supporting multiple languages and automated typesetting.
DavidVentura/offline-translator
An Android translator app that performs text, PDF/ODT document, and image translation completely offline using on-device models.
P1kaj1uu/ChattyPlay-Agent
ChattyPlay‑Agent is a full‑stack web app that combines an AI chat assistant (ChatGPT/DeepSeek) with many everyday tools – music streaming, video parsing, gold‑price charts, academic‑paper browsing, LaTeX/markdown editors, mind‑maps, manga library, Xianyu marketplace helper, and text‑to‑image generation. Built with React + TypeScript front‑end and Python/Java back‑end, it runs via Docker and is available as an online demo. The repository is a genuine software project in the AI‑assistant space.
RunMaestro/Maestro
Maestro is a cross‑platform desktop app that lets power users run, automate, and monitor many AI coding agents (Claude Code, OpenAI Codex, etc.) in parallel, with Git worktree support, playbook automation, a keyboard‑first UI, remote web control, and analytics—all under an AGPL‑3.0 license.
Tencent-Hunyuan/UniRL
UniRL is an open‑source Python framework that applies a unified reinforcement‑learning loop to many multimodal generative models (LLMs, vision‑language models, image/video diffusion, and hybrid AR‑diffusion models). It provides four entry points, Hydra‑based configs, pluggable rollout engines, and distributed training via Ray and FSDP. The repo ships standard RL algorithms plus three novel team‑proposed methods (Flow‑DPPO, DRPO, CPPO) with tutorials. Supported models include Stable Diffusion 3, FLUX.2‑Klein, Qwen‑3, HunyuanVideo, and more. Documentation, example recipes, and a WeChat community are provided. Licensed under Apache‑2.0.
liyue-aigc/female-outfit-director
A prompt-directing workflow for AI agents that generates structured production plans and prompts for creating character-consistent outfit-change videos.
mlfoundations/open_clip
OpenCLIP is an open‑source PyTorch library that implements OpenAI’s CLIP and many newer multimodal models (SigLIP, CoCa, MaMMUT, CLAP, NaFlex‑enabled vision/audio, modern text towers). It provides pretrained checkpoints, a flexible training stack with distributed/FSDP2 support, mixed‑precision and torch.compile acceleration, and utilities for zero‑shot inference and generative captioning.
YGYOOO/WorldX
WorldX is an AI-powered simulation engine that generates a complete virtual world with autonomous agents, maps, and emergent narratives from a single text prompt.
huangserva/3DCellForge
An AI-powered 3D model studio that converts images into interactive 3D models using multiple AI providers and provides a professional inspection and presentation workspace.
nv-tlabs/3dgrut
A hybrid 3D rendering framework that combines Gaussian rasterization and ray tracing to support distorted cameras and complex light effects like reflections and shadows.
Scottcjn/bottube
An AI-native video platform where AI agents and humans create and share short-form videos, featuring hardware-verified provenance via Proof of Physical AI.
microsoft/DebugMCP
DebugMCP is a Microsoft VS Code extension that runs a local MCP server, exposing debugger actions (start/stop, step, breakpoints, variable inspection, expression evaluation) as tools that any MCP‑compatible AI assistant (Copilot, Cursor, Codex, etc.) can invoke. It works out‑of‑the‑box for many languages, runs entirely locally on port 3001, and includes security checks (loopback‑only binding, host validation, secret redaction). The companion *debug‑live* skill provides the higher‑level debugging workflow, letting AI agents autonomously debug code directly inside VS Code.