MoonshotAI/Kimi-K3
Kimi K3 is an open-weight, 2.8T-parameter native multimodal MoE model designed for long-horizon coding, deep research, and complex agentic knowledge work.
modelscope/DiffSynth-Studio
An open-source diffusion model engine that enables high-performance inference and training of image, video, and audio generative models on consumer-grade GPUs.
solidSpoon/DashPlayer
An AI-enhanced video player for English learners that provides local AI subtitle generation, sentence breakdown, and vocabulary tools to facilitate immersive language acquisition.
flyteorg/flyte
Flyte 2 is an open‑source Python framework for defining, orchestrating and serving ML pipelines and AI agents. It lets you write ordinary Python functions as tasks, compose them into workflows, and run them locally or on a future Kubernetes‑native backend. The project includes a CLI, a developer‑focused TUI, and first‑class support for FastAPI model serving.
Jamailar/Beav
Beav (formerly RedBox) is a cross‑platform desktop AI workstation for self‑media creators. It lets you collect web/social‑media material, store it in a local knowledge base, generate text, images and video with AI, edit voice‑over clips, add animations, and schedule publishing—all from one UI. The app ships with a Chrome/Edge scraper, daily operation reports, reusable image templates, a multi‑track video editor, and optional integration with external agents. It runs on macOS, Windows and Linux, uses either the built‑in AI service or any OpenAI‑compatible endpoint, and is released under an MIT‑NC (non‑commercial) license.
smixs/visual-skills
A toolkit of AI agent skills that applies professional filmmaking and dramaturgy principles to generate high-precision prompts for AI video and image models.
ErickWendel/localstudio
A browser-native, local-first slide editor that uses Web AI and WebGPU to provide AI-powered slide generation, translation, and presentation tools without a backend.
verl-project/verl-omni
A general RL training framework for multimodal generative models, providing fast rollouts and stable post-training for diffusion and omni-modality models.
ModernRelay/omnigraph
Omnigraph is a Rust‑based, lakehouse‑style graph database that stores data in the columnar Lance format on any S3‑compatible object store. It is built for fleets of AI agents: each agent can write to its own isolated branch, then submit changes for review and Git‑style merge. Queries fuse graph traversal, vector ANN, and full‑text search in one runtime, and Cedar policies enforce fine‑grained security on every mutation. The project ships a CLI, an Axum HTTP server, TypeScript SDKs, and a “MCP” bridge for LLM hosts, making it a full‑stack solution for agentic memory, company‑brain knowledge graphs, dev‑graph automation, and versioned ML data layers.
NVIDIA/cosmos
An open platform of omnimodal world models and tools for building Physical AI, enabling the joint processing and generation of text, vision, audio, and action sequences for robotics and autonomous systems.
TrenTorch/TrenTorch
Tren⚡Torch is an educational, NumPy‑only deep‑learning framework that re‑implements TinyTorch and adds 20 progressive modules—from tensors and autograd to CNNs, transformers, quantization and benchmarking—plus a CLI and historic‑milestone scripts for hands‑on learning.
Yuliang-Liu/MonkeyOCRv2
A visual-text foundation model for Document AI that provides a multilingual vision encoder and specialized models for document parsing and understanding.
apify/mcpc
mcpc is a lightweight, cross‑platform CLI that gives full access to the Model Context Protocol (MCP) – the standard API for AI agents’ tools, prompts, tasks, and payments. It lets agents and scripts call any MCP operation from a Bash shell, supports persistent sessions, OAuth 2.1 stored in the OS keychain, JSON‑friendly output for piping, dynamic tool discovery via grep, and experimental x402 payment integration.
nexmoe/VidBee
An open-source desktop app that downloads video and audio from 1000+ sites, transcribes them locally, and uses AI to summarize and search the content.
kgoedecke/doop
Doop is an open‑source, real‑time design canvas where humans and AI agents (via Claude Code or a built‑in Doop Agent) edit HTML frames together. It offers live multiplayer features, AI‑driven design tasks, private sharing, and can be self‑hosted with a single Docker command or `bun run dev`. The AI backend works with Anthropic (free tier) or user‑provided OpenAI/ChatGPT keys.
869413421/ai-moive-studio
A full-stack AI content creation workbench that transforms text into AI movies and short videos using an infinite canvas and agent-assisted workflows.
Huanshere/VideoLingo
An AI-powered video translation and dubbing tool that automates transcription, translation, and voiceover generation through a unified Streamlit interface.
nicobailon/pi-mcp-adapter
A Pi agent plugin that proxies MCP servers through one lightweight tool, loading servers lazily and discovering tools on demand to save context window.
taylorwilsdon/google_workspace_mcp
Google Workspace MCP Server is an open‑source Python service that exposes 120+ Google Workspace actions (Gmail, Drive, Docs, Calendar, etc.) through the Model‑Client‑Protocol, enabling LLM assistants like Claude or ChatGPT to read and write Workspace data. It supports multi‑user OAuth 2.1, three tool‑tiers, stateless container deployment, and a full CLI, while keeping all traffic limited to Google’s APIs and offering MIT‑licensed, self‑hostable code.
localai-org/kimodo.cpp
A GGML/C++ implementation of NVIDIA's Kimodo model that generates character and robot motion animations from text prompts.
OpenBMB/MiniCPM-V
A series of pocket-sized multimodal LLMs designed for efficient on-device deployment, enabling high-performance image, video, and real-time omnimodal interaction on mobile phones.
BasedHardware/omi
An open-source AI memory system that captures screen and audio data from wearables and desktop apps to provide real-time transcription, summaries, and a searchable AI chat.
tracel-ai/burn
Burn is a Rust‑native deep‑learning framework that unifies training and inference under a single API. It offers PyTorch‑like ergonomics, automatic differentiation, JIT‑compiled kernel fusion, and a flexible backend system (CUDA, ROCm, Metal, Vulkan, WebGPU, CPU, and no‑std). With tools for ONNX import, weight loading, a terminal training dashboard, and support for embedded and browser environments, Burn lets you write a model once and run it anywhere.
huggingface/diffusers
A modular library for using and training state-of-the-art diffusion models to generate images, audio, and 3D molecular structures.
koharu-rs/koharu
An ML-powered manga translator written in Rust that automates text detection, OCR, translation, and generative inpainting for a local-first workflow.
off-grid-ai/OGAM
A comprehensive offline AI suite for mobile and Mac that provides text, image, and vision AI, and voice transcription, all running natively on-device for total privacy.
icebird1998/scientific-illustrator
A Codex plugin that converts reference images into editable diagrams by automatically redrawing them in Microsoft PowerPoint, WPS Presentation, or draw.io.
modelscope/FunClip
An open-source automated video clipping tool that uses ASR and LLMs to extract video segments based on spoken text, speaker identity, or AI-driven content analysis.
bytebase/dbhub
DBHub is a minimal, token‑efficient MCP server that lets LLM‑based tools query PostgreSQL, MySQL, SQL Server, Oracle, SQLite, and MariaDB safely. It provides only two core tools (SQL execution and schema search) in ~1.4 k tokens, optional extra tools, guardrails, and a built‑in web UI, making it a lightweight bridge between AI agents and real databases.
datascale-ai/opentalking
An open-source orchestration framework for real-time digital-human conversations, integrating LLMs, TTS, STT, and video rendering via WebRTC.
diffusionstudio/lottie
An open-source framework that enables coding agents to generate production-ready Lottie animations from text prompts and SVG assets.
Tencent-Hunyuan/Hunyuan3D-WorldClaw
WorldClaw is an agentic 3D generation framework designed to create large-scale, open-world 3D environments.
pollinations/pollinations
An open-source generative AI platform providing a unified, OpenAI-compatible API for text, image, video, audio, 3D, and embedding generation.
TencentARC/Pixal3D
Pixal3D is a high-fidelity 3D asset generator that creates detailed geometry and PBR textures from a single image or multiple views using pixel-aligned back-projection.
ahujasid/ableton-mcp
A Model Context Protocol (MCP) server that connects Claude AI to Ableton Live, enabling prompt-driven music production, track manipulation, and autonomous song arrangement.
google-deepmind/alphagenome
A unifying model and API for deciphering the regulatory code of DNA sequences by predicting gene expression, splicing, and chromatin features.
liuzhao1225/YouDub-webui
An open-source video localization tool that automates transcription, translation, and dubbing for videos from YouTube, Bilibili, or local files.
neavo/LinguaGacha
LinguaGacha is a cross‑platform desktop app that uses large‑language‑model APIs (OpenAI, DeepSeek, Anthropic, Google, or local models) to translate subtitles, e‑books, and game scripts in dozens of languages with a single click. It offers an “AGENT” workflow for automatic glossary extraction, batch translation, and auto‑review, preserving formatting and code‑style elements.
oil-oil/oil-motion
Oil Motion is an AI‑agent skill that turns design prompts and source images into interactive web animations (scroll‑driven, mouse‑follow, touch, etc.). It automates key‑frame creation, AI video generation, frame‑level cleaning, and packaging into MP4 or WebP assets, then provides ready‑to‑embed front‑end code.
SamurAIGPT/Generative-Media-Skills
A multimodal toolset and MCP server that enables AI agents to generate and edit professional images, videos, and audio using a schema-driven architecture.
jnMetaCode/agency-orchestrator
Agency Orchestrator is an npm‑based tool (CLI + web UI) that automatically composes and runs multi‑agent AI workflows from a single sentence or a tiny YAML file. It ships 276 pre‑built Chinese (and 191 English) role definitions, supports 15 LLM providers (many key‑less), builds a DAG for parallel execution, validates outputs, and produces shareable HTML reports. Ideal for solo entrepreneurs, product analysts, and developers who want an "AI team" without writing code.
jd-opensource/JoyAI-VL-Interaction
An 8B-scale vision-language interaction model and system that enables AI to proactively respond to real-time video streams in under a second without waiting for user prompts.
diivi/aseprite-mcp
An MCP server that gives AI assistants full control over Aseprite, enabling them to create and animate professional-grade pixel art through a comprehensive set of 104 tools.
jaredrhod/barehands
A webcam-based hand-tracking interface that lets you manipulate notes, images, and 3D models spatially, designed to be wired into an AI agent as a visual and interactive body.
777genius/agent-teams-ai
Agent Teams AI is a free cross‑platform desktop app that lets you build and monitor teams of AI agents (Claude Code, Codex, OpenCode, Cursor, SuperGrok, GitHub Copilot, Z.AI, MiniMax, Kiro, etc.). Agents can create tasks, communicate, run code and terminal commands, and perform code reviews while you watch a Kanban board. The app tracks token usage, budgets, and per‑agent resource stats, and offers a built‑in terminal, Git‑aware editor, multilingual UI, and plug‑in support via an `mcp-server` component.
IBM/mcp-context-forge
ContextForge is an open‑source gateway that federates MCP, A2A, REST and gRPC services into a single, centrally governed endpoint for AI agents and tools. It provides auth, rate‑limiting, OpenTelemetry tracing, a plug‑in system, and an admin UI, and can be run via PyPI, Docker, or Helm.
google-gemini/gemini-live-api-examples
A collection of examples for the Gemini Live API, enabling developers to build low-latency, real-time voice and video AI agents with multimodal input.
X-PLUG/MobileAgent
A family of multi-platform GUI agents and foundation models that automate tasks across mobile, desktop, and web interfaces using visual perception and planning.