YouMind-OpenLab/awesome-seedance-2-prompts
A curated, CC‑BY‑4.0 collection of 6 405 community‑submitted prompts for ByteDance’s Seedance 2.0 multimodal video‑generation model, with featured examples, a searchable web gallery, and contribution guidelines.
niedev/RTranslator
An offline, real-time translation app for Android that uses on-device AI models like Whisper and NLLB to enable private, internet-free conversations.
HUANGCHIHHUNGLeo/claude-real-video
A local video processing pipeline that extracts meaningful keyframes and transcripts for LLMs, replacing fixed-interval sampling with intelligent scene-change detection.
op7418/guizang-yingzao-skill
An AI-powered design skill that transforms architectural and cultural photos into art-directed editorial posters with integrated Chinese typography and spatial awareness.
fengshao1227/ccg-workflow
CCG‑Workflow is a Node‑JS CLI that makes Claude Code the central orchestrator for multiple code‑generation models (Codex, Gemini, etc.). It injects JavaScript hooks to keep context, creates persistent task folders, runs built‑in strategies (e.g., full‑collaborate), applies security/quality gates, and can even deploy to Baota panels. Install with `npx ccg-workflow` (full mode) or as a Claude plugin for just the reusable skills.
hgmzhn/manga-translator-ui
A comprehensive manga translation tool that automates text detection, OCR, translation, and typesetting using AI models and a visual editor.
jingyaogong/minimind-o
MiniMind‑O is a genuine open‑source, tiny (≈0.1 B) end‑to‑end omni model that jointly handles text, speech and images as inputs and produces text and streamed speech outputs. It provides full PyTorch code, training scripts, mini/full datasets, pretrained weights, voice‑cloning prompts, and demo interfaces (CLI, Gradio, phone‑mode UI). The repository is clearly an AI/ML project, not an off‑topic list.
xberg-io/xberg
Xberg is an open‑source, Rust‑based document‑intelligence engine that extracts clean, structured text (including OCR, tables, code structure, and optional LLM‑driven enrichment) from 107 formats. It ships with bindings for 15 languages, a CLI, Docker image, REST API, and an MCP server for AI coding assistants.
GMGNAI/gmgn-skills
GMGN Skills equips AI agents with real-time, multi-chain on-chain data and trading capabilities for meme tokens via CLI commands, enabling automated research, analysis, and execution without scraping.
ooboqoo/interview-coder-cn
An AI-powered screenshot assistant that analyzes on-screen questions in real-time to provide answers while remaining invisible to screen-sharing software.
taoufik123-collab/claude-watch
A Claude skill that enables the AI to watch videos from URLs or local files by extracting key frames and transcripts for multimodal analysis.
zilliztech/memsearch
memsearch is a Python library and CLI that records AI coding‑assistant chats as Markdown, builds a hybrid (BM25 + dense) Milvus index, and provides cross‑platform recall via plugins for Claude Code, Codex, DeepSeek Harness, OpenClaw, and OpenCode. It supports local ONNX embeddings, optional Zilliz Cloud backend, background project/user notes, and automatic skill distillation.
argosopentech/argos-translate
Argos Translate is an open‑source Python library for offline neural‑machine translation. It downloads pre‑trained OpenNMT models packaged as ".argosmodel" files, supports pivoting through intermediate languages, offers GPU acceleration via CTranslate2, and can be used via a Python API, CLI, or a separate desktop GUI. The project also powers the LibreTranslate web/API service.
nyldn/claude-octopus
Claude Octopus is a Claude Code plugin that adds a multi‑LLM orchestration layer. It stays dormant until you run an `/octo:*` command, then can call up to twelve external providers (Codex, Copilot, Ollama, etc.) to run consensus‑gated reviews, adversarial councils, or a full “spec‑to‑software” pipeline. It ships with dozens of personas, slash commands, and method‑aware workflows, integrates with persistent memory tools, and works in Claude Code, Codex CLI, Cursor IDE (via an MCP server), and OpenCode.
open-multi-agent/open-multi-agent
Open Multi‑Agent (OMA) is a TypeScript framework for orchestrating dynamic teams of LLM agents inside Node.js apps. It builds a task DAG at runtime from a high‑level goal, executes the plan with optional human approvals, records the full trace for inspection or replay, and supports any LLM provider, tool gating, budgets, and OpenTelemetry integration.
Ch921-cell/Remember-R1
Remember-R1 is a reinforcement learning framework designed to prevent multimodal models from forgetting visual information during long-context reasoning.
llm-as-a-verifier/llm-as-a-verifier
llm-as-a-verifier is a Python framework that turns LLMs into fine‑grained, probabilistic verifiers for AI agent outputs. It scores candidate trajectories, selects the best‑of‑N via an efficient Probabilistic Pivot Tournament, tracks progress step‑by‑step, and supports multimodal (image) inputs. The library achieves state‑of‑the‑art results on several agentic benchmarks and includes scripts for reproducible evaluation, a cache‑optimised token usage tracker, and a Claude Code plugin (TurboAgent) for seamless integration.
jingyaogong/minimind-v
A minimal implementation of a Vision Language Model (VLM) that allows users to train a 65M parameter multimodal model on a single consumer GPU in just 2 hours.
50kg/image-to-slice
An AI-powered UI decomposition tool that converts flat UI images into editable Figma layers and HTML/CSS code by intelligently slicing assets and inpainting backgrounds.
halcyon-video/halcyon-video
A walkable 3D video store interface for media libraries like Jellyfin, Plex, and Emby that turns digital catalogs into an immersive physical browsing experience.
Blizaine/Maestro
A local AI creative studio and video editor that uses an LLM-directed workflow to automate the production of music videos and short films.
hero8152/Infinite-Canvas
A visual, infinite canvas workspace that integrates multiple AI APIs and local ComfyUI workflows for collaborative generative AI content creation.
JohnKinyanjui/sprite-maker
A local-first AI workbench for creating, animating, and exporting 2D game art, featuring identity-preserving AI animation and a deterministic Rust-based rigging engine.
shootthesound/Fizgig
A LoRA training and refinement studio that allows users to fine-tune and repair AI image and video models on consumer GPUs.
Adam-CAD/CADAM
An open-source text-to-CAD web app that uses AI to transform natural language and images into parametric 3D models via OpenSCAD.
tover0314-w/opentypeless
An open-source AI voice input tool for macOS, Windows, and Linux that provides app-aware dictation, voice Q&A, and text rewriting.
MCPJam/inspector
MCPJam Inspector is an open‑source testing and evaluation platform for servers that implement the Model‑Chat‑Protocol (MCP) used by chat‑based AI products. It provides a web playground, OAuth debugger, server‑debug tools, reusable “skills”, shared workspaces, eval test‑case runner, a CLI, an SDK, and CI/CD integrations to catch regressions across 16 client configurations and 170+ LLM models.
csyqlz/VOZEB-PRO
A full-stack AI creation platform that integrates a multimodal agent, a visual canvas, and a short-drama production pipeline into a single commercial-ready SaaS application.
AgnesAI-Labs/AgnesAI-Models
A unified OpenAI-compatible API gateway providing access to high-performance multimodal foundation models for text, image, and video generation.
penecho/penecho
PenEcho is a visual canvas app that integrates with LLM agents via the Model‑Canvas‑Protocol (MCP). It lets you draw, annotate, embed widgets, and have AI agents read and edit the canvas, creating a tight feedback loop between conversation and visual work. Available as a desktop app or npm package, it supports both local and cloud‑hosted models and offers versioned projects, sharing, and a built‑in assistant.
oomol-lab/pdf-craft
A Python library that converts scanned PDFs into editable Markdown or EPUB files using OCR and LLM-powered translation.
digitalsamba/claude-code-video-toolkit
An AI-native video production workspace for Claude Code that automates scriptwriting, asset generation, and rendering for low-cost, professional explainer videos.
NomaDamas/CozyClay
A browser-based previs studio for blocking scenes and posing characters to create high-control reference assets for AI video generation.
modstart-lib/aigcpanel
A one-stop AI digital human desktop application that integrates lip-syncing, voice cloning, and audio-visual tools for easy content creation.
Blaizzy/mlx-vlm
MLX‑VLM is a Python library that enables inference, fine‑tuning, and serving of vision‑language (and multimodal) models on Apple Silicon via the MLX runtime. It provides a CLI, FastAPI server, Gradio chat UI, and advanced speed‑up techniques such as speculative decoding (DFlash, Gemma‑4 MTP, EAGLE‑3). The package supports dozens of pre‑wrapped models, low‑bit quantisation, thinking‑budget control, and an agent‑skills bundle for easy development.
QwenLM/Qwen3.8-Flash-Next
A multimodal MoE model that serves as an architectural preview for Qwen4, offering high efficiency in coding and office tasks with significantly reduced training and inference costs.
NVIDIA-NeMo/Nemotron
A family of open, high-efficiency multimodal models and training recipes designed for agentic AI, providing tools for the full lifecycle from data curation to deployment.
jordanrendric/claude-video-vision
A Claude Code plugin that enables video understanding by extracting frames and transcribing audio via local or cloud backends.
Tencent-Hunyuan/HY-World-2.0
A multi-modal world model framework that generates and reconstructs editable 3D assets (meshes and Gaussian Splattings) from text, images, or video for use in game engines.
rome-os/rome
Rome is an open‑source “agentic OS” that lets humans and AI agents collaborate in a persistent, self‑hosted environment. It introduces **Rome Apps**—git‑tracked packages containing actions, agents, skills, UI, and data—so capabilities survive beyond a single chat. You can run it locally via Docker/Electron or use the preview Rome Cloud, build apps with the provided SDKs, and share them through an app store. The platform focuses on composable, reusable software rather than just model prompts, positioning itself between persistent‑agent services and personal‑software platforms.
ahujasid/camera-to-blender
A tool that converts photos of real-world objects into 3D models using AI and automatically imports them into Blender.
WecoAI/aideml
AIDE ML is an open‑source Python library that implements an LLM‑guided tree‑search agent for automatically writing, debugging, and optimizing machine‑learning pipelines. Users describe a dataset, a goal, and an evaluation metric in natural language; the agent iteratively generates code, evaluates it, and prunes the search tree until the metric is maximized. The package includes a CLI, a Streamlit UI, HTML visualisation, and works with any OpenAI‑compatible LLM (including local Ollama models).
mizorewww/course2md
A tool that converts YouTube, Bilibili, and local videos into illustrated Markdown or HTML notes using speech recognition.
GenielabsOpenSource/spine-animation-ai
An AI-powered toolset for Spine 2D that automates character rigging, asset positioning, and animation generation from raw images.
alibaba/lumenx
An AI-native motion comic and video creation platform that transforms scripts into finished videos through an integrated pipeline of script analysis, asset generation, and composition.
anymouschina/TapCanvas
An agent-native infinite canvas for AI film and multimodal content production that integrates scripts, assets, and storyboards into a traceable workflow.
raojiacui/prompt-lens
An AI video analysis tool that reverse-engineers existing videos to extract prompts and scripts, helping creators replicate and iterate on successful AI video styles.
laravel/mcp
Laravel MCP is a Laravel package that provides a ready‑made Model Context Protocol (MCP) server, enabling AI clients to interact with a Laravel app’s data and logic via a standard protocol.