kyutai-labs/unmute

A system that enables text LLMs to listen and speak by wrapping them in low-latency speech-to-text and text-to-speech models.

Episkey-G/GrokSearch-rs

GrokSearch‑rs is a Rust‑based MCP server that adds live, cited web‑search tools to LLM agents. It can run locally via stdio or be exposed as a multi‑tenant HTTP service. The server queries a configurable chain of providers (Tavily, Exa, TinyFish, Firecrawl) and includes key‑free specialist extractors for GitHub, StackExchange, arXiv, and Wikipedia. Features include response budgeting, caching, OAuth support for Grok, and a health‑check `doctor` tool. Install via npm (pre‑built binary) or Docker, configure with env vars or a global TOML file, and integrate with Claude‑style assistants through the MCP protocol.

LYL1015/JarvisHub

An open harness for canvas-native multimodal creative agents that uses an editable canvas as a shared project state for long-horizon creative work.

NVIDIA-AI-IOT/live-vlm-webui

A universal web interface for real-time interaction and benchmarking of Vision Language Models using live webcam or IP camera streams.

dnhkng/GLaDOS

A multi-modal, autonomous AI persona framework that creates a proactive voice assistant with vision, long-term memory, and a low-latency response pipeline.

Everless321/dYm

A desktop application for Douyin that combines watermark-free video downloading with AI-powered content analysis to automatically tag and summarize videos.

pipecat-ai/pipecat-examples

A collection of intermediate and advanced example applications for building voice and multimodal AI agents using the Pipecat framework.

meangrinch/MangaTranslator

A Gradio-based web application that automates manga and comic translation by detecting speech bubbles, cleaning original text, and rendering AI-translated text.

OpenTSLM/OpenTSLM

OpenTSLM is an open‑source library that adds native time‑series handling to Llama 3.2 and Gemma LLMs, enabling natural‑language prompting and reasoning over medical signals (ECG, EEG, accelerometer, etc.). It provides pre‑trained checkpoints, a simple `OpenTSLM` Python API, demo scripts for inference on several benchmark tasks, and a curriculum‑learning training pipeline covering MCQ, captioning, and chain‑of‑thought reasoning stages. Install via `pip install opentslm`, load models from Hugging‑Face, and fine‑tune or evaluate using the supplied scripts. Licensed under MIT.

livekit/agents-js

A Node.js framework for building real-time, multi-modal voice AI agents that can see, hear, and understand using a plugin-based architecture for STT, LLM, and TTS.

uezo/aiavatarkit

A modular framework for building low-latency conversational AI avatars with integrated VAD, STT, LLM, and TTS pipelines across multiple channels.

Nanako0129/pilotfish

pilotfish is a policy‑driven orchestration layer for Claude Code that automatically delegates low‑effort coding tasks to cheaper Claude models (Haiku, Sonnet) while keeping high‑level planning and final judgment in a main Opus session. It installs via a macOS/Linux plugin or a legacy global config, defines role agents in markdown, and uses a dispatch flow to route work based on interaction shape and risk. The project includes documentation, benchmarks, and a MIT license.

IvanWng97/pixtuoid

pixtuoid is a Rust‑based terminal UI that visualises multiple AI coding agents as animated pixel‑art coworkers in a multi‑floor office. It supports many agent CLIs (Claude Code, Codex, etc.), shows token usage, tool activity, and offers themes, weather, pets, and a lofi soundtrack—all locally with no telemetry.

Djdefrag/QualityScaler

A Windows AI upscaler that enhances, upscales, and de-noises images and videos locally using DirectX12 compatible GPUs.

Tencent-Hunyuan/HY-Motion-1.0

A series of billion-parameter text-to-3D human motion generation models that create skeleton-based animations from text prompts using Diffusion Transformers and Flow Matching.

huohua325/Memslides

MemSlides is an open‑source, hierarchical‑memory LLM agent framework that creates personalized slide decks and supports multi‑turn, localized revisions. It maintains user‑profile, working, and tool memories to keep preferences consistent and avoid re‑doing work, offering a CLI and Docker image for easy experimentation.

AHEKOT/ComfyUI_VNCCS_Utils

A collection of advanced ComfyUI utility nodes for 3D Gaussian Splatting, infinite-canvas image editing, and professional 3D character posing and lighting.

AnubhavChaturvedi-GitHub/jarvis-ai-assistant

A modular, voice-activated AI desktop assistant that combines speech recognition, LLM reasoning, image generation, and system automation to control your computer hands-free.

yangjian102621/geekai

A one-stop AI multimodal content creation platform that integrates text, image, audio, and video generation tools into a commercially ready workspace.

timmyy123/LLM-Hub

An open-source mobile app for Android and iOS that enables private, on-device LLM chat, image generation, and video generation using hardware acceleration.

resend/resend-mcp

Resend MCP Server is an open‑source Node.js package that implements the MCP interface, letting AI agents (Claude, Cursor, Copilot, etc.) manage the full Resend email platform—sending messages, handling contacts, templates, broadcasts, domains, webhooks, and more—via either a hosted remote server or a locally run stdio/HTTP server.

NickPittas/DirectorsConsole

A unified AI VFX production pipeline that ensures cinematographic accuracy in image and video generation by grounding prompts in real-world camera, lens, and lighting constraints.

tensorforger/FluxRT

A real-time stream editing pipeline for FLUX.2 that transforms webcam or video feeds with low latency and interactive prompt updates on consumer GPUs.

OpenGVLab/Ask-Anything

A family of multimodal large language models designed for chat-centric video and image understanding, enabling natural language conversations about visual content.

web-infra-dev/midscene-skills

A vision-driven automation library that enables natural-language control of UI across web, desktop, and mobile platforms using screenshots.

ZJUI-AI4H/Hulu-Med

Hulu‑Med is an open‑source, multimodal medical vision‑language model that handles text, 2‑D images, 3‑D scans, and videos. It offers several model sizes (4 B‑235 B), full training code, and achieves state‑of‑the‑art results on dozens of medical benchmarks. The repo provides easy installation, HuggingFace‑compatible loading, and optional vLLM serving for high‑throughput inference.

pytorch/benchmark

PyTorch Benchmarks is a collection of standardized, lightweight versions of popular deep‑learning models (BERT, ResNet, Stable Diffusion, etc.) that can be run to measure the performance of different PyTorch builds, CUDA versions, and back‑ends. It offers a uniform API, install scripts, and multiple runners (simple test, pytest‑benchmark, custom userbenchmark) plus utilities for low‑noise tuning on AWS g4dn.metal instances.

agentuniverse-ai/agentUniverse

agentUniverse is an Apache‑2.0 Python framework for building and orchestrating LLM‑powered multi‑agent applications. It offers ready‑made collaboration patterns (PEER, DOE), supports dozens of LLM providers, includes tools for domain knowledge injection, observability via OpenTelemetry, and a visual workflow UI. The library is used in real‑world financial products at Ant Group and provides extensive docs, sample apps, and a community on GitHub/Discord.

landing-ai/ade-cli

A command-line tool for agentic document extraction that converts complex documents into grounded Markdown and structured data with verifiable evidence.

meme-search/meme-search

A self-hosted meme search engine that uses AI vision models to index images by content and text for semantic and keyword retrieval.

tonyd2wild/DeepSeek-v4-Flash-Vision-Exp-DSpark-1M-NVFP4-KV-2x-DGX-Spark

A deployment recipe for DeepSeek-V4-Flash-Vision-Exp on DGX Spark clusters, enabling native vision support, 1M token context, and high-throughput speculative decoding via vLLM.

hawk86104/three-vue-tres

An AI-collaborative Web 3D engineering ecosystem built on Three.js and Vue 3 for creating production-ready digital twins and industrial visualizations.

pengchujin/jzsub

An automated tool that downloads high-quality videos from multiple platforms and uses GPT to generate and burn in bilingual subtitles.

thanhkeke97/RSTGameTranslation

A real-time screen translation tool for Windows games that uses OCR and LLMs to translate on-screen text and audio into the user's language.

OpenGVLab/InternVideo

A series of video foundation models and large-scale datasets designed for multimodal video understanding, long-context modeling, and contextual reasoning.

horang-labs/tessera

Tessera is a local desktop/web app that lets you run multiple Claude Code, Codex, or OpenCode AI coding agents in parallel, each in its own isolated Git worktree. It provides a Kanban board, split‑pane UI, chat view for terminals, full Git integration, and mobile remote access, all while using the provider CLIs already installed on your machine.

apple-aiml-research/ml-4m

A framework for training any-to-any multimodal foundation models that uses tokenization and masked modeling to scale across tens of modalities and tasks.

inclusionAI/UI-Venus

A general-purpose foundation GUI agent that unifies mobile, web, and desktop interaction through a closed-loop perception-reasoning-action system.

Anionex/dsh-vision-toolkit

A vision toolkit for DeepSeek Harness that gives text-only models visual capabilities for tasks like UI restoration, long-screenshot OCR, and GUI automation.

google/spatial-media

A collection of specifications and tools for 360-degree video and spatial audio to ensure immersive media is correctly recognized and played back.

yincongcyincong/MuseBot

MuseBot is an open‑source Go bot that connects Telegram, Discord, Slack, Lark, DingTalk, Enterprise WeChat, QQ and WeChat to a wide range of LLM APIs (OpenAI, DeepSeek, Gemini, OpenRouter, etc.). It streams AI replies, supports image/voice input, function calls, RAG, admin UI, metrics and cron jobs, and can be run locally or via Docker.

strnad/CrewAI-Studio

CrewAI Studio is a Streamlit‑based GUI that lets you design, run and export multi‑agent AI crews built on the CrewAI framework, supporting several LLM providers and custom tools, with easy installation via Conda, virtualenv, Docker or one‑click deployment.

facebookresearch/brain2qwerty

A system for decoding natural sentences from non-invasive brain recordings, enabling the reconstruction of text from brain activity during typing.

redis/mcp-redis

Redis MCP Server is a Python‑based, Docker‑ready server that lets AI agents interact with Redis via natural‑language commands. It implements the Model Content Protocol, exposing tools for every Redis data type (strings, hashes, lists, sets, streams, JSON, vector indexes, pub/sub, etc.) and supports Azure Entra ID authentication. Install from PyPI or Docker, configure via CLI flags or env‑vars, and plug it into Claude Desktop, OpenAI Agents SDK, or any MCP client.

wjq-learning/CBraMod

A foundation model for EEG decoding designed for clinical and Brain-Computer Interface (BCI) applications through a pretraining and fine-tuning pipeline.

awwaiid/ghostwriter

An AI assistant for reMarkable tablets that captures handwritten input via screenshots and responds by drawing or typing directly back onto the screen.

nvidia-cosmos/cosmos-transfer2.5

A multi-controlnet platform for Physical AI that transforms video modalities (like depth and segmentation) into high-fidelity training data for robots and autonomous vehicles.

bytedance/Lance

Lance is a 3B parameter unified multimodal model that integrates image and video understanding, generation, and editing into a single framework.