filippogiruzzi/voice_activity_detection
A deep learning-based Voice Activity Detection system using a 1D-ResNet and MFCC features to classify audio signals as speech or noise.
dictation-toolbox/dragonfly
A Python speech recognition framework that allows users to create custom voice commands to automate computer activities and program by voice.
algolia/voice-overlay-ios
An iOS library that provides a polished voice-to-text overlay UI, handling permissions and speech recognition using Apple's native SFSpeechRecognizer.
judahpaul16/gpt-home
A self-hosted AI home assistant for Raspberry Pi that uses LiteLLM and LangGraph to integrate LLMs with home automation and personal productivity tools.
sveinbjornt/hear
A command line interface for macOS that enables transcription of live microphone input and audio files using the system's built-in speech recognition.
Picovoice/rhino
Rhino is Picovoice’s on‑device speech‑to‑intent engine that converts spoken commands into structured intents and slots in real time. It runs locally on everything from micro‑controllers to phones, browsers, and desktops, supports multiple languages, and offers SDKs for Python, .NET, Java, Flutter, React Native, Android, iOS, Web, Node.js, and C.
openspeech-team/openspeech
A framework for building end-to-end automatic speech recognition systems, providing reference implementations of 20+ ASR models and recipes for multiple languages.
yeyupiaoling/MASR
A PyTorch-based automatic speech recognition framework that supports streaming and non-streaming inference across multiple model architectures and languages.
mybigday/whisper.rn
React Native bindings for on‑device Whisper (and NVIDIA Parakeet) speech‑to‑text, with GPU/Core ML acceleration, voice‑activity detection, and realtime streaming support.
TheStageAI/TheWhisper
A high-performance speech-to-text solution based on fine-tuned Whisper models, optimized for low-latency streaming and on-device inference on NVIDIA GPUs and Apple Silicon.
TensorSpeech/TensorFlowASR
A TensorFlow-based framework for Automatic Speech Recognition that implements various ASR architectures and supports TFLite conversion for efficient deployment.
ardha27/AI-Waifu-Vtuber
An AI-powered VTuber assistant that integrates speech recognition, LLMs, and text-to-speech to create an interactive virtual character for live streaming and personal use.
sooftware/conformer
A PyTorch implementation of the Conformer architecture that combines CNNs and Transformers to improve speech recognition by capturing both local and global audio dependencies.
modal-labs/quillman
A voice chat application powered by the Moshi speech-to-speech model, providing low-latency, bidirectional audio streaming for human-like interaction.
tensorflow/lingvo
A modular TensorFlow framework designed for building and scaling neural networks, particularly sequence-to-sequence models and giant language models.
HeyWillow/willow
Willow is a self-hostable inference server for fast language tasks including speech-to-text, text-to-speech, and LLM processing.
flashlight/wav2letter
wav2letter++ is an end-to-end automatic speech recognition framework that provides recipes and pre-trained models to implement state-of-the-art speech-to-text architectures.
Uberi/speech_recognition
A Python library that provides a unified interface for speech recognition, supporting multiple online and offline engines like OpenAI Whisper, Google Speech, and Vosk.
elevenlabs/elevenlabs-js
The official Node.js SDK for ElevenLabs, enabling developers to integrate lifelike AI voice synthesis, real-time audio streaming, and voice-powered AI agents into their applications.
wildminder/ComfyUI-VoxCPM
A ComfyUI custom node integration for VoxCPM, a tokenizer-free TTS system that enables expressive speech generation, natural language voice design, and advanced voice cloning.
codeforequity-at/botium-speech-processing
A unified API for open-source and cloud-based Speech-To-Text and Text-To-Speech services, simplifying audio processing for chatbots and voice applications.
tongjingqi/Thinking-with-Video
A new multimodal reasoning paradigm and benchmark (VideoThinkBench) that evaluates the ability of video generation models to solve complex visual and textual reasoning tasks.
AceDataCloud/Nexior
Nexior is an MIT‑licensed, Vue‑based app that aggregates 40+ chat, image, audio and video AI models into a single self‑hostable UI. It offers one‑click Vercel or Docker deployment, BYOK or a single AceData key, and includes built‑in user accounts, payments and referral mechanics, making it a ready‑made SaaS starter for AI products.
EvolvingLMMs-Lab/lmms-engine
A unified, high-performance training engine for multimodal models that integrates distributed training, kernel fusion, and advanced optimizers to enable scalable pretraining and fine-tuning.
awekrx/ChatGPT-MidJourney-prompt
A Python library that uses LLMs to automatically generate detailed and creative prompts for MidJourney images based on simple text hints.
HorizonWind2004/reconstruction-alignment
A self-supervised post-training method for unified multimodal models that improves image generation and image editing by training models to reconstruct images from their own visual features.
shiimizu/ComfyUI_smZNodes
Custom ComfyUI nodes (CLIP Text Encode++ and Settings) that replicate AUTOMATIC1111's prompt parsing and embedding, enabling identical image generation between stable‑diffusion‑webui and ComfyUI.
biagiomaf/smart-comfyui-gallery
An open-source digital asset manager and gallery built for ComfyUI and AI production libraries to organize, search, and remix generations.
Windsander/ADI-Stable-Diffusion
A C++ library and CLI tool that enables high-performance, Python-free inference for Stable Diffusion models using ONNXRuntime across multiple platforms.
Licoy/GoAmzAI
A private, multimodal AIGC platform for individuals and enterprises that integrates text, image, video, music, and PPT generation into a unified management system.
mrhan1993/Fooocus-API
A FastAPI-powered REST API for Fooocus that allows developers to programmatically generate high-quality images without manual parameter tweaking.
Acly/comfyui-tooling-nodes
A set of ComfyUI custom nodes and API extensions that enable the use of ComfyUI as a backend for external tools, featuring optimized image transfer, regional prompting, and model inspection.
melMass/comfy_mtb
A collection of custom nodes for ComfyUI that extends the image generation workflow with additional utility tools.
a616567126/GPT-WEB-JAVA
A Java-based web platform that integrates GPT, Midjourney, and Stable Diffusion into a commercial AI service with user management and payment systems.
vual/ChatGPT-Next-Web-Pro
ChatGPT‑Next‑Web‑Pro is an open‑source, self‑hosted web UI for OpenAI‑style LLMs that adds multimodal (image, audio, video) generation, file parsing, S3 storage, and an optional backend with user accounts, subscription plans, and admin panels. Deployable via a single Docker run (no‑backend) or a Docker‑Compose stack (with‑backend).
basetenlabs/truss
A CLI tool for packaging and deploying AI/ML models to production, automating containerization, GPU configuration, and dependency management.
xxxily/hello-ai
An AI-driven navigation hub that uses autonomous agents to discover, evaluate, and organize high-quality open-source AI projects from GitHub into an auto-updating directory.
lobehub/sd-webui-lobe-theme
A modern, highly customizable interface framework for Stable Diffusion WebUI that improves the workflow and visual experience of AI image generation.
mikekeith52/scalecast
A time series forecasting library that provides a uniform interface for ML and deep learning models, simplifying model selection, validation, and pipeline creation.
EternalEvan/Astra
Astra is an interactive world model that uses an autoregressive diffusion transformer to generate realistic, action-conditioned long-horizon video predictions.
Ma-Zhuang/OmniNWM
OmniNWM is a research codebase that builds a panoramic, multi‑modal world model for autonomous‑driving simulation. It jointly generates RGB, semantic, depth, and 3‑D occupancy maps from a vehicle trajectory, uses normalized Plücker ray‑maps for precise control, and provides occupancy‑based dense rewards for closed‑loop policy evaluation. The repo includes installation steps (including a manual patch to `transformers`), data preparation instructions for nuScenes, pretrained checkpoint download links, and scripts for inference, out‑of‑distribution testing, and staged training.
forloopcodes/contextplus
An MCP server that turns large codebases into searchable, hierarchical feature graphs using RAG, AST parsing, and spectral clustering for high-accuracy AI coding context.
opensumi/core
A framework for quickly building AI-native IDEs, supporting cloud, desktop, and web-based development environments.
Pimzino/spec-workflow-mcp
An MCP server that enables structured, spec-driven development for AI agents by enforcing a sequential workflow of requirements, design, and tasks with a real-time monitoring dashboard.
VibiumDev/vibium
A verification layer for AI coding agents that provides browser automation tools to navigate, interact with, and verify web-based tasks via CLI, MCP, or client libraries.
KlaatAI/klaatcode
Klaat Code is a terminal‑based AI coding assistant that talks to the hosted Klaatu‑o1 router. The router automatically selects the cheapest model tier for each request and escalates only when needed, cutting costs dramatically. The client indexes your project into a semantic call‑graph, runs tool calls (file edits, shell commands, web searches) for free, and manages context with smart compaction and verification. Features include a rich UI, slash commands, Git integration, sub‑agent delegation, extensive safety/permission checks, and reproducible benchmarks showing $0.027 per solved task at 23 s median time.
hengruiyun/AI-Stock-Master
AI Stock Master is a Windows/macOS desktop app that blends quantitative stock‑analysis algorithms (RTSI, TMA, MSCI) with a locally‑run LLM (Mini Ollama) to generate natural‑language market reports and trading signals for Chinese, Hong‑Kong and US equities. It’s a research‑oriented tool, not a commercial trading platform.
24kchengYe/MemoMind
MemoMind is a locally hosted, GPU‑accelerated memory system for AI coding agents. It stores extracted facts from chats, documents, and daily life events in PostgreSQL + pgvector, builds a knowledge graph, and offers fast 4‑way hybrid retrieval (semantic, BM25, graph, temporal). The agent can retain new information, recall relevant memories, and reflect across the whole store. All data stays on the user’s machine, works with any OpenAI‑compatible LLM, and includes a web dashboard for browsing and exporting memories.