ogulcancelik/pi-extensions

A set of npm‑based plug‑ins that extend the terminal‑based coding assistant *Pi* with UI footers, session search, auto‑permissions, sub‑agents, worktree handling and other utilities.

asklokesh/loki-mode

Loki Mode is an open‑source CLI/SDK that turns a product spec into a fully built, tested, and verified codebase. It runs an autonomous LLM‑driven workflow (Reason‑Act‑Reflect‑Verify) with eight quality‑gate checks, and emits a cryptographically‑bound “Evidence Receipt” that anyone can re‑verify. Install via bun/npm/Homebrew/Docker, then use `loki quickstart` for new apps or `loki modernize heal` to assess and improve existing repos. The tool requires an Anthropic API key (or compatible model) and focuses on auditable, production‑grade code generation.

kaixxx/noScribe

A free, open-source application for local audio transcription and speaker identification, designed for researchers and journalists to ensure data privacy.

knowsuchagency/mcp2cli

mcp2cli is a runtime CLI generator that turns MCP servers, OpenAPI specs, or GraphQL endpoints into instant command‑line tools. It supports OAuth, secret‑safe flags, baked configurations, usage‑aware ranking, and machine‑readable JSON output, enabling LLM agents and developers to call APIs with minimal token overhead.

ken107/read-aloud

A browser extension for Chrome and Firefox that uses text-to-speech technology to convert webpage text into audio for improved accessibility and multitasking.

plasma-umass/scalene

Scalene is a fast, line‑level Python profiler that measures CPU, GPU (NVIDIA) and memory usage, separates Python from native code, visualizes stitched stacks, and can ask LLMs (OpenAI, Azure, Bedrock, Ollama, etc.) for optimization suggestions directly from the UI.

ZaxbyHub/opencode-swarm

OpenCode Swarm is an OpenCode plugin that orchestrates a team of specialised AI agents (architect, coder, reviewer, tester, security critic, etc.) to turn a single AI coding prompt into a gated, production‑ready workflow. It enforces review, testing, security scans and documentation before any code is merged, supports 13 languages, offers resumable sessions, PR monitoring, and multiple safety/ speed modes. Install with `bunx opencode‑swarm install`; configure via `~/.config/opencode/opencode‑swarm.json`; use slash commands like `/swarm agents`, `/swarm status`, `/swarm auto‑proceed` to control the process.

librosa/librosa

A Python library for audio and music signal processing that provides the foundational tools for building music information retrieval systems.

SamadhiFire/xinqingnian-maoxuan-skill

A Claude/Codex skill that guides an LLM through a four‑step “Mao‑style” analysis of stuck projects, team conflicts, or life decisions, outputting a concise text summary and an optional single‑file HTML report.

TaoLiveAIGC/TaoMate

TaoMate is a research‑grade, real‑time audio‑video digital‑human generation system. The repo provides inference code, multi‑GPU launchers, and a four‑GPU interactive browser demo. It stitches together a 22 B LTX‑2.3 transformer, a Gemma‑3 text encoder, and optional Gemma‑4 dialogue model, requiring high‑memory NVIDIA GPUs and separate model checkpoints from Hugging Face.

diodiogod/TTS-Audio-Suite

TTS Audio Suite is a ComfyUI extension that unifies 19 text‑to‑speech, voice‑conversion and audio‑editing engines (e.g., F5‑TTS, DramaBox, Higgs Audio, Qwen3‑TTS, RVC). It offers subtitle‑aware generation, per‑segment tags, visual waveform analysis, silent‑speech timing, and built‑in RVC model training, all with runtime isolation for different engine dependencies.

pytorch/audio

A PyTorch-based audio library that provides tools for processing audio data for machine learning, featuring GPU acceleration and specialized audio transforms.

estebanstifli/LocalText2Voice

A desktop app for creating long-form audiobooks and podcasts using local or cloud AI text-to-speech engines, featuring built-in quality verification and audio mixing.

sdatkinson/NeuralAmpModelerPlugin

A VST3/AudioUnit plug-in that integrates the Neural Amp Modeler engine into DAWs for AI-driven guitar amplifier modeling.

genomoncology/biomcp

BioMCP is a unified CLI/MCP server that lets users and AI agents query ~30 biomedical databases (PubMed, ClinVar, ClinicalTrials.gov, OncoKB, etc.) with a simple command grammar, pivot between entities, run local study analytics, and perform gene‑set enrichment. Installable via script, uv/pip, Homebrew, Docker, or source, it provides provenance‑rich JSON or terminal output and integrates directly with Claude Code, Codex, and Claude Desktop.

vllm-project/vime

Vime is an open‑source RL post‑training framework for large language models. It combines the slime/Megatron distributed training stack with vLLM inference (via vllm‑router) to provide high‑throughput rollouts and flexible data‑generation pipelines. Supported models include Qwen, DeepSeek V3, and LLaMA 3. Users can run agentic RL, coding‑agent RL, or custom RL algorithms by configuring Megatron, vLLM, and router arguments. The repo offers quick‑start docs, several ready‑made examples, and a clear code‑reading path for developers.

tile-ai/tilelang-ascend

TileLang‑Ascend is a Python‑style DSL built on TVM for writing high‑performance AI kernels on Huawei Ascend NPUs. It provides a compiler that emits Ascend‑specific code, supports automatic vectorization, memory planning, and synchronization, and ships as a pip‑installable wheel with many ready‑made examples (GEMM, Flash‑Attention, etc.).

matrixorigin/memoria

Memoria is a Git‑style version‑control layer for AI‑agent memory, built on the MatrixOne database. It offers snapshots, branches, merges, and rollbacks for every memory mutation, hybrid vector/full‑text search, automatic contradiction detection, and a privacy‑first deployment model (cloud or self‑hosted). Agents interact via MCP commands or a REST API, making memory safe to change, auditable, and reusable across sessions.

SUC-DriverOld/MSST-WebUI

A web-based interface for Music-Source-Separation-Training that allows users to isolate vocals and instruments from music tracks using various AI models.

bolna-ai/bolna

Bol na is an open‑source framework that orchestrates speech‑to‑text, LLM, and text‑to‑speech services to build voice‑first assistants that can place and receive phone calls. It ships Docker‑compose containers for the core server, telephony webhook (Twilio/Plivo), Redis, and ngrok, and provides a Python SDK (`Assistant`) for building streaming pipelines. The platform is provider‑agnostic (supports Deepgram, Azure, OpenAI, ElevenLabs, AWS Polly, etc.) and is extensible to new telephony or audio services.

royshil/obs-localvocal

An OBS plugin that provides real-time, local speech-to-text transcription and translation using OpenAI's Whisper, ensuring privacy and no cloud costs.

DeadWaveWave/opencove

OpenCove is a cross‑platform Electron app that provides an infinite 2‑D canvas where you can place terminals running AI coding agents (Claude Code, Codex), notes, tasks, and images. The layout, terminal output, and agent state persist across restarts, and you can snapshot canvases for later reuse. It aims to make AI‑assistant workflows spatial rather than hidden in tabs or long chat logs.

elevenlabs/elevenlabs-python

The official Python SDK for ElevenLabs, enabling developers to integrate lifelike AI text-to-speech, voice cloning, and real-time conversational AI agents into their applications.

prathoshap/vagdhenu

A production-grade Sanskrit chant text-to-speech system that generates traditional melodic recitation instead of flat read-aloud speech.

spotify/pedalboard

A high-performance Python library for reading, writing, and applying studio-quality audio effects and third-party VST3/Audio Unit plugins to audio files and live streams.

LearnPrompt/humanize-ppt

Humanize PPT is an open‑source Python orchestration tool that turns markdown into an audience‑state‑transfer (AST) slide plan, decides per‑slide media, hands off rendering to downstream HTML or PPTX skills, runs an automatic presentation‑check to flag “pretty‑but‑can’t‑talk” slides, and generates a lightweight presenter shell. It integrates with AI agents (Claude‑Code, Codex, Hermes) via a simple skill install, works with downstream renderers like `guizang-ppt-skill`, `frontend-slides`, `ppt‑master`, and media generators (`baoyu-image-gen`, Remotion). MIT‑licensed.

flybirdxx/ComfyUI-Qwen-TTS

ComfyUI custom nodes for high-quality speech synthesis, zero-shot voice cloning, and voice design based on the Qwen3-TTS model.

mbailey/voicemode

A voice interface for Claude Code and other MCP-capable agents that enables natural, hands-free conversations using either cloud or local speech services.

Dicklesworthstone/coding_agent_session_search

A Rust‑based terminal UI and CLI that aggregates, normalises, and indexes the conversation logs of dozens of AI coding assistants into a local SQLite archive. It provides sub‑60 ms lexical search and optional on‑device MiniLM semantic search, with a stable JSON robot‑mode API for automation and diagnostics.

KoljaB/RealtimeSTT

A Python speech-to-text library that provides fast, real-time transcription with integrated voice activity detection and wake word support.

di-sukharev/opencommit

OpenCommit is a Node‑JS CLI that uses LLMs (OpenAI, Anthropic, Ollama, etc.) to auto‑generate conventional, emoji‑enhanced Git commit messages from staged changes. It supports global and per‑repo configuration, multiple providers, local model servers, a Git hook, and a GitHub Action for automatic commit‑message improvement.

richardr1126/openreader

An open-source, self-hosted text-to-speech document reader that provides synchronized read-along playback for EPUB, PDF, TXT, MD, and DOCX files.

OpenMOSS/MOSS-Audio

MOSS-Audio is an open-source audio understanding model that unifies speech, environmental sound, and music analysis with strong temporal awareness and reasoning capabilities.

kigner/audio.cpp-webui

A high-performance C++ audio inference framework based on ggml that enables fast, portable local execution of TTS, ASR, and voice cloning models across multiple hardware backends.

riba2534/happyclaw

HappyClaw is a self‑hosted platform that runs Anthropic Claude Code agents as long‑living, multi‑user services. It wraps the Claude Agent SDK (TypeScript) and Claude Code CLI, offering web‑based and eight instant‑messaging channel interfaces, isolated workspaces (host or Docker), full Claude Code capabilities (file access, shell, browser, plugins), scheduling, usage tracking, and enterprise‑grade RBAC and backup.

hkjarral/AVA-AI-Voice-Agent-for-Asterisk

An open-source AI voice agent for Asterisk and FreePBX that provides a modular pipeline for STT, LLM, and TTS providers to create production-ready voice bots.

magenta/magenta-realtime

An open-weights model and inference engine for real-time music generation, optimized for streaming audio on Apple Silicon.

Eddycrack864/UVR5-UI

A user-friendly graphical interface for UVR5 (via python-audio-separator) that allows users to separate vocals and instruments from audio and video files using various AI models.

f-is-h/Usage4Claude

Usage4Claude is a native macOS menu‑bar app that monitors Anthropic Claude (and optional OpenAI Codex) subscription quotas in real time. It supports all Claude products, multiple accounts, adaptive refresh, colour‑coded usage rings, notifications, and stores credentials securely in the macOS Keychain. Install via a pre‑built DMG or build from source with Swift/SwiftUI.

andimarafioti/faster-qwen3-tts

A high-performance inference engine for Qwen3-TTS that uses CUDA graph capture to enable real-time, low-latency voice cloning and speech generation.

FireRedTeam/FireRedASR2S

An industrial-grade all-in-one ASR system that combines speech recognition, voice activity detection, language identification, and punctuation prediction, with specialized support for Chinese dialects.

JordyZomer/lemmalog

Lemmalog is a Rust‑based Datalog engine that turns LLM‑extracted triples into a provably correct, incrementally updatable memory store. It supports confidence‑weighted facts, temporal validity, entity canonicalization, proof‑tree generation, and a hybrid BM25‑plus‑entity retrieval layer. Integrated via a CLI, REPL, or JSON‑RPC MCP server, it powers Claude Code/Kimi agents and scores competitively on MemEval‑style benchmarks while using far fewer tokens than full‑context baselines.

sdatkinson/NeuralAmpModelerCore

A core C++ DSP library that provides the neural network logic needed to run neural amplifier models for audio plugins.

langchain-ai/social-media-agent

An LLM‑driven agent that scrapes a URL, drafts Twitter/LinkedIn posts, lets a human review them, and then publishes via Arcade or native OAuth. Built on LangGraph, it supports optional Slack, Supabase, GitHub, and YouTube integrations and is configurable through prompt files.

speechbrain/speechbrain

An open-source PyTorch toolkit for Conversational AI that provides recipes and tools for speech and text processing, including speech recognition and synthesis.

Kieirra/murmure

Murmure is an open‑source desktop app that provides offline, privacy‑first speech‑to‑text using NVIDIA’s Parakeet TDT 0.6B v3 model. It runs locally on Windows, macOS, and Linux, supports 25 European languages, and offers a simple push‑to‑talk UI with no telemetry.

stemrollerapp/stemroller

StemRoller is a free application that uses the Demucs algorithm to separate vocals and instrumental stems from songs found via integrated YouTube search.

azuwis/pianotrans

A GUI and packaging tool for ByteDance's piano transcription system that converts piano audio and video recordings into MIDI files with pedal data.