altic-dev/FluidVoice
An open-source voice-to-text dictation app for macOS that enables on-device AI transcription and system control via voice commands.
index-tts/index-tts
A zero-shot text-to-speech system that clones voices from a single audio clip with fine-grained control over emotion, speed, and pronunciation across five languages.
fastapi/fastapi
FastAPI is a high‑performance, type‑hint‑driven Python web framework for building APIs. It provides automatic validation, serialization, and interactive OpenAPI docs, while being fast to code and production ready.
bethington/ghidra-mcp
Ghidra MCP Server is a production‑ready bridge that exposes Ghidra’s full reverse‑engineering engine (decompilation, P‑code emulation, live debugging, structure creation, etc.) via the Model Context Protocol. It provides 253 read/write tools, batch‑optimized operations, AI‑focused documentation workflows, and supports headless, Docker, or GUI deployment with stdio, HTTP, or SSE transports.
OpenByteInc/QuantDinger
QuantDinger is an open‑source, self‑hosted AI‑augmented trading operating system. It lets you write Python strategies, back‑test them, run paper‑trading simulations, and optionally execute live orders on crypto exchanges and traditional brokers. The platform integrates multiple LLM providers for market research, offers a Flask‑Gunicorn API, Docker‑compose deployment, optional Prometheus/Grafana observability, and a secure token‑based Agent Gateway for AI‑driven automation.
Significant-Gravitas/AutoGPT
AutoGPT is an open‑source platform for building, deploying, and managing autonomous AI agents. Users can describe a goal in plain English or assemble a workflow visually; the system creates an LLM‑driven agent that interacts with 45+ integrations (Gmail, Slack, GitHub, etc.) to complete tasks such as drafting reports, triaging incidents, or generating marketing copy. The project offers a paid hosted service and a free self‑hostable Docker appliance, with the core platform under a Polyform Shield license and the classic agent code under MIT.
stemdeckapp/stemdeck
A free, local audio stem separation tool that splits songs into isolated tracks like vocals and drums using Demucs, providing a DAW-style mixer for local playback and export.
QwenAudio/qwen-audio-agent
A realtime voice runtime that allows AI agents to maintain a continuous conversation with the user while simultaneously executing complex background tasks.
Gentleman-Programming/gentle-shell
gentle-shell™ is an npm‑distributed terminal UI that runs a coding‑agent (LLM‑backed) inside the Pi development environment. It provides a unified workspace showing tasks, changes, and agent hierarchy, supports lightweight Organic Driven Development (ODD) with optional formal Specification‑Driven Development (SDD), and offers native review of concrete diffs before they are applied. Extensible via optional companion packages, it is installed with `pi install npm:gentle-pi@2.6.0` and launched via `pi`.
risa-labs-inc/BossConsole
BossConsole (BOSS) is an open‑source, cross‑platform desktop harness for AI agents. Built with Kotlin Multiplatform and the JVM, it provides a native editor, embedded Fluck browser, shareable terminal, Git/Docker/K8s tools, and a plug‑in “Toolbox”. All features are exposed as MCP tools (`mcp__boss__*`) that any model (Claude Code, Codex, Gemini, OpenCode, etc.) can call, with fine‑grained RBAC, per‑tool kill‑switches, and signed plug‑ins. Hot‑reload lets agents evolve new tools at runtime. Install via Homebrew, script, or OS packages; licensed Apache‑2.0.
MengTo/Skills
A Markdown‑based library of 123 "skills"—step‑by‑step, versioned prompt/workflow recipes—for AI coding agents (Codex, Claude, Cursor, etc.). Each skill lives in `agent-skills/<category>/<skill>/SKILL.md` and may include demos, assets, and reference links. The repo lets developers treat prompts as reusable code assets, enabling repeatable UI design, web‑GL, Three.js game, and media‑generation workflows. MIT‑licensed and ready to be loaded as context by any LLM‑driven development tool.
kirodotdev/KiroCrew
Kiro Crew is an open‑source, on‑premise AI‑agent platform that keeps agent state (memory, lessons, skills) alive across restarts, lets you run long‑running or scheduled tasks unattended, and is reachable via a desktop app, web dashboard, CLI, or chat integrations (Slack, Discord, etc.). Install with a one‑line script, Docker, or the native desktop packages; the Gateway process stores all data locally and enforces sandboxing and approval policies.
modelscope/FunASR
An industrial speech recognition toolkit for offline, streaming, and edge deployment, providing a unified pipeline for ASR, VAD, punctuation, and speaker diarization.
zubair-trabzada/geo-seo-claude
A Claude Code skill that audits a website for AI‑search (GEO) visibility, scoring citability, brand signals, content quality, technical SEO and schema, then generates markdown or PDF client reports.
thewh1teagle/vibe
A private, offline audio and video transcription tool that uses AI models like Whisper to convert speech to text locally on your device.
Gentleman-Programming/engram
Engram is a single‑binary Go program that gives AI coding agents a persistent, searchable memory store (SQLite + FTS5). It works with any MCP‑compatible agent (Claude Code, Gemini CLI, VS Code Copilot, Cursor, etc.) via CLI, HTTP API, MCP, or a terminal UI. The tool is zero‑dependency, supports local‑first storage with optional cloud replication, and provides a rich set of memory commands (`mem_save`, `mem_search`, `mem_context`, `mem_session_summary`, …) to let agents remember decisions, bug fixes, and design rationales across sessions and teams.
modelbus/one-api-pro
OneAPI Pro is an open‑source, Go‑based AI API gateway with a Vue 3 admin UI. It unifies access to many LLM providers, adds token/role management, subscription billing, usage dashboards, and supports decentralized multi‑node clustering. Deployable as a static binary or Docker image.
ggml-org/whisper.cpp
whisper.cpp is a pure C/C++ implementation of OpenAI’s Whisper speech‑to‑text model. It runs offline on CPUs and a wide range of accelerators (Apple Silicon, NVIDIA CUDA, AMD ROCm, Intel OpenVINO, etc.), supports mixed‑precision and quantized models, and provides a simple C‑API plus language bindings for easy integration across desktop, mobile, embedded, and web platforms.
echo-loop/Echo-Loop
An AI-powered English listening and speaking training app that automates a scientific learning loop of intensive listening, shadowing, and retelling with spaced repetition.
ace-step/ACE-Step-1.5
ACE‑Step 1.5 is an open‑source music‑generation foundation model that combines a language‑model planner with a Diffusion‑Transformer audio decoder. It can create full songs (10 s‑10 min) in under 10 seconds on a RTX 3090, runs with <4 GB VRAM for base models, and supports LoRA fine‑tuning from a handful of tracks. The repo provides Gradio and REST interfaces, launch scripts for all major OSes, multiple model sizes (including a 4 B XL version), and extensive control over lyrics, style, tempo, and instrumentation.
michael-denyer/pstack-claude
pstack‑claude is a MIT‑licensed plugin that ships 54 reusable “Agent Skills” (e.g., TDD, CI fixing, PR summarising) for Claude Code and, via a shared skill directory, for Codex, Prime Agent, opencode and Gemini‑CLI. It auto‑injects a poteto‑mode mandate on Claude sessions, provides generated slash‑command stubs for Codex, and includes CI to keep the skill set in sync with the upstream Cursor version.
lightpanda-io/browser
Lightpanda Browser is a lightweight, Zig‑written headless browser built for AI agents and large‑scale web automation. It runs a full V8 JavaScript engine, offers CDP and WebDriver BiDi interfaces, and includes an “agent” mode that lets LLMs control the browser via plain‑English tasks, recording deterministic JavaScript scripts. Benchmarks claim ~16× lower memory and ~9× faster page loads than headless Chrome. Installable via Homebrew, AUR, direct binaries, or Docker, it supports Linux/macOS (WSL2 for Windows) and integrates with major LLM providers. The project is MIT‑licensed, actively tested (unit, end‑to‑end, Web Platform Tests), and backed by a community Discord.
m-bain/whisperX
A fast automatic speech recognition system that enhances OpenAI's Whisper with batched inference, word-level timestamps, and speaker diarization.
rafal-qa/slopo
Slopo is a Python‑based CLI that uses code‑embedding models (Jina AI, Voyage AI, or local Ollama) to find semantically similar code fragments across a repository. It indexes source files, computes vector embeddings, clusters similar snippets, and outputs Markdown or compact reports for human review or AI coding agents. Installation is via `uv tool install slopo`; configuration is done with `slopo init`. The tool targets hidden, non‑exact code duplication to aid refactoring and reduce AI token usage.
NVIDIA/skills
NVIDIA Agent Skills is a catalog of portable instruction sets (skills) that teach AI coding agents (Claude Code, Codex, Cursor, etc.) how to install, configure, and use NVIDIA software such as cuDF, cuOpt, DeepStream, Jetson, NeMo, DOCA, and many others. Skills are installed with the open‑source `skills` CLI (`npx skills add nvidia/skills …`) and can be targeted to specific agents. The repo mirrors skill definitions from product repos and is kept up‑to‑date via an automated sync pipeline.
OpenMinis/OpenMinis
OpenMinis is an open‑source mobile AI agent (iOS + Android) that lets you run any LLM (Claude, GPT, Gemini, etc.) on‑device. It ships a sandboxed Alpine Linux shell (iSH on iOS, PRoot on Android) so the agent can install packages, run scripts, and manipulate files, while also exposing native device services (Health, Calendar, HomeKit, Bluetooth, clipboard, media, etc.) as tools. Users create reusable “skills” (folders with a `SKILL.md` and optional scripts) and organise work into separate workspaces. Typical uses include nutrition logging, morning briefings, task extraction from chats, syncing with Obsidian, and share‑to‑event automation. The codebase is Swift/SwiftUI for iOS, Kotlin/Compose for Android, and is GPL‑v3 licensed.
Gnosil/semantix
Semantix is a Go‑based CLI that adds cross‑session memory to LLM coding agents. It extracts reusable “slices” from finished sessions, stores them locally, and injects matching slices as byte‑stable prefixes to improve provider prompt‑cache hit rates and cut inference cost. It can run as a full agent (`semantix‑agent`) or as a kernel attached to existing agents via tool registration, middleware, or a gateway. The project reports high cache hit rates (up to 99.8 %) and ~80 % cost savings in demos, and integrates with Claude Code, LangChain, and any OpenAI‑compatible client.
MaxHu-xuan/chat-archive-guard
ChatArchiveGuard is a Python‑only CLI that locally scans chat export files (text, JSON/JSONL, SQLite) for secret/PII patterns, format errors, and SQLite integrity, reporting aggregate counts and coverage flags without modifying or uploading data.
trailofbits/skills
Trail of Bits Skills is a Claude Code/Codex plugin marketplace that adds dozens of security‑focused “skills” (e.g., static analysis, smart‑contract auditing, YARA rule authoring, mutation testing). Install it with a single `/plugin marketplace add trailofbits/skills` command, then invoke individual plugins from the LLM to run concrete analyses and get structured results.
cubeplexai/cubeplex
CubePlex is an open‑source, cloud‑native platform that lets teams run managed AI agents in isolated, persistent workspaces. It offers multi‑model chat, reusable skills, scoped memory, tool connectors, IM integrations, and built‑in governance, and can be deployed with Docker Compose or Helm/Kubernetes.
coldteadotai/pr-lens
PR Lens is a GitHub‑integrated tool that turns a pull‑request diff into animated architecture and data‑flow diagrams. It can be used via a GitHub App, a CI Action, a local CLI, or a coding‑agent skill, and is configurable through a `.github/pr‑lens.yml` file. The visual output helps reviewers understand the scope and impact of a change before diving into code.
k2-fsa/sherpa-onnx
A portable framework for running speech-to-text, text-to-speech, and other audio AI functions locally across diverse hardware and operating systems.
vshulcz/deja-vu
deja‑vu is a local‑only memory layer for code‑generation agents (Claude Code, Codex, Cursor, Copilot, etc.). It indexes the transcript files those agents already write, builds a fast searchable index, and automatically injects relevant past decisions into any agent session. Features include retroactive natural‑language search, cross‑agent recall, compaction‑aware persistence, state tracking, staleness warnings, sync/handoff between machines, and automatic redaction of secrets. Install via a one‑line script, Homebrew, Scoop, Go, or NPM; `deja install --auto` wires a tiny MCP server into detected agents. Use commands like `deja "jwt refresh token"`, `deja wip`, `deja blame <file>`, `deja sync ssh <host>`, and optional semantic embedding via a local LLM. All processing stays on the machine, with built‑in privacy safeguards.
Javis603/token-monitor
Token Monitor is a cross‑platform desktop widget that reads local logs and API‑key balances from 35+ AI coding assistants, shows live token usage, cost, and quota limits, and can sync those summaries across multiple machines via a self‑hosted or Cloudflare hub. It offers per‑session breakdowns, historical heatmaps, export to CSV/JSON, and a highly customizable UI, all while keeping raw prompt data private.
chidiwilliams/buzz
Buzz is an offline desktop app that transcribes and translates audio/video (or YouTube links) using OpenAI’s Whisper. It supports live microphone captioning, speaker identification, speech‑separation, multiple hardware‑accelerated back‑ends (CUDA, Apple‑silicon, Vulkan), export to subtitle formats, a searchable viewer, folder‑watch automation, CLI, and a plugin system. Install via DMG (macOS), installer (Windows), Flatpak/Snap/AppImage (Linux), or pip (Python). MIT‑licensed.
Arize-ai/phoenix
Arize Phoenix is an open‑source, self‑hosted AI observability platform that lets you trace, evaluate, experiment on, and debug LLM applications. It works with many frameworks (LangChain, OpenAI Agents, Claude, etc.) and providers (OpenAI, Anthropic, Bedrock, Google GenAI, etc.) via OpenTelemetry/OpenInference instrumentation.
HAKORADev/VODER
VODER is a local, offline, professional-grade voice processing toolkit that combines 8 modes—speech-to-text, text-to-speech, voice conversion, music generation, speech enhancement, sound effects, vocal separation, and speaker diarization—plus a Project Eva DLC for image/video/3D generation and uncensored chat. It runs entirely on your machine, works with or without a GPU, and uses open-source models like Whisper, Qwen3-TTS, Seed-VC, and more.
resemble-ai/chatterbox
A family of open-source text-to-speech models providing low-latency voice cloning and multilingual speech generation across various model sizes for GPU and CPU deployment.
XiaoMaColtAI/math-modeling-skill
An open‑source AI‑agent skill that structures math‑modeling competition work into three stages (analysis, coding, paper), supports Python/MATLAB, generates publication‑grade figures, enforces reproducibility, and outputs Word/LaTeX papers with built‑in quality‑check subagents. Includes a DeepSeek Harness plugin and a full test suite.
huggingface/speech-to-speech
A low-latency, modular voice-agent pipeline that coordinates VAD, STT, LLM, and TTS components to enable real-time voice conversations.
makecindy/cindy
Cindy is an open‑source, cross‑platform AI‑agent client (Electron desktop + React‑Native mobile) that integrates multiple LLM harnesses (Claude Code, Codex, etc.) with local file access, browser/OS control, persistent memory, and extensible “skills”. The repo contains the client code, shared TypeScript packages, and tooling to build/run the apps; the backend service lives elsewhere. It supports both cloud‑backed usage (via a Cindy account) and fully local agents (skip sign‑in). Licensed under Apache‑2.0.
QwenAudio/CosyVoice
CosyVoice is an advanced, LLM-based text-to-speech system designed for high-quality, zero-shot multilingual speech synthesis and voice cloning.
yuanbw2025/storyforge
StoryForge is an open‑source, browser‑based AI‑assisted storytelling suite. It lets you write long‑form or short novels, convert them into scripts or comics, and turn worlds into playable experiences (run‑table, chat, text‑adventure). All data lives locally (IndexedDB) and you bring your own LLM API key. The platform includes a version‑aware “Harness” layer that records every AI call, manages memory retrieval, and enforces author‑confirmed adoption, ensuring long‑term consistency across chapters and derived products.
visualbruno/3DGenStudio
3D Gen Studio is an open‑source, AI‑driven desktop/web app that lets you build 3‑D assets end‑to‑end. It combines a Kanban board or node‑graph UI with ComfyUI workflows and external 3‑D generation APIs, offering image‑to‑mesh pipelines, auto‑UV/retopo, AI‑auto‑rigging (SkinTokens + optional Kimodo motion), LOD/optimization, and game‑ready checks—all stored locally on disk.
breezeblue-ai/breeze-tts
Breeze TTS 2 is an open-weight, bilingual text-to-speech model designed for real-time interaction, featuring natural-language voice design and ultra-low latency streaming.
leo-lilinxiao/codex-autoresearch
Codex‑autoresearch is a Git‑backed Codex skill that runs an autonomous loop: propose a single code change, commit it, measure a numeric metric (e.g., test failures), keep the change if it improves the metric and passes optional guards, otherwise revert. It repeats until a user‑specified target is reached, storing an immutable event log and HTML report. Supports foreground or background execution, works with any measurable outcome, and enforces strict safety (all trials are real commits, failures abort with clear errors).
liliu-z/stashbase
StashBase is a local‑first desktop app that builds a linked Markdown wiki from your personal files (PDFs, docs, images, recordings) and lets AI agents use that wiki as context via chat. It offers semantic search, OCR/transcription, and integrates with OpenQuill, Claude Code, Codex, or any MCP‑compatible client, while keeping all original files on your machine.
OpenMOSS/MOSS-Transcribe-Diarize
An end-to-end audio understanding model that jointly performs speech transcription and speaker diarization for long-form multi-speaker audio in a single pass.