fastapi/fastapi
FastAPI is a high‑performance, type‑hint‑driven Python web framework for building APIs. It provides automatic validation, serialization, and interactive OpenAPI docs, while being fast to code and production ready.
jnMetaCode/agency-agents-zh
A Chinese‑language collection of 276 ready‑to‑use AI expert agents (personas, workflows, deliverables) covering 20 business domains. Includes 63 China‑specific roles and integrates with 18 AI coding tools via simple install scripts. Works with the companion “Agency Orchestrator” (CLI or native desktop app) to compose multi‑agent workflows without writing code. MIT‑licensed and published on npm.
altic-dev/FluidVoice
An open-source voice-to-text dictation app for macOS that enables on-device AI transcription and system control via voice commands.
Gentleman-Programming/gentle-ai
Gentle‑AI is a Go‑based configurator that enhances existing AI coding agents (Claude Code, Cursor, etc.) with persistent memory, skill libraries, planning (Spec‑Driven Development), security guardrails and optional evidence‑based review (Receipt‑Driven Development). It installs via a simple script, runs an interactive TUI to select agents and components, writes native config files, and provides a health‑check command. The tool does not ship its own model; it merely augments agents you already use.
NeoLabHQ/context-engineering-kit
A plug‑in marketplace for LLM‑based coding assistants (Claude Code, Gemini, Antigravity, etc.) that supplies token‑efficient, context‑engineering patterns—reflection loops, sub‑agent judges, specification‑first development, code‑review agents, and Git helpers—to make AI‑generated code more reliable and easier to integrate into CI pipelines.
harbor-framework/terminal-bench-science
Terminal‑Bench‑Science is a community‑run benchmark that evaluates autonomous AI agents on real scientific research workflows (70+ tasks across life, physical, earth, math, and engineering). It uses the Harbor execution framework to run each task in a terminal environment, compares agent outputs to oracle solutions, and publishes scores on a public leaderboard.
0xShug0/audio.cpp
A high-performance C++ audio inference framework built on ggml that provides a portable, native runtime for running diverse local audio models including TTS, ASR, and voice conversion.
browseros-ai/BrowserOS
BrowserOS is an open‑source Chromium fork that embeds a locally‑run AI assistant for humans; BrowserOS neo is a companion browser that lets external AI agents (Claude Code, Codex, Cursor, etc.) drive real, logged‑in tabs. Both run on your machine, import Chrome data, support any LLM via API keys or local models, and keep all session data private.
denizsafak/abogen
A text-to-speech conversion tool that turns ePub, PDF, and text files into high-quality audio with synced subtitles using the Kokoro-82M model.
index-tts/index-tts
A zero-shot text-to-speech system that clones voices from a single audio clip with fine-grained control over emotion, speed, and pronunciation across five languages.
kizuna-ai-lab/sokuji
A cross-platform live speech translation app that enables real-time two-way translation for bilingual meetings using either cloud APIs or fully offline on-device inference.
ggml-org/whisper.cpp
A high-performance C/C++ port of OpenAI's Whisper model for efficient, offline, on-device automatic speech recognition across diverse hardware.
echo-loop/Echo-Loop
An AI-powered English listening and speaking training app that automates a scientific learning loop of intensive listening, shadowing, and retelling with spaced repetition.
risa-labs-inc/BossConsole
BossConsole is an open‑source, JVM‑native desktop console that lets any LLM agent work with a real browser, terminal, editor, git, Docker/Kubernetes, and more. It exposes ~100 tools via a Model Context Protocol, offers fine‑grained RBAC and per‑tool kill‑switches, and supports hot‑reloading plugins so agents can evolve their own capabilities.
huiliyi37/Tianshu-harness
Tianshu Harness is a TypeScript‑based AI coding‑assistant runtime. It provides a shared core engine for both a terminal TUI and a Tauri desktop GUI, adds a “Cognitive Virtual Machine” with 72 safety hooks, supports multi‑agent orchestration, prefix‑cache for low token cost, Zen/Plan modes, and a star‑domain system for different cognitive styles. Install via a one‑click script, npm, or pre‑built binaries for macOS/Windows/Linux, and use the `rivet` CLI to interact with LLMs for multi‑turn coding tasks.
EverMind-AI/EverOS
EverOS is a Python library that provides a local‑first, Markdown‑backed memory layer for LLM agents. It stores conversations and files as editable `.md` files, syncs them to SQLite and LanceDB indexes, and offers a REST API for adding, flushing, and searching memories. With just an OpenRouter API key you can run a server, ingest text, and retrieve it via keyword search; optional embedding, rerank, and multimodal extensions add hybrid search and image/PDF handling. The project positions itself as the persistent memory backend for many EverMind‑ecosystem agents and demos.
OpenMinis/OpenMinis
OpenMinis is an open‑source mobile AI agent (iOS + Android) that lets you run any LLM (Claude, GPT, Gemini, etc.) on‑device. It ships a sandboxed Alpine Linux shell (iSH on iOS, PRoot on Android) so the agent can install packages, run scripts, and manipulate files, while also exposing native device services (Health, Calendar, HomeKit, Bluetooth, clipboard, media, etc.) as tools. Users create reusable “skills” (folders with a `SKILL.md` and optional scripts) and organise work into separate workspaces. Typical uses include nutrition logging, morning briefings, task extraction from chats, syncing with Obsidian, and share‑to‑event automation. The codebase is Swift/SwiftUI for iOS, Kotlin/Compose for Android, and is GPL‑v3 licensed.
xbtlin/ai-berkshire
AI Berkshire is a set of Claude Code / Codex skills that orchestrate multiple LLM agents—each embodying a famous value‑investor’s perspective—to produce structured, data‑validated investment research reports, checklists, industry scans, and portfolio tools.
Moyf/moys-asr-workflow
A subtitle generation and editing workflow that uses AI ASR APIs to transcribe local media and provides a dedicated editor for refining and exporting subtitles.
TokenRhythm/opensquilla
OpenSquilla is a token‑efficient, micro‑kernel AI‑agent framework that routes each turn to the cheapest capable LLM. It bundles a local model router, persistent memory, sandboxed tools, web search, and on‑device embeddings, and runs the same turn loop across a desktop GUI, CLI, and chat‑channel integrations. Install via a pre‑built desktop installer, a one‑line `uv` wheel install, or from source. Supports 20+ LLM providers, many chat channels, and optional extras like Matrix or PDF generation. Licensed Apache 2.0, with privacy‑focused telemetry that can be disabled.
thewh1teagle/vibe
A private, offline audio and video transcription tool that uses AI models like Whisper to convert speech to text locally on your device.
k2-fsa/sherpa-onnx
A portable framework for running speech-to-text, text-to-speech, and other audio AI models locally across diverse hardware and operating systems.
Beingpax/VoiceInk
A native macOS application that uses local AI models to provide fast, private, and context-aware voice-to-text transcription.
presenton/presenton
Presenton is an open‑source, self‑hosted AI presentation generator. It lets you create, edit, and export fully editable PowerPoint decks using any LLM or image model (OpenAI, Anthropic, Gemini, Ollama, etc.). Available as a Docker container, a native Electron desktop app, or a cloud/enterprise deployment, it offers drag‑and‑drop slide editing, custom HTML/Tailwind templates, and an API for automated generation—all under an Apache 2.0 license.
kyutai-labs/pocket-tts
A lightweight, CPU-efficient text-to-speech application that enables local voice cloning and multi-language audio generation without requiring a GPU.
chidiwilliams/buzz
An offline transcription and translation tool powered by OpenAI's Whisper that supports local files, YouTube links, and live audio.
AutoArk/GPA
GPA is a unified auto-regressive transformer model that integrates speech recognition (ASR) and text-to-speech (TTS) into a single system with near-SOTA performance.
Alpha-Dojo/DojoAgents
DojoAgents is an open‑source AI‑agent framework for personal investing. It runs a large‑language‑model‑driven “Agent Loop” that fetches market data, parses news, executes Python‑based calculations, and stores reusable analysis skills. A FastAPI backend and React dashboard let users view portfolio metrics, get daily market overviews, run news‑impact analyses, and even diagnose a portfolio from a screenshot. Installable via `uv pip install dojoagents`, it works with any LLM API (OpenAI, Gemini, Anthropic, local Ollama, etc.) and can push automated insights to chat apps.
OpenByteInc/QuantDinger
QuantDinger is an open‑source, self‑hosted AI‑powered trading operating system. It lets you write Python strategies, back‑test them, run paper or live trades on crypto exchanges and traditional brokers, and optionally call LLM providers for market research. The stack is containerised (Docker Compose), uses PostgreSQL, Redis, Celery, and optional Prometheus/Grafana monitoring, and includes a secure Agent Gateway for external AI agents. It ships with a one‑command installer, detailed security hardening, and Apache‑2.0 licensing.
sv-number/skills
A minimal, MIT‑licensed *skill* that lets AI agents obtain disposable phone numbers and read SMS verification codes via a commercial API, enabling automated sign‑ups and 2FA without human interaction.
modelscope/FunASR
An industrial speech recognition toolkit for offline, streaming, and edge deployment, providing high-performance ASR, VAD, and speaker diarization pipelines.
tphakala/birdnet-go
A self-hosted, real-time soundscape analyzer for birds, wildlife, and bats that provides local AI inference and a web dashboard for detections.
charmbracelet/crush
Crush is a cross‑platform terminal application that lets you converse with any LLM (OpenAI, Anthropic, local Ollama, etc.) to write, edit, and run code. It supports multiple sessions, LSP‑based code understanding, extensible tool back‑ends via Model Context Protocol, and a Bash‑style `crushrc` for per‑project or global configuration.
verl-project/verl
verl is an open‑source reinforcement‑learning library for large language models. It provides a modular, high‑performance framework (Hybrid‑Flow) that integrates with FSDP, Megatron‑LM, vLLM, SGLang, and Hugging‑Face, supporting many RLHF algorithms (PPO, GRPO, DAPO, etc.), multi‑modal training, and scaling up to 671 B‑parameter models across hundreds of GPUs.
ace-step/ACE-Step-1.5
ACE‑Step 1.5 is an open‑source music‑generation foundation model that combines a language‑model planner with a Diffusion‑Transformer audio decoder. It can create full songs (10 s‑10 min) in under 10 seconds on a RTX 3090, runs with <4 GB VRAM for base models, and supports LoRA fine‑tuning from a handful of tracks. The repo provides Gradio and REST interfaces, launch scripts for all major OSes, multiple model sizes (including a 4 B XL version), and extensive control over lyrics, style, tempo, and instrumentation.
GreyDGL/PentestGPT
PentestGPT is an MIT‑licensed Python CLI that chains LLMs (Claude Code, Codex, or many other providers) into an autonomous penetration‑testing pipeline. It runs staged phases (recon → exploit → walkthrough for CTF, or asset discovery → vulnerability identification → report for real pentests), can persist sessions, and offers a legacy interactive mode where a user guides three cooperating LLMs. Install with `make install` (Python 3.12+, uv) or via Docker, then run `pentestgpt --target <IP>` (or `pentestgpt-legacy` for the interactive version). The tool reports optional anonymous telemetry and achieved ~86 % success on a published benchmark.
Edge0-AI/Audio8_TTS
A compact multilingual text-to-speech model with zero-shot voice cloning, available in 0.6B and 0.1B parameter versions for high-efficiency speech synthesis.
fishaudio/fish-speech
A state-of-the-art multilingual text-to-speech system that uses a Dual-AR architecture to provide realistic voice cloning and fine-grained emotional control via natural language tags.
huggingface/speech-to-speech
Speech‑to‑Speech is a Hugging Face open‑source pipeline that turns spoken input into spoken output. It chains VAD → STT → LLM → TTS, each component being swappable (Silero VAD, Parakeet/Faster‑Whisper/Whisper STT, any OpenAI‑compatible LLM or local Transformers/MLX model, Qwen3‑TTS or other TTS back‑ends). The whole system speaks the OpenAI Realtime protocol over WebSocket/WebRTC, so existing OpenAI Agents SDK code works. Install with a single `pip install speech-to-speech`, then run `speech-to-speech serve` (server) and `speech-to-speech talk` (client) or `speech-to-speech local` (both together). It supports local‑only operation, mixed local/hosted LLMs, Docker deployment, and optional extras for alternative TTS/STT models. Designed for voice‑assistant robots (e.g., Reachy Mini) and low‑latency spoken AI applications.
rullerzhou-afk/clawd-on-desk
Clawd on Desk is a cross‑platform desktop‑pet that watches AI‑coding assistants (Claude Code, Copilot, Gemini, Cursor, etc.) and animates in real time—thinking, typing, building, handling permission requests, and celebrating task completion. It offers hook‑based integration for dozens of agents, permission‑bubble UI, remote approvals (Telegram/Feishu), a dashboard/HUD, and a read‑only mobile PWA companion.
QwenAudio/CosyVoice
CosyVoice is an advanced, LLM-based text-to-speech system designed for high-quality, zero-shot multilingual speech synthesis and voice cloning.
lemonade-sdk/lemonade
Lemonade is a community‑driven local AI server (and embeddable binary) that runs LLM, speech, TTS, and image models on CPUs, GPUs, or AMD NPUs. It mimics OpenAI‑style APIs, includes a model manager, hardware‑auto‑optimizing back‑ends, and a marketplace of ready‑made integrations.
muscriptor/muscriptor
A multi-instrument music transcription model that converts audio recordings into MIDI and professional sheet music using a transformer decoder architecture.
m-bain/whisperX
A fast automatic speech recognition system that enhances OpenAI's Whisper with batched inference, word-level timestamps, and speaker diarization.
MemTensor/MemOS
MemOS 2.0 is a Memory Operating System for LLM‑based AI agents. It offers a unified API to store, retrieve, edit, and delete multi‑modal memories (text, images, tool traces) organized as a graph. The system can run as a hosted cloud service (API‑key access) or be self‑hosted via Docker (Neo4j + Qdrant) or a local SQLite plugin. Ready‑made plugins integrate MemOS with OpenClaw, Hermes, and DeepSeek Harness. Benchmarks show it outperforms existing commercial memory products. Licensed under Apache 2.0.
akdeb/ElatoAI
A framework for deploying real-time voice AI models on ESP32 devices, enabling low-latency speech-to-speech conversations via edge functions.
OpenMOSS/MOSS-Transcribe-Diarize
An end-to-end audio understanding model that jointly performs speech transcription and speaker diarization for long-form, multi-speaker recordings in over 50 languages.
resemble-ai/chatterbox
A family of open-source text-to-speech models providing low-latency voice cloning and multilingual speech generation across various model sizes for GPU and CPU deployment.