bigsk1/voice-chat-ai

🎙️ Speak with AI - Run locally using Ollama, OpenAI, Anthropic or xAI - Speech uses SparkTTS, OpenAI, ElevenLabs, Kokoro, Typecast or xAI

What it solves

Voice Chat AI provides a way to interact with AI characters through speech, enabling real-time, expressive conversations. It solves the problem of limiting AI interactions to text-based interfaces by integrating multiple speech-to-text (STT) and text-to-speech (TTS) providers to create a seamless voice-driven experience.

How it works

The system integrates various LLM providers (OpenAI, xAI, Anthropic, Ollama) with a range of TTS engines (Spark-TTS, OpenAI TTS, ElevenLabs, Kokoro TTS, Typecast) and STT options (OpenAI, Local Faster Whisper). It can be run as a Web UI or a terminal-based CLI. For high-performance, real-time interaction, it utilizes OpenAI's WebRTC Realtime API. It also includes sentiment analysis to adjust AI responses based on the user's mood.

Who it’s for

Users looking for immersive AI companionship, role-play, or interactive storytelling. It is also for developers who want a flexible framework to mix and match different AI chat and voice providers locally or via API.

Highlights

  • Multi-Provider Support: Mix and match LLMs from OpenAI, xAI, Anthropic, and Ollama.
  • Real-time Interaction: Supports OpenAI's WebRTC Realtime API for instant, interruptible conversations.
  • Llocal Voice Cloning: Local zero-shot voice cloning via Spark-TTS.
  • Interactive Content: Includes 15+ built-in game types and immersive story adventures.
  • Mood Analysis: Adjusts AI responses based on user sentiment analysis.
  • Visual Integration: Ability to analyze screen captures using models like LLaVA.
  • Flexible Deployment: Supports native installation for best audio performance or Docker for easier setup.

Related

  • Project
  • Project
  • Project
  • Project
  • Project