xikhar/persona

Bringing real-time voice to life.

What it solves

Persona provides a visual identity for voice-based AI assistants. It creates a desktop character (avatar) that reacts in real-time to audio output from other applications, solving the problem of having a "voice without a face" during AI conversations.

How it works

Persona is a cross-platform desktop application that listens to the system's audio playback streams (using PipeWire on Linux, WASAPI on Windows, and Core Audio on the macOS). When it detects audio from a selected voice app, the avatar performs "Speaking" or "Idle" animations. It can also be integrated with AI agents via a Model Context Protocol (MCP) server, allowing an agent to trigger specific custom animations based on the conversation context.

Who it’s for

Users who want to add an expressive visual presence to their local AI voice assistants or any desktop voice experience.

Highlights

  • Real-time audio reactivity: Automatically triggers speaking animations based on system audio output.
  • Crosspatform support: Native listeners for Windows, macOS, and Linux.
  • Customizable avatars: Supports .vrm and .vrma files for custom models and animations.
  • MCP Integration: Allows connected AI agents to control the avatar's window visibility and trigger specific animations.

Related

  • Project
  • Project
  • Project
  • Project
  • Project