worm128/AI-YinMei

AI吟美-人工智能主播-Vtuber-桌宠-智能体

What it solves

AI-YinMei provides an all-in-one platform for creating AI-driven live-streaming avatars and bots. It solves the complexity of integrating multiple AI modalities—such as LLMs, text-to-speech (TTS), vision, and character animation—into a single system that can interact with audiences across various platforms (Bilibili, TikTok, QQ, etc.) in real-time.

How it works

The system acts as a central hub that aggregates inputs (like live stream comments or chat messages) and processes them through a pipeline of AI models. It uses a multi-head attention mechanism for intent and emotion analysis, connects to OpenAI-compatible LLMs for dialogue, and employs streaming TTS (like CosyVoice2 or GPT-SoVITS) for low-latency voice responses. It integrates with VTube Studio or its own desktop pet software for visual representation and uses Neo4j for "diffuse thinking" (knowledge graph relationships) and FastGPT for RAG-based knowledge management.

Who it’s for

It is designed for live streamers, content creators, and hobbyists who want to deploy an interactive, anime-style AI persona (VTuber) that can sing, draw, and chat with viewers autonomously.

Highlights

  • Multi-Platform Integration: Aggregates comments from Bilibili, TikTok, Kuaishou, Douyu, Huya, WeChat, and QQ.
  • Low Latency: Features full-streaming dialogue and TTS, achieving response times as low as 600ms.
  • Extensive Toolset: Includes 27 built-in tools for tasks like web searching (via SearXNG), image generation (Stable Diffusion), and computer control.
  • Advanced Memory: Combines short-term MongoDB persistence with long-term memory that is selectively awakened based on time-related keywords.
  • Multimodal Capabilities: Supports AI singing (voice cloning), image generation, and vision-based camera monitoring.

Related

  • Project
  • Project
  • Project
  • Project
  • Project