worm128/AI-YinMei
AI吟美-人工智能主播-Vtuber-桌宠-智能体
What it solves
AI-YinMei provides an all-in-one platform for creating AI-driven live-streaming avatars and bots. It solves the complexity of integrating multiple AI modalities—such as LLMs, text-to-speech (TTS), vision, and character animation—into a single system that can interact with audiences across various platforms (Bilibili, TikTok, QQ, etc.) in real-time.
How it works
The system acts as a central hub that aggregates inputs (like live stream comments or chat messages) and processes them through a pipeline of AI models. It uses a multi-head attention mechanism for intent and emotion analysis, connects to OpenAI-compatible LLMs for dialogue, and employs streaming TTS (like CosyVoice2 or GPT-SoVITS) for low-latency voice responses. It integrates with VTube Studio or its own desktop pet software for visual representation and uses Neo4j for "diffuse thinking" (knowledge graph relationships) and FastGPT for RAG-based knowledge management.
Who it’s for
It is designed for live streamers, content creators, and hobbyists who want to deploy an interactive, anime-style AI persona (VTuber) that can sing, draw, and chat with viewers autonomously.
Highlights
- Multi-Platform Integration: Aggregates comments from Bilibili, TikTok, Kuaishou, Douyu, Huya, WeChat, and QQ.
- Low Latency: Features full-streaming dialogue and TTS, achieving response times as low as 600ms.
- Extensive Toolset: Includes 27 built-in tools for tasks like web searching (via SearXNG), image generation (Stable Diffusion), and computer control.
- Advanced Memory: Combines short-term MongoDB persistence with long-term memory that is selectively awakened based on time-related keywords.
- Multimodal Capabilities: Supports AI singing (voice cloning), image generation, and vision-based camera monitoring.
Related
- Project
- Project
- Project
- Project
- Project