aliyun/alibabacloud-bailian-speech-demo

Sample Repository for the AlibabaCloud Bailian Speech SDK

What it solves

This repository provides a comprehensive set of example codes for developers to integrate speech-related AI capabilities from the Alibaba Cloud Bailian platform. It simplifies the process of implementing speech recognition (ASR), speech synthesis (TTS), and advanced multimodal interactions like real-time voice chat and music generation.

How it works

The project serves as a a collection of implementation samples using the DashScope SDK. It demonstrates how to call various models such as Qwen-Audio-3.0 (ASR, TTS, and Realtime), CosyVoice, and Fun-ASR. Developers can implement these capabilities by configuring their Alibaba Cloud account and API keys to interact with the Bailian model services.

Who it’s for

Developers looking to build applications featuring voice-driven AI, such as customer service bots, real-time meeting transcription services, audio-visual analysis tools, and interactive voice assistants.

Highlights

  • Comprehensive Speech Suite: Supports streaming and non-streaming ASR, voice cloning (TTS), and multi-language synthesis.
  • Real-time Interaction: Includes examples for end-to-end real-time voice dialogue via WebSockets, supporting function calling and reasoning routes.
  • Multimodal Integration: Demonstrates how to combine speech models with LLMs for video chat, translation, and audio summarization.
  • Creative AI: Provides capabilities for generating music from prompts or lyrics.
  • Enterprise Use-cases: Includes specific examples for call center quality inspection and professional text normalization for natural-sounding speech.

Related

  • Project
  • Project
  • Project
  • Project
  • Project