LongCat-2.0 MoE Model: 1.6T Parameters and ASIC-Based Training

LongCat-2.0 is a large-scale Mixture-of-Experts (MoE) model with 1.6 trillion total parameters and 48 billion active parameters. Its primary significance lies in its successful training and deployment on large-scale clusters of AI ASIC superpods, demonstrating a viable alternative to the Nvidia GPU ecosystem for training trillion-parameter models.

Architecture and Training Scale

LongCat-2.0 utilizes a Mixture-of-Experts (MoE) architecture to balance model capacity with computational efficiency. While the model possesses 1.6 trillion total parameters, only 48 billion are active during any single inference pass, reducing the compute requirements per token.

Key training statistics include:

  • Dataset Size: Pretraining spanned over 35 trillion tokens.
  • Compute Infrastructure: The model was built on clusters of tens of thousands of AI ASIC superpods. Community analysis suggests the use of Huawei Ascend 910C chips, with some estimates placing the scale at approximately 1,024 superpods (roughly 50,000 chips).
  • Development Origin: The model is associated with Meituan, a major Chinese food delivery and commerce company.

Technical Innovations and Infrastructure

Beyond the parameter count, the development of LongCat-2.0 highlights a shift toward specialized hardware for AI training. The team invested significant effort into building a stable, secure, and scalable infrastructure to compensate for the less developed software ecosystem surrounding AI ASICs compared to the mature Nvidia CUDA environment.

Some technical discussions regarding the model's design include:

  • N-gram Embedding: The model incorporates N-gram embedding, a technique designed to improve how the model handles tokenization and semantic representation.
  • Architectural Influence: Observers note that while LongCat-2.0 builds upon work similar to the DeepSeek architecture, it introduces distinct architectural contributions beyond simple post-training or fine-tuning.

Community Evaluation and Observations

Early user testing and community discussions have raised several points regarding the model's performance, accessibility, and alignment:

Performance Benchmarks

In comparative testing against other frontier models, some users report mixed results. In one specific test involving nuclear physics (comparing U-235 and Pu-241 fuels), a user noted that LongCat-2.0 provided a well-reasoned but incorrect answer, while Qwen 3.7 Plus and Gemini Flash provided correct answers with stronger arguments.

Accessibility and Openness

There is significant community skepticism regarding the "open" nature of the model. Users have reported that Hugging Face and GitHub links provided by the developers have returned 404 errors, leading to concerns about the availability of the model weights and the full technical report.

Content Filtering and Alignment

Users have observed strict content filtering on sensitive political topics. Queries regarding the Tiananmen Square protests or Chairman Mao resulted in the model refusing to answer or returning "Too many requests" errors, indicating heavy alignment with regional regulatory requirements.

Language Behavior

Some users reported that when using the "Search" feature with the application set to English, the model occasionally returned results in Chinese, suggesting a strong influence of the original training data context on the model's output.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch