jeinlee1991/chinese-llm-benchmark

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。

What it solves

It addresses the difficulty of evaluating Chinese Large Language Models (LLMs) by providing a scalable, structured benchmark called ReLE (Really Reliable Live Evaluation). The project aims to help users avoid "blindly choosing" models by offering detailed performance data across diverse domains and a massive library of model defects to assist in research and improvement.

How it works

ReLE evaluates hundreds of commercial and open-source LLMs across seven primary domains and approximately 300 fine-grained dimensions. These domains include:

  • Education (Primary, Middle, and High School subjects, and Gaokao)
  • Medical and Mental Health
  • Finance
  • Law and Administrative Civil Service
  • Reasoning and Mathematical Calculation
  • Language and Instruction Following
  • Agents and Tool Calling

It provides comprehensive leaderboards, a defect library containing over 2 million entries, and a model selection tool that allows users to upload their own test data to find the most cost-effective model for their specific scenario.

Who it’s for

  • LLM Developers: To diagnose capability anisotropy and improve models using the defect library.
  • Enterprise Users: To select the most performant and cost-effective LLM for specific business needs (e.g., reducing costs by up to 90%).
  • Researchers: To analyze the capabilities of various Chinese and global LLMs through structured benchmarks.

Highlights

  • Massive Scale: Covers 398+ models, including major commercial and open-source options.
  • Granular Evaluation: Tests across ~300 dimensions, from specialized fields like dentistry to general high school Chinese.
  • Huge Defect Library: Includes over 2 million documented model failures for community analysis.
  • Cost-Optimization Tool: Enables scenario-specific testing to identify the cheapest model that meets performance requirements.

Related

  • Project
  • Project
  • Project
  • Project
  • Project