QVQ-Max Visual Reasoning Model Release

Qwen has officially released QVQ-Max, a visual reasoning model designed to move beyond simple image understanding to active analysis and problem-solving. QVQ-Max integrates visual perception with deep reasoning to solve tasks ranging from complex mathematical problems to creative content generation.

Scaling Thinking Process for Mathematical Accuracy

QVQ-Max demonstrates a direct correlation between the length of its internal thinking process and its accuracy on the MathVision benchmark. By adjusting the maximum length of the model's thinking process, Qwen observed a continuous improvement in accuracy, indicating that the model's reasoning capabilities scale with the computational budget allocated to "thinking."

Core Technical Capabilities

QVQ-Max is built to function as an assistant that is both "sharp-eyed" and "quick-thinking," focusing on three primary functional areas:

Detailed Observation

QVQ-Max parses complex visual inputs, including charts and daily snapshots, to identify key elements, textual labels, and minute details that may be overlooked by standard models.

Deep Reasoning

The model combines visual observations with background knowledge to derive conclusions. This includes solving geometry problems based on diagrams and predicting future events in video clips based on current scenes.

Flexible Application

Beyond analysis, QVQ-Max applies its reasoning to creative and practical tasks, such as:

  • Refining rough sketches into complete illustrations.
  • Generating short video scripts.
  • Creating role-playing content.
  • Providing critiques or interpretations of uploaded photos.

Practical Application Scenarios

QVQ-Max is designed for utility across three main domains:

  • Workplace: Assisting with data analysis, information organization, and code generation.
  • Education: Solving complex math and physics problems that rely on diagrams and explaining intricate concepts intuitively.
  • Daily Life: Providing practical advice, such as recommending outfit combinations from wardrobe photos or guiding users through recipes based on images.

Future Development Roadmap

Qwen has identified three primary areas for the evolution of QVQ-Max:

  1. Enhanced Observation Accuracy: Implementing grounding techniques to validate visual observations.
  2. Visual Agent Capabilities: Improving the model's ability to execute multi-step complex tasks, including operating computers, smartphones, and playing games.
  3. Multimodal Interaction: Expanding interaction beyond text to include tool verification and visual generation.

Sources