QVQ-Max Visual Reasoning Model Release
Qwen has officially released QVQ-Max, a visual reasoning model designed to move beyond simple image understanding to active analysis and problem-solving. QVQ-Max integrates visual perception with deep reasoning to solve tasks ranging from complex mathematical problems to creative content generation.
Scaling Thinking Process for Mathematical Accuracy
QVQ-Max demonstrates a direct correlation between the length of its internal thinking process and its accuracy on the MathVision benchmark. By adjusting the maximum length of the model's thinking process, Qwen observed a continuous improvement in accuracy, indicating that the model's reasoning capabilities scale with the computational budget allocated to "thinking."
Core Technical Capabilities
QVQ-Max is built to function as an assistant that is both "sharp-eyed" and "quick-thinking," focusing on three primary functional areas:
Detailed Observation
QVQ-Max parses complex visual inputs, including charts and daily snapshots, to identify key elements, textual labels, and minute details that may be overlooked by standard models.
Deep Reasoning
The model combines visual observations with background knowledge to derive conclusions. This includes solving geometry problems based on diagrams and predicting future events in video clips based on current scenes.
Flexible Application
Beyond analysis, QVQ-Max applies its reasoning to creative and practical tasks, such as:
- Refining rough sketches into complete illustrations.
- Generating short video scripts.
- Creating role-playing content.
- Providing critiques or interpretations of uploaded photos.
Practical Application Scenarios
QVQ-Max is designed for utility across three main domains:
- Workplace: Assisting with data analysis, information organization, and code generation.
- Education: Solving complex math and physics problems that rely on diagrams and explaining intricate concepts intuitively.
- Daily Life: Providing practical advice, such as recommending outfit combinations from wardrobe photos or guiding users through recipes based on images.
Future Development Roadmap
Qwen has identified three primary areas for the evolution of QVQ-Max:
- Enhanced Observation Accuracy: Implementing grounding techniques to validate visual observations.
- Visual Agent Capabilities: Improving the model's ability to execute multi-step complex tasks, including operating computers, smartphones, and playing games.
- Multimodal Interaction: Expanding interaction beyond text to include tool verification and visual generation.
Sources
- OriginalQVQ-Max: Think with Evidence