Qwen3.8-Max Release: 2.4T‑Parameter Model, Open Weights Next Week, and Broad Autonomous Capabilities

Qwen3.8‑Max Overview

Qwen3.8‑Max is the most capable model in the Qwen family to date, scaling to 2.4 trillion parameters (95 B active) and delivering comprehensive improvements across coding, work, research, and long‑horizon tasks. The release marks the first time a Qwen‑Max‑class model will have its weights open‑sourced, with the open weights scheduled for release next week.

Coding Capabilities

Autonomous Self‑Evolving Harness

Qwen3.8‑Max was tasked with building the oh‑my‑cli project from scratch and, over a 10+ day long‑horizon autonomous coding run, constructed a self‑evolving harness that continuously incorporates user feedback, community practices, and self‑test results. As of July 30 2026, after approximately 16 days of fully autonomous operation, the repository had accumulated 265 commits, 127 PRs, and 151 issues【...】.

Key elements of the harness include an issue state machine, dispatcher, monitor, and watchdog forming an execution loop; self‑testing that routes abnormal states back to the relevant issue/PR; and multi‑source evolution that converts community experience and feedback into executable work.

Reproducing and Improving a Research Paper

Given the paper “Unified Data Selection for LLM Reasoning” (arXiv:2605.22389) and only a set of GPUs, Qwen3.8‑Max spent about five days (~125 hours) writing roughly 7,600 lines of code, performing 33 rounds of GPU training, and reproducing the paper’s six main findings. It then entered a self‑improving research loop over the next ~88 hours, testing 18 improvement ideas across four rounds. The final method—counting hard decision points (“nhighgate”)—achieved a +2.7‑point gain on the AIME24 benchmark over the paper’s baseline【...】.

Beating Human Teams in a 24‑Hour Contest

In the WWW2025 Multimodal Dialogue Intent Recognition Challenge (hosted on Alibaba Cloud’s Tianchi platform) with 526 human teams, Qwen3.8‑Max operated autonomously under a strict 24‑hour limit. It built a solution that fine‑tuned and ensembled BERT, MacBERT, and RoBERTa for text, and a Qwen2.5‑VL‑7B vision‑language model backed by Chinese‑CLIP for screenshots, fused via a weighted‑voting system. Over 45 submissions, its accuracy rose from 0.60 to 0.853, beating 458 of the 526 teams (87 % of the field)【...】.

These three cases illustrate Qwen3.8‑Max’s ability to stay focused on open‑ended goals for days, generate its own ideas, and turn them into working results without human intervention.

Real‑World Work Competence

Scaling Real‑World RL Systems

Qwen3.8‑Max’s working ability is improved by jointly scaling RL environments and compute. Three coupled challenges were addressed:

  1. Continuously scaling decoupled real environments along independent axes—Task (single‑task → multi‑task → multi‑day), Workspace (multi‑file → hierarchical → complex heterogeneous folders), and Harness (category, version, skills)—so environment growth compounds combinatorially.
  2. A universal reward system that internalizes heterogeneous verification (execution‑based checks, rubric‑conditioned adjudication over text/visual output, agentic inspection) under automatically scalable rubrics.
  3. An online data balancer that shapes each batch to keep distribution over tasks, difficulty, workspaces, and harnesses highly balanced, suppressing inter‑batch gradient variance and sustaining stable RL compute scaling.

Together these provide breadth, reliable reward, and stability, yielding a measurable horizontal lift in real‑world working ability.

Breadth Across Hundreds of Professions

Stress‑testing showed Qwen3.8‑Max delivering production‑quality results in real workflows for many high‑value professions:

  • Corporate compliance counsel: surfaced 1,284 relevant clauses across hundreds of documents in under an hour, versus a paralegal team’s typical week.
  • UI/UX designer: produced a high‑fidelity, interactive prototype for the digital‑banking app NOVA (8 screens) in one shot with zero rounds of human revision, versus the conventional 3–5 rounds.
  • Restaurant brand founder: read over a hundred ingredient‑supply briefs and produced a complete 26‑dish menu in one pass, each dish annotated with average caloric value and ingredient provenance, holding the food‑cost ratio at 33.8 %.
  • Structural engineer: reconstructed the seismic structural model of a 30‑story office tower in the browser, with natural period, base shear, and inter‑story drift ratio available on hover, a task that normally takes over a week in specialized software.
  • Rehabilitation therapist: turned a 2D paper assessment form into a 3D interactive demo with freely rotatable viewing angles and layer‑by‑layer anatomical overlays, a task previously outsourced to a medical‑animation studio at 2–4 weeks’ lead time and thousands of dollars.
  • Sports data analyst: parsed ~8,400 offensive/defensive possessions per player into a ready‑to‑use player tactical profile and coaching report in tens of minutes, versus several working days for a traditional analytics team.

Building a Profitable End‑to‑End Quant Strategy in a Single Session

Powered by its Dynamic Workflows construction capability, Qwen3.8‑Max can turn a single conversation into an end‑to‑end quant‑research loop. From a one‑line task description it autonomously planned a complex dynamic workflow, built the data system, constructed base factors, and orchestrated multi‑round greedy iteration while dynamically analyzing backtests and correcting course. It also parallelized factor mining: from six short descriptions spanning classic factor families, it decomposed each into 50 research directions, dispatched ~330 sub‑agents, completed ~6,000 backtests, and continuously adapted the workflow mid‑run. Selected factors achieved excess Sharpe ratios of 0.64–1.48 with IC uniformly positive ranging from 0.010 to 0.014【...】.

Long‑Horizon Task Performance

Autonomous Chip Design and Closed‑Loop Optimization

Qwen3.8‑Max independently executed the full silicon design flow for a GCD/RSA cryptographic hardware accelerator. Starting from minimal inputs—a task description, a stub RTL workspace, and an evaluation script—it operated completely autonomously, managing RTL editing, simulation debugging, synthesis analysis, redundancy localization, and iterative datapath re‑architecting. Over a single continuous run it completed approximately 500 turns and 71 evaluations across 13 key milestones, reducing the synthesized gate count from 8,298 gates to 678 gates (an 81 % reduction in physical die area) while achieving timing closure at 500 MHz (+0.66 ns slack)【...】.

Key milestones included:

  • Algorithmic rewrite: modulo divider → iterative shift‑subtract (8,298 → 2,010 gates, Turn 22).
  • Redundancy elimination & bitwidth trimming (2,010 → 1,304 gates, Turns 35–48).
  • Register & control FSM pruning (1,304 → 907 gates, Turns 60–113).
  • Module fusion & logic sharing (907 → 765 gates, Turns 170–252).
  • Gate‑level refinement (765 → 678 gates, Turns 443–500).

Physical layout verification via OpenROAD showed the die shrinking from 106×106 µm² to 46×46 µm², wirelength dropping from 33,369 µm to 4,187 µm, and successful timing closure.

Continuous Learning in Long‑Term Operations (E‑Commerce Bench)

In the 365‑day e‑commerce operation simulation benchmark, Qwen3.8‑Max started with ¥100,000 capital and had to manage product selection, supply chain negotiation, inventory, dynamic pricing, and returns handling across 12 store types, 60 product categories, nearly 600 suppliers, and 7,000 products. It demonstrated continuous learning in supplier negotiations, achieving progressive price reductions and steady profit increases, causing negotiation efficiency to expand over time. It also generalized this experience to similar products while other models plateaued.

Against hidden risks (152 fraudulent merchants embedded in the supplier matrix) and seasonal demand swings, Qwen3.8‑Max invested heavily early, accelerating its asset growth curve and achieving a net profit exceeding ¥100,000 during the year‑end major promotion—nearly 2.4 times that of the second‑place GLM 5.2. It ultimately reached the highest total balance of ¥416,252 (a 4.16× return), surpassing GLM 5.2 by 38 % and representing a 152 % improvement over its previous flagship, Qwen3.7‑Max【...】.

Multimodal Agents and Visual Feedback Loops

Qwen3.8‑Max treats vision as a native feedback loop across planning, execution, verification, and iteration. It can understand and generate visual content, inspect its own intermediate results, detect deviations (e.g., a television facing the wrong direction, a misaligned interface), and autonomously revise its plan and correct the output.

This capability enables Hybrid Agent behavior: pairing coding (efficient, scalable heavy lifting) with GUI operation (reaching what a human can see and touch) to verify against a real, running application.

To measure this, the RecreationBench benchmark spans five platforms (desktop Ubuntu/macOS/Windows, mobile Android, and web). The model observes a real, running application only as a black box—no source code, no internet access—and must rebuild the whole application from scratch through repeated cycles of iterative coding and interactive feedback. Qwen3.8‑Max already demonstrates frontier‑level Hybrid Agent capability on this benchmark【...】.

For easier integration, Qwen‑MM‑Plugins provides a harness extension library offering image and video processing, multimodal memory, dynamic‑resolution support, visual tool use, and specialized capabilities for tasks such as video editing, Blender, and CAD. Any existing agent harness can be extended into a more naturally multimodal‑native system.

User Feedback and Community Reaction

Comments on Hacker News discussion highlighted several points:

  • The announcement of a Qwen3.8‑27B open‑weight release next week was noted as significant, with one comment calling it "the real news here"【...】.
  • Some users pointed out that a Qwen3.8‑Max‑Preview had already debuted on Alibaba’s Token Plan, Qoder, and QoderWork on July 19th, questioning what the August 3rd announcement adds【...】.
  • There was curiosity about cost and efficiency: the model supports a reasoning_effort parameter (xhigh, medium, low) to adjust reasoning depth and control cost, with hopes that it will be significantly cheaper than alternatives【...】.
  • A few commenters expressed difficulty running the 27B model locally on modest hardware, noting GPU memory constraints【...】.
  • Questions were raised about token‑ and reasoning‑efficiency, with requests for data on token usage in benchmarks【...】.
  • The model’s vision capabilities were praised, though some noted timeout issues when using it for visual web‑development tasks【...】.
  • Discussions touched on broader implications: the potential narrowing of the window for banning open‑weight models, the model’s possible pro‑West bias in distilled outputs, and whether the model could be stripped down to a language‑specific variant for lighter local use【...】.
  • Several users expressed excitement about integrating Qwen3.8‑Max with agent frameworks such as Claude Code, Codex, Qoder CLI, Qwen Code, and OpenClaw, sharing example configuration snippets【...】.

Availability and Usage

Qwen3.8‑Max is accessible via QwenCloud at https://www.qwencloud.com/. The model weights will be open‑sourced on Hugging Face and ModelScope next week.

API Usage

The official API supports a reasoning_effort argument to trade off depth for cost:

  • xhigh (default): for complex tasks demanding thorough analysis
  • medium: balances accuracy and speed
  • low: optimizes for speed and cost

preserve_thinking is enabled by default for the best out‑of‑the‑box experience.

Example call using the OpenAI‑compatible client:

from openai import OpenAI
import os
api_key = os.environ.get("DASHSCOPE_API_KEY
if not api_key:
    raise ValueError("DASHSCOPE_API_KEY is required. Set it via: export DASHSCOPE_API_KEY='your-api-key'
)
client = OpenAI(
    api_key=api_key,
    base_url=os.environ.get(
        "DASHSCOPE_BASE_URL",
        "https://dashscope-intl.aliyuncs.com/compatible-mode/v1",
    ),
)
messages = [{"role": "user", "content": "Write a Python function to merge two sorted linked lists."}]
completion = client.chat.completions.create(
    model="qwen3.8-max",
    messages=messages,
    extra_body={
        "enable_thinking": True,
        # "preserve_thinking": True,
    },
    reasoning_effort="xhigh",
    stream=True,
)
# ... process streaming chunks for reasoning_content and answer_content

More details are available in the API documentation.

Integration with Coding Assistants

  • Claude Code: set ANTHROPIC_MODEL="qwen3.8-max" and point ANTHROPIC_BASE_URL to a QwenCloud compatible endpoint.
  • Codex: add a model entry to ~/.codex/model-catalog.local.json with slug: "qwen3.8-max", context_window: 1000000, and appropriate reasoning levels; then configure ~/.codex/config.toml to use the ModelStudio provider.
  • Qoder CLI: install via curl -fsSL https://qoder.com/install | bash and run qoder.
  • Qwen Code: install via npm install -g @qwen-code/qwen-code@latest and run qwen.
  • OpenClaw: follow the installation script, set DASHSCOPE_API_KEY, and configure ~/.openclaw/openclaw.json to point to the QwenCloud endpoint.

Conclusion

Qwen3.8‑Max pushes the frontier of autonomous AI systems by combining massive scale (2.4 T parameters) with open‑weight availability, strong self‑evolving coding abilities, broad real‑world work competence, long‑horizon optimization in hardware design and business simulation, and multimodal visual feedback loops. The upcoming release of its weights, together with the smaller Qwen3.8‑27B variant, offers the community a powerful foundation for building agents that can operate with minimal human oversight across coding, professional workflows, research, and creative tasks.

Sources

Related

  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch
  • Dispatch