Ox Alpha: A Stealth Reasoning Model for Coding and Agentic Workflows
Ox Alpha is a multimodal reasoning model for production engineering
Ox Alpha is a reasoning model designed for coding, sustained agentic work, and production workloads. It is specifically optimized for long-horizon software engineering, complex reasoning tasks, and workflows that require the integration of text with visual context. Released on August 20, 2026, the model is currently available as a "stealth" model on OpenRouter, meaning it is developed and operated by an anonymous third-party provider.
Technical Specifications and Capabilities
Ox Alpha provides high-capacity context and multimodal input support to facilitate complex software development and agentic tasks.
Context and Modalities
- Context Window: 1,048,576 tokens.
- Maximum Output: 131,072 tokens.
- Input Modalities: Supports text, images, and video.
- Output Modalities: Returns text.
Feature Support
- Tool Calling: Supports
toolsandtool_choicefor function calling. - Structured Outputs: Supports
response_formatfor JSON output (without JSON-schema enforcement). - Reasoning: Supports reasoning-enabled requests, allowing users to access the model's internal thinking process via the
reasoning_detailsarray in the API response.
Performance and Pricing
Ox Alpha is currently offered for free through OpenRouter, with no charges for prompt or completion tokens.
Latency and Throughput
Based on OpenRouter's P50 metrics:
- Throughput: 33 tokens per second (tps).
- Latency: 3.76 seconds.
- Availability: 98.47% over a three-day period.
- Cache Hit Rate: Average of 72.97%.
Production Usage
High-volume traffic to Ox Alpha is driven by agentic coding tools and autonomous agents, including:
- Claude Code: 27.7B tokens
- Hermes Agent: 21.7B tokens
- Oh-My-Pi: 19.9B tokens
- DeepSeek Harness: 17.2B tokens
- ZCode: 13.4B tokens
Community Analysis and Speculation
Because the model's provider remains anonymous, the developer community has used stylometry and behavioral analysis to speculate on its origin.
Probable Origin: GLM Series
Multiple users suggest the model is a variant of the GLM (General Language Model) series, citing specific reasoning patterns and stylistic markers.
"Based on its indecisive and far-too-lengthy thinking traces when given complex instructions... as well as a rudimentary stylometry... this is almost certainly a GLM model."
Other theories suggest it could be a GLM-5.3 Flash or a vision-enabled variant of GLM-5.3, potentially intended to compete with DeepSeek's flash models.
Behavioral Observations
- Knowledge Cutoff: Testing suggests a knowledge cutoff around mid-2025.
- Content Filtering: Users report the model refuses to answer questions regarding Tiananmen Square but is more permissive regarding technical instructions for electronic warfare that other frontier models typically block.
- Coding Performance: While some users find it capable for general coding, others report failures in specific tasks, such as implementing CSS
color-mix()functions, where the model allegedly replaced dynamic logic with hardcoded hexadecimal values. - Creative Tasks: Some users report the model outperforms other high-end models (like "K3") on creative and "looser" tasks, though visual reasoning is noted as a weaker point.
Data Privacy and Terms
Ox Alpha is governed by OpenRouter's Stealth Model Terms. While the provider states that prompts and completions are retained but not used for training, the anonymity of the provider has led to community caution regarding the submission of proprietary or confidential data.
Sources
- HNOx Alpha
Related
- Dispatch
- Dispatch
- Project
- Dispatch
- Dispatch