Raven: The Harness of Harnesses for Recursive Self‑Improvement
What Raven Claims to Deliver
Raven positions itself as a "harness of harnesses" that enables recursive self‑improvement (RSI) for AI agents. It bundles four built‑in agents—Research, Code, Design, and On‑call—under a Host Agent that can orchestrate complex, multi‑step workflows. The platform is built on EverOS, providing persistent memory across sessions, and includes a Curator component that can modify an agent’s planning, capability, memory, and action modules between rounds.
"Raven is the harness of harnesses, built for recursive self‑improvement (RSI)." – Raven README
End‑to‑End Project Showcases
Raven demonstrates complete, autonomous delivery of three distinct projects:
- THRESHOLD game – a first‑person shooter built in Godot 4 over four days, involving 42 planning‑development‑verification cycles. The deliverable includes the playable game, poster, presentation, and website.
- Raven RSI benchmark – a self‑improving pipeline that ran 172 nano‑chat pre‑training experiments across seven rounds, reducing validation bits‑per‑byte (
val_bpb) by 5.8 % within a 20‑minute, single‑GPU budget. It also performed a dam‑break CFD simulation (error reduced from 1e0 to 1.36e‑10) and an FEA limit‑load bisection search over eight rounds. - Product launch kit – a browser‑based physics mini‑game, a 16‑slide overview, bilingual posters, and a README, all generated by Raven.
Each showcase includes video demos, website links, and downloadable slide decks, illustrating Raven’s ability to drive a team of specialized agents from brief to final deliverable without human intervention.
Benchmark Performance of Built‑In Agents
Raven’s four core agents are evaluated on domain‑specific benchmarks:
- Raven‑Research – ranks competitively on the DeepResearch Mixed benchmark (accuracy, token usage, and cost).
- Raven‑Code – achieves state‑of‑the‑art scores on SWE‑bench Pro, SWE‑bench Verified, WorkBuddy‑Code Reward, and SWE‑Refactor. It also tops DataAgentBench (Pass@1 = 0.8762) for data‑analysis tasks.
- Raven‑Design – leads PresentBench for slide generation and scores highly on ArtifactsBench visual‑design metrics.
- Raven‑Oncall – outperforms Claude Code on AI4AI and AI4S internal benchmarks, delivering higher quality and lower cost for unattended workflow automation.
These results are presented as static images in the repository; the README does not provide raw numbers or links to the underlying evaluation scripts.
Runtime Self‑Evolution Architecture
Raven’s agent loop is split into four decoupled strategy modules:
| Module | Role |
|---|---|
| Memory | Determines what information the agent sees each turn. |
| Planning | Generates the plan for the current iteration. |
| Capability | Selects which tools or external services the agent may invoke. |
| Action | Executes the chosen operation and judges its outcome. |
A Curator can rewrite any of these modules for a specific agent, effectively changing the agent’s prompt, toolset, skills, and judgment code. Changes are only installed after passing validation checks, and failed checks revert the agent to the previous version. The Curator is shipped as an experimental component in the repository rather than as a packaged binary.
A Persona is generated first: the user describes the desired assistant, and the Curator assembles a lead role plus specialist sub‑agents drawn from the built‑in suite. The resulting harness defines task division, tool access, and verification criteria.
"Describe what you need once. Raven assembles the assistant, keeps improving it while you work, and afterwards a single sentence is enough to put it to work again."
Integration of Third‑Party Agents
Raven can orchestrate external agents via three mechanisms:
- ACP (Agent Communication Protocol)
- CLI wrappers
- OpenAI‑compatible APIs
Presets for 13 third‑party agents (e.g., Claude Code, GitHub Copilot, Qwen Code) are included, allowing users to mix built‑in and external capabilities within a single workflow.
Installation and Usage
Raven offers multiple installation paths:
- One‑liner script for Linux/macOS/WSL2:
curl -fsSL https://raven.evermind.ai/install.sh | bash - PowerShell script for native Windows.
- Docker compose for containerised deployment.
- Editable source checkout for development (
./install.sh).
After installation, the command raven web launches a browser‑based UI where users can create tasks, monitor sub‑agents, inspect memory, and browse skills.
Community Reaction on Hacker News
The HN discussion highlighted several themes:
- Skepticism about star count – users questioned how a relatively unknown repo amassed thousands of GitHub stars.
- Critique of “one prompt, one result” claim – commenters argued that real‑world product development requires iterative prompting and steering, which may limit the practicality of a fully autonomous harness.
- Clarification of RSI – a comment noted that “RSI” stands for recursive self‑improvement, not “repetitive strain injury”.
- Comparison to existing meta‑harnesses – users referenced similar projects such as Omnigent, Paseo, and DeepSeek Harness, asking for comparative benchmarks.
- Performance concerns – some pointed out modest gains on coding benchmarks (e.g., 0.8 % improvement on SWE‑bench Verified) and questioned token efficiency.
- Positive impressions – a few commenters praised the visual design of the README and the impressive THRESHOLD showcase.
Overall, the thread reflects a mix of enthusiasm for the concept of a self‑evolving multi‑agent platform and caution about the hype surrounding fully autonomous AI project delivery.
Core System Summary
| System | Function |
|---|---|
| Agent Orchestration | Manages DAG‑based task dependencies and parallel execution. |
| Evolver | Diagnoses failures, tests candidate harness changes, and adopts improvements that outperform baselines. |
| EverOS Memory | Provides long‑term, Markdown‑native memory across sessions. |
| SkillForge | Retrieves over 114 k curated skills from the SkillCorpus for on‑demand expertise. |
| Proactivity | Monitors events and schedules actions to anticipate user needs. |
Licensing and Citation
Raven is released under the Apache 2.0 license. Users are asked to cite the technical report (v1, Sep 2026) when employing Raven in academic work.
This article synthesises the information publicly available in the Raven GitHub repository and the Hacker News discussion thread linked above. No additional data or unpublished details have been introduced.
Sources
Related
- Project
- Project
- Project
- Project