mini-AGI: Continual Learning Byte-Level Model for Consumer Hardware

mini-AGI is a byte-level language model designed to be trained from scratch on a single consumer GPU with 8GB of VRAM. Unlike traditional frozen models, mini-AGI implements a continual learning system that allows it to learn from a stream of data without the catastrophic forgetting typically associated with online training.

Architecture: Paged Mixture-of-Experts and Adaptive Depth

mini-AGI uses a non-traditional transformer architecture that prioritizes memory efficiency and dynamic computation. Instead of a fixed stack of layers, it employs two dense prelude blocks followed by a recurrent block applied up to 24 times.

Dynamic Computation and Routing

  • Adaptive Depth (PonderNet): The model uses a halting head to determine how much computation to spend on each character. Simple characters may require only one row of computation, while complex ones can take up to 24. This ensures compute resources are allocated based on the difficulty of the input.
  • Recurrent MoE Routing: For each application of the recurrent block, the model selects the top-8 experts from a shared pool. Because routing happens per block-application rather than per character, a single character can engage a wide variety of experts across different depths.
  • Byte-Level Processing: The model operates directly on 256 byte values. By eliminating the tokenizer, the model can read any data type without needing a new vocabulary.

Memory Management via Paging

To fit a model larger than the available VRAM, mini-AGI implements a paging system that treats disk space as the primary storage for weights:

  • Disk Storage: Every expert's weights and Adam optimizer moments are stored as individual files on disk.
  • VRAM Working Set: Only a small subset of experts (the working set) is resident in VRAM at any time.
  • Demand-Based Swapping: The model predicts which experts will be needed for the next chunk of text based on the routing patterns of the previous chunk. Experts are swapped between disk, RAM cache, and VRAM to ensure the necessary parameters are available for the forward pass.

Mitigating Catastrophic Forgetting

Continual learning from a single stream of data often leads to "catastrophic forgetting," where new information overwrites old knowledge. mini-AGI addresses this through two primary mechanisms:

Trunk Learning Rate Decoupling

The "trunk" of the model—consisting of embeddings, attention, routers, and the halting head—is trained at a significantly lower learning rate (0.1x) than the experts.

Experimental data shows that while using the same learning rate for both leads to significant loss (forgetting), the 0.1x trunk rate retains 99.84% of progress against chance. This prevents the core routing and structural logic from being overwritten by the specific content of a single data stream.

Structural Isolation

Because the model uses a Mixture-of-Experts (MoE) approach, only a fraction of the model is updated during any given pass. During a probe of 524,000 characters of chess data, only 54 of 136 experts received gradients. This structural isolation ensures that learning a new subject does not necessarily interfere with parameters used for other subjects.

Model Growth and Pruning

mini-AGI is designed to grow and shrink its capacity dynamically based on the data it encounters:

  • Growth via Recombination: When the model lacks capacity, new experts are created by recombining hidden units from existing experts. This ensures new experts start with useful, trained components rather than random noise.
  • Pruning via Usage: Experts that are not selected by the router for extended periods are considered "dead" and are deleted from disk to reclaim space.
  • Growth Constraints: New experts are only added if specific conditions are met, including available disk/VRAM space and the stability of the held-out loss (the "honest" brake).

Performance and Scaling

At the time of reporting, the model had read approximately 318.1 million characters and consisted of 169 experts. Its held-out loss across eight subjects was 0.8336 nats/char (1.2026 bits/byte).

Data Scaling Trends

Scaling analysis indicates the model follows a healthy power law (L ∝ D^-0.239). The author suggests that reaching 0.80 BPB (bits per byte) would require approximately 1.92 billion bytes of data, which is achievable within a few weeks of training on a single laptop GPU.

Community Insights and Critiques

Discussion among technical users highlights both the potential and the limitations of the current implementation:

  • Generalization vs. Memorization: Some users questioned whether the model is truly generalizing or simply performing high-efficiency memorization. One user noted that the model's current bits-per-byte (BPB) performance is higher (worse) than some very small dense models trained on specific datasets like enwik9.
  • Architectural Potential: There is interest in the model's ability to act as a "proto-AGI" due to its continual learning nature, with suggestions to explore nested reinforcement learning or self-similar recursive architectures.
  • Learning Rate Concerns: Critics pointed out that the trunk learning rate reduction may only delay catastrophic forgetting rather than eliminate it, suggesting that the trunk will eventually undergo forgetting over a very long time horizon.

"The expert pool is not what prevents forgetting. Freezing the working set... costs only 0.3571 nats... Preserving most of the weights is not sufficient."

Technical Specifications Summary

Component Detail
Vocabulary 265 (256 bytes + 9 markers)
Context Window 4,096
Core Architecture RMSNorm, RoPE, SwiGLU, Flash Attention
Depth 2 prelude blocks + 1 recurrent block (up to 24 applications)
VRAM Requirement Minimum 8 GB
Parameter Management Paged MoE (Disk $\rightarrow$ RAM $\rightarrow$ VRAM)

Sources

Related