mizorewww/laya-coreml

Local Laya typed decisions on Apple Core ML and Neural Engine. Validated ports, ~5 ms short decisions on M3 Max, reproducible speed and energy benchmarks.

What it solves

Laya-CoreML provides a way to run "typed decision" models on Apple Silicon with extremely low latency and high energy efficiency. It eliminates the need for autoregressive decoding (generating tokens one by one), which allows the model to return probabilities, scores, and boolean answers instantly without the overhead of parsing generated JSON or text.

How it works

The project ports the Laya model to Apple's Core ML framework, specifically optimizing for the Apple Neural Engine (ANE). By using a non-generative approach to decision-making, it can produce a single-shot prediction for a given input. The ANE-optimized versions use specific architectural adjustments (like BC1L activations and 1x1 projections) to ensure the workload runs on the Neural Engine rather than just the CPU or GPU, significantly reducing power consumption and increasing speed.

Who it’s for

  • Developers building real-time agents on macOS (e.g., the included Snake game demo).
  • Applications requiring fast, low-power classification or decision-making on Apple hardware.
  • Users who want to run multilingual decision models locally without needing heavy dependencies like PyTorch or Transformers.

Highlights

  • Neural Engine Optimization: Achieves significantly lower energy per decision compared to MLX FP16 implementations.
  • Zero Token Generation: Returns direct probabilities and typed answers instead of generating text.
  • High Performance: Capable of sustaining ~50 decisions per second in a game loop on M3 Max hardware.
  • Standalone Inference: Does not require PyTorch, Transformers, or MLX to run inference.

Related

  • Project
  • Project
  • Project
  • Project