Laya CoreML Demo on Apple Silicon M4 – Offline Snake Game Running via CoreML
Laya CoreML runs offline on Apple Silicon M4 with minimal CPU load
Takeaway: The laya-coreml package can execute a multilingual LLM locally on a Mac equipped with an M4 chip, using CoreML to offload inference to the Neural Engine and keeping memory usage under 800 MiB while playing a Snake game demo.
Quick start script
The gist author posted a reproducible shell script (laya.sh) that sets up a fresh Python environment, installs the demo, downloads a pre‑converted CoreML model, and launches the Snake demo:
mkdir test-laya
cd test-laya
uv init # create an isolated Python environment
uv add 'laya-coreml[demo]' # install the library and demo extras
hf download aac6fef/laya-multilingual-coreml-ane --local-dir models/snake
uv run laya-coreml-snake --model models/snake
Running the script on macOS 27.0 (Apple Silicon M4) launches a window where the LLM controls the Snake agent in real time.
Performance metrics from the author
The author captured a sample snapshot of the Python process while the demo was running:
Physical footprint: 560.4 MiB
Physical footprint (peak): 778.0 MiB
The report shows that the process never exceeded 800 MiB of RAM, confirming that the model fits comfortably within the unified memory pool of typical M‑series Macs.
CoreML off‑loading to the Neural Engine
A commenter (speedping) observed that the inference runs almost entirely on the Neural Engine rather than the GPU:
"So cool. I've fired up pumas (energy monitor) and it seems to run almost fully on the neural engine and not the GPU so it plays really nicely with CoreML."
This confirms that laya-coreml leverages Apple’s hardware acceleration stack, reducing CPU load and power consumption.
Example API usage via Cloudflare
Another community member (coezbek) wrapped the local Laya model behind a Cloudflare Workers endpoint (typesafe/jev API) and demonstrated a simple JSON request:
curl https://laya.inference.zaitlabs.com/ai/run \
-H "Content-Type: application/json" \
-d '{
"state": "Help! My payouts have been failing for 3 days.",
"questions": {
"is_urgent": {"type": "noul", "instructions": "Does this convey urgency?"}
}
}'
The response returned a probability score for urgency (0.7894), showing that the model can be exposed as a low‑latency inference service.
Community perspective
- Deterministic tasks: A comment by
imranqnotes that Laya shines on tasks with some training data, where deterministic behavior is valuable, but may lag behind zero‑shot models like Jev for purely novel queries. - Local LLM future:
PaulRobinsonexpressed enthusiasm for local LLMs in control‑problem domains, suggesting that hardware‑accelerated inference (as demonstrated) could drive broader adoption. - Memory curiosity:
altanoasked about memory consumption on an M3 Max with 128 GB unified memory. The author’s sample data (peak ~778 MiB) indicates that even on lower‑end M‑series chips the footprint is modest. - Model scope: Several commenters (
tentacleuno,frag) queried whether the Snake demo used a fine‑tuned Laya model. The gist does not specify fine‑tuning; it uses the pre‑publishedlaya-multilingual-coreml-anemodel from Hugging Face.
What Laya‑CoreML is
laya-coreml is a Python package that wraps a multilingual LLM (approximately 300 M parameters) and provides a CoreML conversion pipeline. The repository (https://github.com/mizorewww/laya-coreml) includes:
- Model conversion scripts for Apple’s CoreML format.
- A demo CLI (
laya-coreml-snake) that runs the model in a reinforcement‑learning‑style Snake game. - Optional extras (
[demo]) that install the necessary dependencies for the demo.
Why this matters
Running a language model entirely offline on consumer‑grade Apple Silicon demonstrates that sophisticated AI workloads no longer require cloud GPUs. The low memory footprint and Neural Engine acceleration make it feasible to embed LLM‑driven features—such as intent classification, urgency detection, or game agents—directly into macOS applications without network latency or data‑privacy concerns.
How to try it yourself
- Prerequisites: macOS 27.0 (or later) on Apple Silicon (M1/M2/M3/M4). Install the
uvpackage manager (pip install uv). - Clone the repo:
git clone https://github.com/mizorewww/laya-coreml.git. - Follow the script: Execute the commands from the "Quick start script" section above.
- Experiment: Replace the Snake demo with your own CoreML‑compatible inference loop, or expose the model via a local HTTP server as shown in the Cloudflare example.
Limitations and open questions
- Model size: At ~300 M parameters, Laya is far smaller than state‑of‑the‑art models (e.g., 7 B+). Expect reduced language understanding breadth.
- Fine‑tuning: The public demo does not include a fine‑tuned version for the Snake environment; users interested in task‑specific performance will need to fine‑tune the base model and re‑export to CoreML.
- Hardware dependence: Performance gains rely on Apple’s Neural Engine; results may differ on non‑Apple hardware or older Mac models lacking the latest Neural Engine.
Bottom line
The laya-coreml demo proves that a 300 M‑parameter multilingual LLM can run offline on an Apple Silicon M4, using CoreML to keep RAM under 800 MiB and off‑load compute to the Neural Engine, enabling low‑latency, privacy‑preserving AI applications on the desktop.
Sources
Related
- Project
- Project
- Dispatch
- Dispatch
- Dispatch