ollaya-dev/ollaya
Run open decision models locally: pull and serve Laya, decider, NLI and GLiClass behind a TypeSafe-compatible API. Ollama for decision models.
What it solves
Ollaya provides a way to run "decision models" locally. Unlike standard LLMs that generate text, decision models are designed to read a state (like an email or a ticket) and return calibrated probabilities for specific, typed questions (such as whether a request is urgent or if a user is frustrated) in a single forward pass. This allows for fast, exact, and calibrated classification and routing without the need for text generation.
How it works
Ollaya operates as a local daemon that manages and serves these models. It uses a CLI similar to Ollama, allowing users to pull, run, and create models. It supports multiple inference engines, including ONNX Runtime (for CPU and NVIDIA GPUs) and llama.cpp (for GGUF models on CPU, CUDA, and Metal).
To ensure efficiency, Ollaya publishes small ONNX graphs that point to original weight files on Hugging Face, verifying them via sha256. It also supports "Modelfiles" to bake specific question sets into a custom model.
Who it’s for
Developers building agents or automated workflows that require fast, calibrated classification, intent detection, and safety guarding without the overhead of text generation.
Highlights
- TypeSafe-compatible: Wire-identical to TypeSafe's
/v1/systemoneAPI, allowing existing clients to switch to local hosting. - Agent Integration: Includes an MCP server for integration with Claude Code, Claude Desktop, and Cursor.
- laya Router: A built-in router that detects language and scripts to route requests to the same-language model.
- High Performance: Capable of extremely low latency (e.g., 8-10ms for five questions on an RTX 4090).
- Flexible Weights: Supports both ONNX and GGUF formats, pulling weights directly from authors' Hugging Face repositories.
Related
- Project
- Project
- Project
- Project