Ollaya: Local Inference for Jev-style Decision Models
Ollaya is an open-source inference engine designed to run decision models locally. Unlike traditional Large Language Models (LLMs) that generate text token-by-token, decision models provide calibrated, typed answers to specific questions about a text or JSON input in a single forward pass, enabling millisecond-level latency.
High-Performance Local Decision Inference
Ollaya allows users to run decision models on their own hardware, ensuring that sensitive data—such as customer tickets and emails—remains private. The system is optimized for speed, with a five-question request to the Laya model taking approximately 10ms end-to-end through the HTTP API on an NVIDIA RTX 4090.
Performance Benchmarks
Median latency for a five-question request via the HTTP API on an NVIDIA RTX 4090 is as follows:
| Model | Median Latency |
|---|---|
| laya:multilingual | 8.1 ms |
| laya:en | 9.6 ms |
| gliclass | 14.7 ms |
| nli | 20.4 ms |
| decider:0.8b | 155 ms |
| decider:2b | 190 ms |
For comparison, the TypeSafe hosted API shows a median request latency of 236–276 ms, including network overhead.
TypeSafe API Compatibility
Ollaya is a drop-in replacement for TypeSafe's API, serving the /v1/systemone and /v1/models endpoints. This allows the official TypeSafe Python SDK (v0.7.1) to work without modification by simply updating the TYPESAFE_BASE_URL and TYPESAFE_API_KEY environment variables.
Example Request and Response
A typical request asks the model to categorize an input string (e.g., "Can I get an invoice for last month?") into a predefined set of criteria (e.g., invoice, refund, other).
Request:
{
"model": "laya",
"state": "Can I get an invoice for last month?",
"questions": {
"intent": {
"type": "choice",
"instructions": "What does the customer want?",
"criteria": {
"invoice": "Needs an invoice or receipt",
"refund": "Wants money back",
"other": "Anything else"
}
}
}
}
Response:
{
"model": "laya:en",
"answers": {
"intent": {
"type": "choice",
"choice": "invoice",
"confidence": 0.9547,
"probabilities": {
"invoice": 0.9698,
"refund": 0.0172,
"other": 0.013
}
}
},
"usage": {
"input_tokens": 43,
"output_tokens": 0
}
}
Model Support and Deployment
Ollaya provides access to Laya models from Convai Innovations, including English-specific, multilingual, and router models.
Platform Availability
Ollaya is distributed as a desktop app, command-line tool, and Docker image. It supports the following hardware configurations:
- Linux (x86-64/ARM64): Full support with NVIDIA GPU acceleration (CUDA 13) on x86-64.
- Windows: Support via WSL 2 for NVIDIA GPU acceleration.
- macOS (Apple Silicon): CPU-only execution.
- Docker: NVIDIA GPU acceleration available for amd64 images.
Community Insights and Technical Discussion
Community members on Hacker News discussed the utility and positioning of Ollaya. While some praised the ease of setup and the ability to use older hardware (e.g., a GTX 970), others questioned the actual decision quality of Laya compared to Jev.
"Has anyone actually seen better or the same results with Laya compared to Jev? From my experience, Laya performs significantly worse. It's less confident and often makes wrong decisions with more complex queries."
Other discussions focused on the distinction between decision models and traditional classifiers or re-rankers. Some users suggested that if decision models become prominent, the primary LLM runner, Ollama, may eventually implement native support for them.
Additionally, a community-developed web UI for the Ollaya API was released to complement the tool: ollaya-web-ui.
Sources
Related
- Dispatch
- Project
- Project
- Project
- Dispatch