NandhaKishorM/laya

Non-autoregressive System 1 decision engine. Typed choice, score and yes/no decisions over any text in a single forward pass, in 100+ languages, with a router that picks the right checkpoint per request.

Laya – Fast, multilingual, non‑autoregressive decision engine

What it is – Laya is a Python library (with a matching TypeScript package) that lets you ask typed questions about any piece of text, JSON, email, ticket, etc. Instead of generating free‑form text, it returns structured decisions such as:

  • choice – pick one label from a list (e.g., which department should handle a request)
  • score – give a numeric rating (e.g., urgency on a 0‑100 scale)
  • noul – a binary yes/no question (e.g., “does the user request a refund?”)

All of this happens in a single forward pass of a BERT‑style encoder, so a single question costs ~33 ms on an NVIDIA T4 GPU (7 ms when batched). The model never generates text, which eliminates hallucinations and makes the output directly usable for downstream automation.


Key ideas

Feature Why it matters
Non‑autoregressive No token‑by‑token generation → deterministic, instantly parsable answers.
Typed decisions Structured output (choice/score/noul) fits well with routing, triage, guardrails, and analytics pipelines.
Multilingual (100+ languages) A single request can be in any language; Laya automatically routes to the best checkpoint.
Router Detects script/language and selects one of three checkpoints (English, multilingual, or a larger “typed‑decisions” model) without extra latency.
Routed batches Groups heterogeneous requests by checkpoint and question schema, scoring many items together while preserving order.
ONNX & torch.compile support Optional fast paths for CPU, GPU, Apple Silicon, and Intel XPU.
Prediction hooks Plug‑in points to audit, redact, cache or gate results – useful for compliance or custom guardrails.
Self‑hosted HTTP server A Jev‑compatible API (/predict, /predict/batch) and a small web GUI for quick testing.
LangChain / LangGraph integration Drop‑in agents for LLM‑orchestrated workflows.
Docker & ARM64 builds Easy to run in containers on cloud or edge hardware.

How it works (high‑level)

  1. Router creation – router = Router(preload=True) loads the three checkpoints (or lazily loads them on first use). The router also contains a lightweight language‑detection module that runs in a separate process, so importing laya does not immediately pull in PyTorch.
  2. Define a question set – a JSON‑compatible dict where each key is a question name and each value specifies:
    • type: choice, score, or noul
    • instructions: natural‑language prompt for the model
    • criteria: either a mapping of label → description (choice) or an ordered list of labels (score).
  3. Predict – router.predict(state, questions) runs the appropriate checkpoint in one forward pass and returns:
    • answers – the model’s decision for each question, with confidence scores.
    • routing – which checkpoint was used and why (script, language, fallback, etc.).
  4. Batching – router.predict_batch(requests) automatically groups compatible requests, so a batch of 10 mixed‑language tickets still only incurs a few forward passes.

Quick start (Python)

from laya import Router

router = Router(preload=True)   # load all checkpoints for sub‑35 ms routing

state = {
    "from": "user@example.com",
    "subject": "Duplicate charge",
    "body": "We were billed twice for March. Please refund it today."
}

questions = {
    "department": {
        "type": "choice",
        "instructions": "Which department should handle this request?",
        "criteria": {
            "billing": "invoices, payments, refunds",
            "technical": "bugs, outages",
            "sales": "pricing, contracts",
            "other": "everything else"
        }
    },
    "urgency": {
        "type": "score",
        "instructions": "How urgent is this?",
        "criteria": ["not urgent", "soon", "critical"]
    },
    "refund_requested": {"type": "noul", "instructions": "Does the user explicitly request a refund?"}
}

result = router.predict(state, questions)
print(result["answers"]["department"]["choice"])   # → billing (confidence 0.94)
print(result["routing"])

The same code works for Hindi, Arabic, or any other language; the router will automatically pick the multilingual checkpoint.


Command‑line usage

After pip install laya you get a laya executable:

# routing only – no model download, returns instantly
laya "I was charged twice, please refund"

# full prediction (downloads the checkpoint on first run)
laya "I was charged twice, please refund" --predict

# force a language (useful when detection is ambiguous)
laya "Mein Konto wurde zweimal belastet" --lang de

# use a preset workflow (triage, moderation, etc.)
laya "Refactor this service" --preset triage

The CLI prints a concise JSON summary of the decisions.


Running a local API / demo UI

pip install "laya[serve]"
python examples/server.py   # starts FastAPI at http://127.0.0.1:8000
  • The web UI lets you paste raw JSON or fill a simple form and visualises each answer as a 0‑100 bar.
  • The same server exposes /predict and /predict/batch endpoints that accept the same state/questions payload as the Python SDK.

Installation notes

  • Python 3.10+ – the package pins transformers>=5, torch>=2.14, and huggingface_hub>=1. Use the official PyTorch installer for CPU‑only or GPU‑specific builds before installing Laya.
  • Optional extras –
    • laya[serve] – FastAPI server and UI.
    • laya[mcp] – optional MCP (model‑control‑plane) server.
    • laya[langchain] – integration helpers for LangChain / LangGraph.
  • Model download – the first call that needs a checkpoint pulls the model from the Hugging Face hub (convaiinnovations/laya*). Subsequent runs cache the files locally.

When to use Laya

Use case Why Laya fits
Customer‑support triage – route tickets to the right department, flag urgency, detect refund requests. Structured outputs, sub‑50 ms latency, multilingual out‑of‑the‑box.
Automated moderation – binary “is this abusive?” or “does this contain personal data?” checks. No generation → no false positives from hallucination; confidence scores let you set thresholds.
Business rule enforcement – evaluate policy compliance on incoming emails or JSON payloads. Decision‑head model can be fine‑tuned on your own labeled data; hooks let you add custom redaction or caching.
Edge or low‑resource inference – runs on CPU, Apple M‑series, or Intel XPU via ONNX. Fast loading (2 s CPU load) and optional torch.compile / ONNX paths.

Roadmap & community

  • The repository follows semantic versioning; the latest release (0.3.11) adds routed batches, prediction hooks, and safer HTTP serving.
  • A TypeScript package (laya-ts) provides the same decision API for Node.js and browser environments.
  • The project is Apache‑2.0 licensed, and the author accepts donations via Buy‑Me‑A‑Coffee.

TL;DR

Laya = fast, typed, multilingual decision‑making without any text generation. Import the Router, define a small set of structured questions, and get instant, confidence‑scored answers for emails, tickets, or any free‑form text. It’s ideal for production‑grade triage, guardrails, and any workflow that needs deterministic, language‑agnostic decisions.

Written about in

Related

  • Project
  • Project
  • Project
  • Project