Mapika/decider

A family of System One-style models fine-tuned from Qwen3.5, designed for one-pass typed decisions with calibrated probabilities.

decider – One‑pass typed decisions with calibrated probabilities

What it is – decider is a family of language‑model checkpoints that don’t generate free‑form text. Instead they take a state (plain text or JSON) and a list of typed questions and, in a single forward pass, output a probability distribution for each question. The supported question types are:

  • Choice – pick one option from 2‑255 candidates.
  • Score – assign a probability to a small ordered set of levels (2‑10).
  • Noul – give the probability that the answer is “yes”.

Because the model never has to decode tokens beyond a fixed set of label tokens, inference is fast and the probabilities are calibrated (the model’s confidence matches observed accuracy after a small RL‑based fine‑tuning stage).


Key capabilities

Feature Why it matters
One‑pass inference All questions are answered together; no iterative prompting or chain‑of‑thought needed.
Typed output Guarantees that the answer is always one of the predefined options – no post‑processing required.
Calibration‑aware RL (v10) Improves the match between confidence scores and true correctness, especially on decision‑making tasks.
Multiple model sizes 0.8 B, 2 B, 4 B (dense) and 35 B mixture‑of‑experts, plus a 2 B vision‑language variant.
Hardware flexibility Runs on CUDA (bf16, optional FP8), Apple Silicon via MPS/MLX, and CPU (eager mode).
HTTP server Simple POST /v1/systemone (TypeSafe wire format) or POST /decide endpoints; schema‑caching for repeated schemas speeds up repeated calls.
Open‑source training pipeline Scripts to download ~95 public decision datasets, build a mixture, and fine‑tune a Qwen‑based backbone.

Typical use cases

  • Customer‑service routing – feed a ticket and policy text, ask “Which department should handle this?” and get a calibrated choice with confidence.
  • Risk scoring – ask “How frustrated is the user?” or “What is the likelihood of fraud?” and receive a probability per score level.
  • Game‑playing agents – the repo includes demos where the model decides moves in text‑based games, Atari Pong (from RAM‑derived text), and Super Mario Bros, all with a single forward pass per move.
  • Vision‑question answering – the decider-2b-vision checkpoint can answer image‑based queries, returning probabilities for each option.
  • Batch classification pipelines – the HTTP server’s schema cache makes it cheap to run the same set of questions over many records.

Quick start (Python)

pip install decider-ai            # or: pip install -e .[serve,train]
from decider.infer import Decider
# Load a checkpoint (downloads from Hugging Face on first use)
model = Decider("Mapika/decider-2b")

state = {
    "ticket": {"messages": [{"from": "customer", "text": "I was charged twice for order A‑104."}]},
    "refund_policy": "Duplicate charges are eligible for a refund."
}
questions = {
    "department": {"type": "choice", "instructions": "Which team should handle this?",
                    "criteria": {"returns": "Exchanges, refunds, wrong or damaged items",
                                 "billing": {"what": "Charges, invoices", "not_for": "delivery"},
                                 "other": None}},
    "refund_requested": {"type": "noul", "instructions": "Does the ticket request a refund?"},
    "frustration": {"type": "score", "instructions": "How frustrated is the customer?",
                    "criteria": ["calm", "frustrated", "very frustrated"]}
}

result = model.system_one(state, questions)
print(result)

The output contains a probability distribution for each typed question, e.g. a choice with confidence and the full vector of option probabilities.


Serving via HTTP

scripts/serve.sh Mapika/decider-2b 8000
  • POST /v1/systemone – compatible with TypeSafe AI SDKs.
  • POST /decide – plain JSON format used in the README examples.
  • The server automatically picks the best device (CUDA → MPS → CPU) and, on CUDA, captures a CUDA graph per (batch, length) shape for zero‑runtime compilation overhead.

Training your own model

The repository ships the full data‑building and fine‑tuning code:

  1. scripts/train.sh full builds the public mixture (≈1.5 M examples, 455 M tokens) and runs one epoch on a Qwen‑3.5 base.
  2. scripts/train.sh delta <existing‑ckpt> continues training from a checkpoint, useful for adding new decision datasets.
  3. An optional RL stage (v8 → v10) further calibrates probabilities using live MiniWoB++ browser tasks; the RL loop lives in a separate research repo but the README links to the required environment.

Known limitations (as stated in the repo)

  • No multi‑step reasoning – the model cannot perform chained arithmetic or multi‑hop inference; split such problems into separate questions.
  • Calibration drops on hard items – especially for knowledge‑heavy multiple‑choice (e.g., GPQA, GSM8K). The 2 B model shows noticeable over‑confidence on the hardest benchmark items.
  • English‑only – all training data and evaluation are in English; performance on other languages is undocumented.
  • Schema‑cache trades accuracy for speed – enabling the cache can reduce answer quality for some schemas.
  • Vision variant still under development – decider-2b-vision uses older text weights and is being retrained.
  • Teacher bias – the custom‑question data is labeled by a 27 B teacher model, which introduces some bias (≈72 % agreement with its own labels).

Where to find more

  • Model cards on Hugging Face: Mapika/decider-2b, decider-4b, decider-35b-a3b, etc.
  • Detailed benchmark tables: docs/RESULTS.md.
  • Change log & RL details: docs/CHANGELOG.md, docs/RL.md.
  • Demo notebooks and example programs: examples/.

Bottom line – decider provides a practical, open‑source alternative to traditional text‑generation LLMs when you need fast, calibrated decisions over a fixed set of options. It is especially useful for routing, scoring, and simple game‑playing tasks, and it can be run on GPUs, Apple Silicon, or even CPUs.

Related

  • Project
  • Project
  • Project
  • Project