Yinsongxu/LLM2Jev

Turn local language models into Jev-style structured decision models. Get results from text and images with prefill alone—no token-by-token decoding required.

🧠 What is LLM2Jev?

LLM2Jev is an open‑source library that lets you run a local large language model (LLM) as a Jev‑style decision engine. Instead of generating text token‑by‑token, it prefills the model once, reads the logits, and directly computes probabilities for a set of answer candidates. The result is a structured decision output (e.g., scores, choices) that can be consumed by applications that expect the Jev /System One API.

The project supports three back‑ends:

  • SGLang – a high‑performance inference server that can cache prefixes (Radix Cache) and score many candidates efficiently.
  • Transformers – the classic Hugging Face pipeline.
  • MLX – Apple‑Silicon‑native inference for text and vision models.

It works with pure text as well as multimodal inputs (text + images) and can be accessed via a Python API or an HTTP endpoint compatible with Jev’s /v1/systemone API.


✨ Core Features (as described in the README)

Feature What it means
Broad backend support Run local LLMs (text‑only or vision‑language) through SGLang, Transformers, or MLX on Apple Silicon.
Prefill‑only inference The model is run once to produce logits; candidate probabilities are derived from those logits without iterative decoding.
Apple Silicon acceleration MLX backend provides quantized models, batching, and prefix‑reuse on M‑series Macs.
Multimodal inputs You can feed images together with text in the request’s state or instructions.
Order‑independent scoring Each candidate is evaluated independently, so shuffling options does not affect scores.
Prefix reuse on cold requests For long prompts with many candidates, the first candidate builds a cache that subsequent candidates reuse, cutting redundant work.
Jev‑compatible HTTP service Exposes a POST /v1/systemone endpoint that mimics the official Jev API.

🚀 Quick‑Start (Linux + NVIDIA GPU example)

# Clone and install extra dependencies for the SGLang backend
git clone https://github.com/Yinsongxu/LLM2Jev.git
cd LLM2Jev
uv sync --extra sglang   # or `pip install -e .[sglang]`
source .venv/bin/activate

# Run the demo script with a local causal model (HF format)
python examples/sglang_inference.py --model-path /path/to/model

The script sends a Choice, Score, and Noul question to the model and prints the JSON response.

For Apple‑Silicon users, the same command works with the MLX backend (see the installation docs). Image‑based requests are covered in docs/multimodal.md.


📦 Installation

The repository provides a detailed guide (docs/installation.md) that lists:

  • Python 3.12+ requirement
  • Optional extras: sglang, transformers, mlx
  • System libraries for CUDA (Linux) or Metal (macOS)
  • Recommended package manager uv (fallback to pip)

📖 Getting Started

The Usage guide (docs/usage.md) walks you through:

  • Offline Python API – direct function calls for scoring.
  • Online HTTP service – start a server that accepts Jev‑compatible requests.
  • Choosing between the staged (prefix‑reuse) and all (no reuse) scoring strategies.
  • How to construct multimodal payloads.

🎮 Demos

Demo Description
Web demo Interactive UI to submit questions and view probability breakdowns. (demos/web/README.md)
Snake demo A tiny game where the LLM decides the snake’s moves based on the current board state. (demos/snake.py)
MuJoCo pick‑and‑place Shows a robot arm picking objects using LLM‑driven decision making. (demos/pick_place/README.md)

Animated GIFs are included in the README to illustrate each demo.


📊 Benchmarks

Performance numbers for the Qwen3‑1.7B model on an RTX 5090 are documented in docs/shared-prefix-benchmarks.md. The benchmarks compare:

  • staged (prefix reuse) vs. all (no reuse)
  • Cold vs. warm cache scenarios
  • Impact of input length and candidate count on latency and throughput.

🗺️ Roadmap (current status)

  • ✅ Interactive web demo completed
  • ✅ Initial local‑image support for Transformers and SGLang
  • ⬜ Expand benchmarks across more model sizes and datasets
  • ⬜ Add more multimodal tasks and demos
  • ⬜ Deeper evaluation of decision quality vs. latency

🧪 Tests

Run the test suite with:

python -m unittest discover -s tests -v

The repository includes unit tests for the scoring pipeline and HTTP service.


📄 License

Apache License 2.0 – free for commercial and academic use.


TL;DR: LLM2Jev turns any locally‑run LLM (text or vision) into a fast, Jev‑compatible decision engine by scoring candidate answers directly from a single prefill pass. It supports SGLang, Transformers, and Apple‑Silicon MLX backends, offers a Python API and an HTTP service, and includes demos ranging from a web UI to a robot‑arm simulation.

Related

  • Project
  • Project
  • Project
  • Project