Jev-like Python Wrapper for LLMs and Vision Models Using Token Logprobs

TL;DR – What the wrapper does and why it matters

The author built a minimal Python wrapper that sends a Jev‑style request (state, questions, optional image attachments) to any Chat Completion endpoint, asks the model to emit one token (a letter representing the chosen option), and reads back the log‑probabilities for all candidate tokens. This technique turns a generative LLM into a cheap, low‑latency classifier that works with both text‑only and vision‑augmented models.


Core idea – Use log‑probs to turn generation into classification

  • Prompt format – The wrapper builds a prompt that lists a state, a question, and a set of options labelled [A], [B], …, then ends with the instruction “Answer with the letter of the best option only.” Example:
    State:
    Inspect this webcam frame. Judge only what is visibly present.
    
    Question: Is a person visible?
    Options:
    [A] true
    [B] false
    
    Answer with the letter of the best option only.
    
  • API parameters – The request includes:
    {
      "max_completion_tokens": 1,
      "logprobs": true,
      "top_logprobs": 20,
      "temperature": 0
    }
    
    Setting max_completion_tokens to 1 forces the model to output a single token, which makes the response tiny and the latency low.
  • Log‑prob extraction – The API returns the top‑N token candidates with their log‑probabilities. By mapping each option letter to a token, the wrapper converts those log‑probs into a normalized probability distribution over the options.
  • Result interpretation – For binary questions (noul type) the wrapper returns the probability of true. For multi‑choice questions it returns the most likely option and the full probability table. For ordinal scores it computes an expected value.

Extending Jev to vision – the attachments field

  • The original Jev spec only supports a text/JSON state. The author added an attachments array that can contain either file paths or base64‑encoded data URLs.
  • When the request is sent to an OpenAI endpoint, each image is added as a type: "input_image" element; for llama.cpp it is added as type: "image_url".
  • The example captures a live webcam frame with OpenCV, encodes it as JPEG, wraps it in a data URL, and places it in attachments before each request.

End‑to‑end Python example (≈150 lines)

The script performs three steps in a loop:

  1. Capture a frame from /dev/video0.
  2. Encode the frame to a base64 JPEG and set data["attachments"].
  3. Submit the request to the chosen backend (local llama.cpp server or OpenAI) using a background thread.
  4. Print a table with the latest answers and the measured FPS.

Key functions:

  • score(data, url, model) – builds the prompt for each question, sends the request, normalizes log‑probs, and returns a structured answer dictionary.
  • main – parses CLI arguments (url and model), starts the webcam, and orchestrates the asynchronous scoring.

The script is deliberately self‑contained; the only external dependency is opencv-python for webcam access. All other libraries are part of the Python standard library.


Performance numbers reported by the author

Backend Model Hardware FPS (frames / second)
llama.cpp (local) Gemma‑4‑12B‑QAT (GGUF) RTX 3090 ≈ 1 fps (three questions per frame)
OpenAI gpt‑6‑luna Cloud ≈ 0.2 fps

The author notes that the OpenAI run is slower partly because each frame creates a new HTTP connection per question, which could be optimized.


Setup instructions for the local llama.cpp server

# 1. Download the 7 GB Gemma‑4‑12B model and its multimodal projector (~175 MB)
mkdir -p ~/models/gemma-4-12b/
cd ~/models/gemma-4-12b/
curl -fL -C - -o gemma-4-12b-it-qat-q4_0.gguf \
  https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf/resolve/main/gemma-4-12b-it-qat-q4_0.gguf
curl -fL -C - -o mmproj-gemma-4-12b-it-qat-q4_0.gguf \
  https://huggingface.co/google/gemma-4-12B-it-qat-q4_0-gguf/resolve/main/mmproj-gemma-4-12b-it-qat-q4_0.gguf

# 2. Install the CUDA‑86 llama.cpp binary
curl -fL -o llama.zst \
  https://huggingface.co/buckets/ggml-org/install.sh/resolve/b11160/x86_64/linux/cuda/86/llama-app.zst
mkdir -p ~/bin/
zstd -d llama.zst -o ~/bin/llama
chmod +x ~/bin/llama

# 3. Start the server on port 8060
~/bin/llama serve --models-dir ~/models/ --port 8060

After the server is running, execute the wrapper:

uv run webcam.py http://localhost:8060/v1 gemma-4-12b
# Or use OpenAI (set OPENAI_API_KEY first)
uv run webcam.py https://api.openai.com/v1 gpt-6-luna

Community reactions (selected HN comments)

@TeMPOraL – “This is how we could get Star‑Trek‑style ambient awareness: infer intent from multimodal context and act automatically.”
The comment highlights the broader vision of using such a wrapper for real‑time intent detection beyond simple classification.

@prathje – “Is this just grammar‑based decoding with a JSON response? If so, Jev is essentially a caching layer on top of that.”
The author’s implementation indeed relies on a deterministic prompt and token‑level logits, which is similar to grammar‑constrained decoding, but adds a lightweight caching strategy for repeated state prefixes.

@frabcus – “Normal LLMs trained with RLHF may decide earlier in the network, so a Jev‑style log‑prob wrapper could be less accurate than a model trained specifically for calibrated decisions.”
This points out a potential limitation: the probability calibration of vanilla models may be sub‑optimal for downstream decisions.

@arcticbull – “It’s like Jev but several orders of magnitude more expensive and slower.”
The comment reflects the trade‑off between the flexibility of a generic LLM and the efficiency of purpose‑built vision classifiers.

@CROON_tv – “I’d like to see tail latency numbers; in our real‑time speech‑endpoint, Jev was faster and less hesitant than a plain LLM with the same prompt.”
Latency is a critical metric for interactive applications, and the wrapper’s single‑token approach helps keep tail latency low.

@czl_my – “I’ve created a Jev wrapper that works with any OpenAI‑compatible endpoint: https://github.com/zhulinchng/jevper.”
An external implementation demonstrates that the idea is gaining traction.


Limitations and open questions

  • Probability calibration – The raw log‑probs from a vanilla model may not be well‑calibrated, especially for rare tokens. Users may need temperature scaling or post‑hoc calibration.
  • Missing token handling – The wrapper treats any option whose token is absent from the top‑N list as zero probability, but raises an error if the omitted mass exceeds 1e‑6. This safety check prevents silent mis‑ranking.
  • Scalability – Sending a separate request per question (three per frame in the demo) multiplies network overhead. Batching multiple questions into a single prompt could improve throughput.
  • Vision model choice – While Gemma‑4‑12B works, dedicated multimodal models (e.g., CLIP, Florence) would likely achieve higher FPS and better visual accuracy.

Takeaway

The posted wrapper demonstrates a practical, language‑model‑agnostic method for turning any Chat Completion API—including vision‑augmented backends—into a fast, deterministic classifier by exploiting token‑level log probabilities. Its simplicity (a single Python function) and the ability to add image attachments make it a useful building block for real‑time multimodal decision systems, albeit with the usual caveats around probability calibration and per‑question request overhead.

Sources