Janus: Local LLM Server with Vulkan Support and OpenAI-Compatible API

Janus is a single Go binary designed to run .gguf models locally on GPU or CPU, providing an OpenAI-compatible API for seamless integration with clients like Cursor and Cline. By leveraging llama.cpp via Vulkan, Janus enables hardware-accelerated inference across AMD, Intel, and NVIDIA GPUs without requiring Python, Docker, or the Ollama runtime.

Core Features and Technical Capabilities

Janus provides a lightweight wrapper around llama.cpp to simplify the deployment of local models. Its primary technical advantages include:

  • Cross-Vendor GPU Acceleration: Uses the Vulkan backend to support a wide range of hardware, including AMD, Intel, and NVIDIA GPUs, with a CPU fallback for systems without compatible graphics hardware.
  • OpenAI-Compatible API: Implements standard endpoints such as /v1/chat/completions and /v1/models, allowing it to act as a drop-in replacement for OpenAI for any compatible client.
  • Dynamic Model Management: Supports hot-swapping .gguf models via the /models/load endpoint without needing to restart the server.
  • Reasoning Model Support: Specifically handles "thinking" models by splitting <think> reasoning blocks into a dedicated reasoning_content field.
  • Automated Configuration: Automatically detects chat templates from GGUF metadata, reducing the need for manual prompt formatting.

System Requirements and Deployment

Janus is designed for minimal dependency overhead, requiring only Go 1.22+ for building from source.

Platform Support

Platform Requirements
Windows Windows 10/11, Go 1.22+, Vulkan-capable GPU recommended
Linux Go 1.22+, Vulkan or CPU
macOS Go 1.22+, CPU backend (Vulkan support varies)

Configuration Parameters

Users configure the server via a .env file. Key variables include:

  • INFERENCE_BACKEND: Set to vulkan, cpu, or openrouter.
  • JANUS_GPU_LAYERS: Controls GPU offloading (-1 for all layers on GPU, 0 for CPU only).
  • JANUS_VRAM_CEILING_MB: Provides a VRAM budget hint to manage memory allocation.
  • JANUS_MODEL_PATH: The absolute or relative path to the .gguf model file.

API Endpoints

Janus exposes several endpoints for health monitoring, model management, and inference:

Method Path Description
GET /health Liveness check (use ?deep=true for detailed status)
GET /v1/models Lists available models
POST /v1/chat/completions Chat interface with streaming support
POST /models/load Hot-swaps the active model
GET /models/list Lists available .gguf files in the models directory
GET /engine/status Returns current VRAM usage and backend information

Community Analysis and Critical Perspectives

While Janus simplifies the setup process for some users, the technical community on Hacker News has raised several points regarding its utility and performance:

Redundancy with llama.cpp

Some users argue that Janus provides little additive value over the native llama-server provided by llama.cpp.

"This seems to be a wrapper around libllama, what's the point? Llama.cpp already ships a web server. I don't see anything in the README that llama.cpp doesn't already support."

Hardware Performance

There are concerns regarding the efficiency of the Vulkan backend, particularly on Intel hardware, where it may introduce significant overhead compared to other backends.

"From what I've seen, Vulkan adds a lot of overhead on Intel hardware."

Lack of Benchmarks

Critics noted the absence of performance data comparing Janus to high-performance inference engines like vLLM, SGLang, or ExLlama, making it difficult to assess its efficiency for production or high-throughput use cases.

Sources

Related

  • Dispatch
  • Project
  • Project
  • Project
  • Project