Janus: Local LLM Server with Vulkan Support and OpenAI-Compatible API
Janus is a single Go binary designed to run .gguf models locally on GPU or CPU, providing an OpenAI-compatible API for seamless integration with clients like Cursor and Cline. By leveraging llama.cpp via Vulkan, Janus enables hardware-accelerated inference across AMD, Intel, and NVIDIA GPUs without requiring Python, Docker, or the Ollama runtime.
Core Features and Technical Capabilities
Janus provides a lightweight wrapper around llama.cpp to simplify the deployment of local models. Its primary technical advantages include:
- Cross-Vendor GPU Acceleration: Uses the Vulkan backend to support a wide range of hardware, including AMD, Intel, and NVIDIA GPUs, with a CPU fallback for systems without compatible graphics hardware.
- OpenAI-Compatible API: Implements standard endpoints such as
/v1/chat/completionsand/v1/models, allowing it to act as a drop-in replacement for OpenAI for any compatible client. - Dynamic Model Management: Supports hot-swapping
.ggufmodels via the/models/loadendpoint without needing to restart the server. - Reasoning Model Support: Specifically handles "thinking" models by splitting
<think>reasoning blocks into a dedicatedreasoning_contentfield. - Automated Configuration: Automatically detects chat templates from GGUF metadata, reducing the need for manual prompt formatting.
System Requirements and Deployment
Janus is designed for minimal dependency overhead, requiring only Go 1.22+ for building from source.
Platform Support
| Platform | Requirements |
|---|---|
| Windows | Windows 10/11, Go 1.22+, Vulkan-capable GPU recommended |
| Linux | Go 1.22+, Vulkan or CPU |
| macOS | Go 1.22+, CPU backend (Vulkan support varies) |
Configuration Parameters
Users configure the server via a .env file. Key variables include:
INFERENCE_BACKEND: Set tovulkan,cpu, oropenrouter.JANUS_GPU_LAYERS: Controls GPU offloading (-1for all layers on GPU,0for CPU only).JANUS_VRAM_CEILING_MB: Provides a VRAM budget hint to manage memory allocation.JANUS_MODEL_PATH: The absolute or relative path to the.ggufmodel file.
API Endpoints
Janus exposes several endpoints for health monitoring, model management, and inference:
| Method | Path | Description |
|---|---|---|
GET |
/health |
Liveness check (use ?deep=true for detailed status) |
GET |
/v1/models |
Lists available models |
POST |
/v1/chat/completions |
Chat interface with streaming support |
POST |
/models/load |
Hot-swaps the active model |
GET |
/models/list |
Lists available .gguf files in the models directory |
GET |
/engine/status |
Returns current VRAM usage and backend information |
Community Analysis and Critical Perspectives
While Janus simplifies the setup process for some users, the technical community on Hacker News has raised several points regarding its utility and performance:
Redundancy with llama.cpp
Some users argue that Janus provides little additive value over the native llama-server provided by llama.cpp.
"This seems to be a wrapper around libllama, what's the point? Llama.cpp already ships a web server. I don't see anything in the README that llama.cpp doesn't already support."
Hardware Performance
There are concerns regarding the efficiency of the Vulkan backend, particularly on Intel hardware, where it may introduce significant overhead compared to other backends.
"From what I've seen, Vulkan adds a lot of overhead on Intel hardware."
Lack of Benchmarks
Critics noted the absence of performance data comparing Janus to high-performance inference engines like vLLM, SGLang, or ExLlama, making it difficult to assess its efficiency for production or high-throughput use cases.
Sources
Related
- Dispatch
- Project
- Project
- Project
- Project