LYiHub/pub-local-jarvis

Windows 本地多模态 AI 桌面桌宠,支持屏幕与音频感知。

What it solves

AI Jarvis is a local, multimodal desktop assistant for Windows that provides real-time companionship and utility without requiring data to be uploaded to the cloud. It solves the problem of AI assistants being disconnected from the user's actual activity by continuously perceiving the screen and system audio to offer contextual help, gaming tips, or educational notes.

How it works

The system uses the MiniCPM-o 4.5 GGUF model, integrating a Large Language Model (LLM), a Vision Model (VPM), and an Audio Model (APM). It captures screen frames via DXGI and system audio via WASAPI approximately once per second. Instead of a simple Q&A format, the model operates in a full-duplex mode, deciding whether to continue listening (LISTEN) or provide a brief, context-aware response (SPEAK) based on the current scene.

Who it’s for

  • Gamers who want real-time, non-intrusive hints and interaction via transparent overlays.
  • Students who need automated course recording, key-frame capture, and Markdown knowledge summaries.
  • General Windows users seeking a privacy-focused AI companion that can "see" and "hear" their desktop activity locally.

Highlights

  • Local Multimodal Inference: Processes screen and audio on-device for speed and privacy.
  • Active Dialogue: Allows users to trigger a chat window (Ctrl+M) to ask questions about the current screen state.
  • Course Recording: Automatically generates Markdown notes with key images and knowledge points from online classes.
  • Gaming Companion: Provides scene-specific tips and interactions via transparent bullet-chat overlays.
  • Local Memory: Organizes activity timelines and daily summaries without storing raw screen or audio data long-term.
  • Privacy Mode: Allows users to pause screen and audio perception with a double-click on the desktop pet.

Related

  • Project
  • Project
  • Project
  • Project
  • Project