szczyglis-dev/py-gpt

Desktop AI Assistant powered by GPT-5, GPT-4, o1, o3, Gemini, Claude, Ollama, DeepSeek, Perplexity, Grok, Bielik, chat, vision, voice, RAG, image and video generation, agents, tools, MCP, plugins, speech synthesis and recognition, web search, memory, presets, assistants,and more. Linux, Windows, Mac

What it solves

PyGPT is a comprehensive desktop AI assistant that brings the capabilities of various large language models (LLMs) to a local environment. It eliminates the need to rely solely on web-based interfaces by providing a unified desktop application for chatting, document analysis, image/video generation, and automation through agents and plugins.

How it works

The application acts as a frontend and orchestrator for multiple AI providers. It connects to cloud-based models (like OpenAI, Google Gemini, Anthropic Claude, xAI Grok, and DeepSeek) via API keys, or to local models (via Ollama, LlamaIndex, or HuggingFace). It integrates LlamaIndex for RAG (Retrieval-Augmented Generation), allowing users to index and chat with their own local files (PDFs, CSVs, etc.). It also features a plugin system for tool use, including a Python code interpreter, web search, and integrations with platforms like GitHub and Slack.

Who it’s for

It is designed for users who want a powerful, configurable AI workspace on their desktop (Windows, Linux, Mac) without needing to use a browser, as well as those who prefer using their own API keys for cost control and privacy, or those who want to run models entirely locally.

Highlights

  • Multi-Model Support: Compatible with a wide array of providers including OpenAI, Google, Anthropic, xAI, DeepSeek, and local Ollama installations.
  • RAG Capabilities: Integrated LlamaIndex support for chatting with various file types (txt, pdf, csv, html, etc.) and web pages.
  • Extensible Tooling: Built-in plugins for file I/O, code execution via a real-time Python interpreter, and web search (DuckDuckGo, Google, Bing).
  • Diverse Modalities: Supports speech-to-text (Whisper), text-to-speech, image generation, video generation, and real-time camera capture for vision analysis.
  • Agentic Features: Includes a node-based Agents Builder and an "Autonomous Mode" for complex task execution.

관련

  • 프로젝트
  • 프로젝트
  • 프로젝트
  • 프로젝트
  • 프로젝트