bytedance/UI-TARS-desktop

A multimodal AI agent stack that enables natural language control of desktop applications and web browsers through visual recognition and precise GUI interaction.

roboflow/supervision

A computer vision toolkit that provides model-agnostic building blocks for data loading, visualization, and dataset management to accelerate the development of vision applications.

gpustack/gpustack

An open-source GPU cluster manager that orchestrates inference engines and provisions GPU instances for scalable AI model serving.

getmaxun/maxun

Maxun is an open‑source, no‑code web‑data platform that lets you scrape, crawl, and extract information (including from PDFs and images) and output it as clean, AI‑ready formats. It provides visual robots, an LLM‑powered extraction mode, a CLI/SDK, and can be self‑hosted, making it a practical tool for building datasets and APIs for LLM applications.

SimoneAvogadro/android-reverse-engineering-skill

A Claude Code skill for reverse engineering Android apps to extract HTTP APIs and recover original Kotlin class names from obfuscated binaries.

Dryxio/reagent

An AI-powered reverse-engineering agent that uses Ghidra and LLMs to reconstruct and validate C/C++ functions from compiled binaries through an autonomous pipeline.

onlook-dev/onlook

An open-source visual-first code editor for Next.js and TailwindCSS that allows users to design UIs visually in the browser and sync changes directly to the codebase using AI.

Forget-C/Jellyfish

An end-to-end AI production workspace for short dramas that manages the full pipeline from script breakdown and asset consistency to video generation.

ant-research/4DAnyone

4DAnyone is a research‑grade neural model that converts a single‑camera video of a person into dozens of synchronized, view‑consistent videos (6‑24‑48 views) for downstream 4‑D Gaussian Splatting reconstruction. It runs on consumer GPUs (≤24 GB VRAM), offers a fast distilled Turbo variant (5.6× speedup), and provides ready‑to‑use inference scripts and nerfstudio integration.

palmier-io/palmier-pro

A macOS video editor with built-in generative AI and MCP support, enabling AI agents to create and edit content directly on the timeline.

Open-LLM-VTuber/Open-LLM-VTuber

An open-source, voice-interactive AI companion featuring a Live2D avatar and visual perception, capable of running entirely offline for private, real-time conversations.

alchaincyf/darwin-skill

An iterative optimization framework for Agent Skills that uses a 9-dimensional rubric and a ratchet mechanism to ensure skills only evolve toward higher quality.

hang-jin/editaplot

An AI-guided workflow for creating editable scientific figures in Origin, ensuring scientific accuracy through human-in-the-loop data confirmation.

junhoyeo/tokscale

A high-performance CLI tool and visualization dashboard for tracking token usage and costs across multiple AI coding agents.

keslr/keslr_connect

A set of TypeScript packages for integrating with the Keslr network to ensure applications are accessed only by verified humans via trust-graph authentication and private network addressing.

tsunehimatoi/psd2live

An automated pipeline and desktop app that converts layered PSD files into fully rigged Live2D models, featuring adaptive meshing, automatic deformer setup, and AI agent integration.

chatboxai/chatbox

A cross-platform desktop and mobile client for ChatGPT, Claude, and other LLMs that provides a unified interface and local data storage for enhanced privacy.

ChenShuo2004/cs-board

cs-board is a locally‑run AI video‑creation workstation that turns Chinese text, a reference audio clip, and optional style/character images into a fully‑rendered whiteboard‑style MP4. It handles voice cloning (via IndexTTS), script segmentation, AI‑generated illustrations, hand‑drawn animation, subtitles, and audio‑video sync, offering 12 visual templates, custom character/style support, dynamic infographics, and LAN‑based collaborative queues.

1CatAI/1Cat-vLLM

An optimized vLLM fork that restores high-performance LLM inference for NVIDIA Tesla V100 GPUs by rebuilding the attention dataflow and adding speculative decoding.

RapidAI/RapidOCR

An open-source OCR tool that converts PaddleOCR models to ONNX format for fast, low-resource offline deployment across multiple platforms and languages.

ArcReel/ArcReel

An open-source, self-hosted AI video production workbench that transforms scripts and novels into consistent short videos via a structured generative pipeline.

verl-project/verl-omni

A general RL training framework for multimodal generative models, providing fast rollouts and stable post-training for diffusion and omni-modality models.

strelov1/freehire

An open-source IT job aggregator that pulls millions of postings directly from company career pages and provides AI-powered tools for CV tailoring and application tracking.

adshao/flounder

Flounder is an open‑source framework that lets LLM coding agents run autonomous, sandboxed security audits. Given a repo, contract address, transaction hash, or similar clue, it prepares a workspace, maps the attack surface, lets the model write and execute PoC tests in an OCI sandbox, verifies findings locally, optionally confirms them against a real target, and generates markdown reports. It supports many scenarios (blind audits, incident investigations, bug‑bounty work) and works with various model providers via a daemon‑based control plane.

zenbu-labs/terminal-code

A tool that brings the full VS Code editor experience directly into the terminal by combining code-server and terminal-browser.

google/adk-python

A code-first Python framework for building, evaluating, and deploying sophisticated AI agents using graph-based workflows and multi-agent orchestration.

ultracontext/ultracontext

An open-source, real-time context infrastructure that synchronizes AI agent sessions, allowing users to share and version context across different AI tools.

google-deepmind/mujoco_menagerie

A curated collection of high-quality robot models for the MuJoCo physics engine, designed to ensure realistic and faithful simulations.

jundizhou/easy-stock

An AI-powered research workbench for A-share investors that automates market analysis, sentiment tracking, and portfolio inspection using a local-first architecture.

jnMetaCode/superpowers-zh

A Chinese-enhanced version of the superpowers framework that provides 20 systematic work methodologies to help 26 different AI coding tools plan, debug, and review code more professionally.

alexgreensh/attention-span

Attention Span provides three markdown‑based output‑style plug‑ins (Attention‑kind, Spartan, Rundown) for Claude Code and other LLM agents. The styles keep the model’s reasoning unchanged but make replies up to ~43 % shorter and easier to skim, helping users with limited attention and reducing token spend. Install via a one‑step Claude Code plugin or manually drop the markdown files, then activate with `/style` or skill commands. Works with other agents after stripping front‑matter. Complementary tools (Token Optimizer, Outsourcerer) address deeper token waste.

henryqin1997/statem

A command-line state machine that turns AI agent workflows into inspectable graphs of states and executable checks to ensure reliability in long-running tasks.

zts212653/clowder-ai

Clowder AI is a platform layer that orchestrates multiple AI agents from different model families, enabling them to collaborate as a team with shared memory and cross-model review.

ExplosiveCoderflome/AI-Novel-Writing-Assistant

An open‑source, monorepo app that uses LangChain/LangGraph agents to turn a single seed idea into a full‑book outline, generate chapters, review and repair them, and optionally create manga or short‑drama adaptations. It runs on a React + Vite front‑end, Express + Prisma back‑end, supports multiple LLM providers, and offers optional RAG via Qdrant. Designed for novice writers and developers interested in long‑chain AI workflows.

julyx10/lap

A private, local-first photo manager for macOS, Windows, and Linux that uses local AI for search and face clustering without requiring cloud uploads.

alexgreensh/token-optimizer

A token management tool for AI coding assistants that reduces API costs by compressing context waste and preventing data loss during session compaction.

Roboparty/roboto_origin

A fully open-source DIY humanoid robot project providing the hardware designs, RL training workflows, and ROS2 deployment framework needed to build a walking and running robot.

modelscope/ms-swift

A scalable infrastructure for the fine-tuning, evaluation, and deployment of over 1,000 text and multimodal large models.

PrefectHQ/prefab

A generative UI framework for Python that allows developers to build interactive interfaces and MCP Apps using a declarative DSL rendered by a bundled React frontend.

googleapis/mcp-toolbox

An open-source MCP server that connects AI agents and IDEs to enterprise databases through prebuilt generic tools or a customizable framework for secure, structured data access.

NomaDamas/CozyClay

A browser-based previs studio for blocking scenes and posing characters to create high-control reference assets for AI video generation.

HKUDS/RAG-Anything

RAG‑Anything is a Python library that builds a multimodal Retrieval‑Augmented Generation system. It parses PDFs, Office files, images, tables, and equations using MinerU, creates a cross‑modal knowledge graph, and retrieves relevant pieces with a hybrid vector‑graph search before passing them to an LLM for answer generation. Install via `pip install raganything`, configure with `RAGAnythingConfig`, ingest documents, and query in natural language.

data-privacy-stack/presidio

A data protection SDK that identifies and anonymizes Personally Identifiable Information (PII) in text and images using customizable recognizers.

DSH-APP/DSHA

An Android launcher that allows users to run the full DeepSeek Harness agent framework on mobile devices without root or Termux, featuring a bundled Ubuntu environment and system-level device control.

FeiZhuLulu/real-api-pricing

A data-driven project that calculates the real unit price of AI model subscriptions and APIs by normalizing token allowances and costs against a standard workload.

yappologistic/Spun

A visual music player for Linux that simulates physical media like vinyl, CDs, and cassettes to provide an immersive playback experience across multiple music sources.

modelscope/sirchmunk

An indexless, agentic search engine that replaces traditional vector databases with a self-evolving knowledge base for real-time retrieval from raw data.

VectifyAI/PageIndex

PageIndex is an open‑source RAG system that replaces vector‑store retrieval with a hierarchical tree index built from a document’s layout. An LLM reasons over this tree to find relevant sections, giving explainable, citation‑ready answers without any vector DB or chunking. The SDK works locally (≈ $0.001 / page indexing cost) or via PageIndex Cloud for OCR‑heavy, large‑scale corpora. Benchmarks show 98.7 % accuracy on FinanceBench and up to 16× lower query cost versus feeding whole PDFs to a model.