comfyui_LLM_party: a comprehensive suite of ComfyUI nodes for building complex LLM and agentic workflows
A comprehensive set of ComfyUI nodes for building LLM workflows, enabling the integration of AI agents, RAG, and VLMs into visual image generation pipelines.
OpenManus-RL: an RL-based agent tuning framework for enhancing LLM reasoning and decision-making
An open-source framework for optimizing LLM agents using reinforcement learning to enhance their reasoning, tool use, and decision-making capabilities.
llm-beginner: a progressive hands-on curriculum for mastering LLMs and AI agents through from-scratch implementations
A comprehensive, hands-on tutorial series for beginners to learn LLMs and AI agents through six progressive implementation tasks, from building a mini-GPT to creating a coding agent.
SkillOpt: an executive strategy for self-evolving agent skills using a deep-learning-style optimization loop
SkillOpt is a framework that treats agent skill documents as trainable states, using a deep-learning-style optimization loop to evolve and improve agent performance without modifying model weights.
mrpt: a comprehensive C++ framework for mobile robotics featuring SLAM, navigation, and 3D visualization
A robust, modular C++ framework for mobile robotics providing essential tools for SLAM, navigation, and sensor integration.
unrealcv: a bridge between Unreal Engine and AI frameworks for creating synthetic computer vision environments
UnrealCV is a plugin for Unreal Engine that allows computer vision researchers to programmatically interact with virtual worlds to generate synthetic data and PyTorch/Tensorflow integration.
paperlib: an open-source academic paper manager with AI-powered summarization and semantic library search
An open-source academic paper management tool that simplifies metadata scraping and organization, featuring LLM-powered summarization and semantic search.
bgslibrary: a comprehensive C++ framework for foreground-background separation in video streams
A comprehensive C++ framework for background subtraction in computer vision that provides over 40 algorithms to detect moving objects in video streams.
emgucv: a cross-platform .NET wrapper for the OpenCV image-processing library
A cross-platform .NET wrapper for the OpenCV image processing library, enabling .NET compatible languages to use OpenCV functions across multiple operating systems.
MaaNTE: an image and audio recognition based automation tool for simplifying repetitive tasks in the game *異環*
An automation tool for the game *異環* (The Ring) that uses image and audio recognition to simulate user interactions to simplify repetitive gameplay.
stitching: a Python package for fast and robust image stitching and panorama creation
A Python package for fast and robust image stitching that combines multiple overlapping images into a single panorama using OpenCV.
CV-CUDA: a high-throughput GPU-accelerated computer vision library for AI preprocessing pipelines
A GPU-accelerated computer vision library from NVIDIA that provides high-throughput, low-latency image and video processing for AI pipelines.
comic-translate: an AI-powered comic translator that leverages LLMs and inpainting to translate text within comic panels
An AI-powered comic translator that uses LLMs, OCR, and inpainting to translate comics from multiple languages while preserving the original artwork.
habitat-lab: a modular framework for training and evaluating embodied AI agents in indoor environments
A modular high-level library for end-to-end development in embodied AI, used to train and evaluate agents performing tasks in indoor environments.
WebPlotDigitizer: a computer vision assisted tool for extracting numerical data from images of data visualizations
WebPlotDigitizer is a computer vision assisted tool that extracts numerical data from images of data visualizations for researchers and scientists.
ComputeLibrary: a collection of low-level machine learning functions optimized for Arm hardware
A collection of low-level machine learning functions optimized for Arm Cortex-A, Neoverse, and Mali GPU architectures to provide high-performance ML inference.
LichtFeld-Studio: a modular workstation for training, editing, and automating 3D Gaussian Splatting scenes
A modular workstation for 3D Gaussian Splatting that integrates training, real-time visualization, editing, and export in a single native application.
anylabeling: an AI-assisted image annotation tool with auto-labeling via YOLOv8 and Segment Anything
An AI-powered image annotation tool that integrates YOLOv8 and the Segment Anything Model (SAM) family to automate and accelerate data labeling for computer vision.
AliceVision: a photogrammetric computer vision framework for 3D reconstruction and camera tracking
A photogrammetric computer vision framework that provides algorithms for 3D reconstruction and camera tracking from photographs or videos.
habitat-sim: a high-performance physics-enabled 3D simulator for embodied AI research
A high-performance physics-enabled 3D simulator for embodied AI research, providing fast rendering and rigid-body dynamics for training agents in realistic 3D environments.
scenic: a JAX-based research library for prototyping large-scale attention-based computer vision models
A JAX and Flax-based library for computer vision research that provides scalable libraries and project templates for developing attention-based models across images, video, and audio.
torchgeo: a PyTorch domain library for deep learning with multispectral geospatial and remote sensing data
A PyTorch domain library that provides datasets, samplers, and pre-trained models specifically designed for geospatial and remote sensing data.
VLMEvalKit: a unified evaluation toolkit for large vision-language models with support for 70+ benchmarks
An open-source evaluation toolkit for large vision-language models that enables one-command evaluation across 70+ benchmarks and 200+ models.
MaaFramework: a cross-platform image-recognition framework for building low-code black-box automation tools
An image-recognition-based automation framework for black-box testing that enables developers to create visual-driven automation tools with low-code pipelines.
obs-backgroundremoval: a virtual green-screen and low-light enhancement plugin for OBS Studio
An OBS Studio plugin that uses neural networks to provide virtual green-screen background removal and low-light enhancement for portrait video and images.
webots: an open-source robot simulator for modeling and programming mechanical systems
An open-source robot simulator providing a complete development environment to model, program, and simulate robots, vehicles, and mechanical systems.
ceres-solver: a mature C++ library for solving large-scale non-linear least squares and general optimization problems
Ceres Solver is a C++ library for modeling and solving large, complex non-linear least squares and general unconstrained optimization problems.
sparrow: an API-first document intelligence platform for structured data extraction and agentic workflows
An API-first platform for enterprise document intelligence that converts invoices, receipts, and statements into structured JSON using Vision LLMs and agentic workflows.
Chinese-CLIP: a large-scale Chinese vision-language model for cross-modal retrieval and zero-shot image classification
A Chinese-language version of the CLIP model trained on 200 million image-text pairs for cross-modal retrieval and zero-shot image classification.
scikit-image: a comprehensive Python library for scientific image processing and analysis
scikit-image is a Python library for image processing that provides a collection of algorithms for analyzing and manipulating images.
Final2x: a cross-platform image super-resolution tool supporting custom models and Nvidia 50 series GPUs
A cross-platform image super-resolution tool that uses AI models to upscale images and improve their resolution.
gocv: Go language bindings for the OpenCV 4 computer vision library
GoCV provides Go language bindings for OpenCV 4, enabling Go developers to integrate advanced computer vision and hardware-accelerated image processing into their applications.
librealsense: a cross-platform library for streaming depth and color data from RealSense cameras
A cross-platform SDK for RealSense depth cameras that enables depth and color streaming and provides calibration data for computer vision and robotics.
rerun: a multimodal data layer and visual debugger for physical AI and robotics
A multimodal data layer and visual debugger for physical AI, enabling developers to log, query, and visualize multi-rate sensor data in real-time.
pcl: a large-scale open-source library for 2D and 3D image and point cloud processing
A large-scale open-source library for 2D/3D image and point cloud processing, used widely in robotics and 3D perception.
segmentation_models.pytorch: a high-level PyTorch library for image semantic segmentation with over 800 pretrained encoders
A PyTorch library providing a high-level API to easily create and train image semantic segmentation models using a wide variety of pretrained encoders and architectures.
colmap: a general-purpose Structure-from-Motion and Multi-View Stereo pipeline for 3D reconstruction
A general-purpose Structure-from-Motion and Multi-View Stereo pipeline for reconstructing 3D structures from image collections.
Meshroom: a node-based visual programming framework for 3D reconstruction and computer vision pipelines
A node-based visual programming framework for 3D reconstruction and computer vision, enabling users to build complex data processing pipelines using AI and photogrammetry.
open_clip: an open-source framework for training and deploying large-scale contrastive language-image and audio-text models
An open-source implementation of OpenAI's CLIP that enables training and using large-scale contrastive language-image models for zero-shot classification and retrieval.
carla: an open-source urban driving simulator for training and validating autonomous driving systems
An open-source simulator for autonomous driving research used to develop, train, and validate autonomous driving systems in simulated urban environments.
labelme: a graphical image annotation tool with AI-assisted masking and multi-format dataset export
A graphical image annotation tool that allows users to to label images using various shapes and AI-assisted tools to create datasets for computer vision models.
cvat: a professional data annotation platform for building high-quality computer vision datasets
An open-source data annotation platform for building high-quality visual datasets for computer vision, supporting image, video, and 3D annotation.
AirSim: a high-fidelity visual and physical simulator for autonomous vehicles and AI research
An open-source simulator for drones and cars built on Unreal Engine, designed as a platform for AI research in deep learning, computer vision, and reinforcement learning.
MaaAssistantArknights: an image-recognition based automation assistant for Arknights that automates daily tasks and base management
An automation assistant for Arknights that uses image recognition and deep learning to automate daily tasks, base management, and resource farming.
vit-pytorch: a comprehensive collection of Vision Transformer (ViT) and its variants implemented in PyTorch
A comprehensive PyTorch library implementing the original Vision Transformer (ViT) and dozens of its advanced variants for image classification.
label-studio: a multi-modal open-source data labeling tool with ML-assisted pre-labeling and active learning
An open-source data labeling tool that allows users to annotate audio, text, images, and video to prepare or improve training data for machine learning models.
espnet: a comprehensive end-to-end speech processing toolkit for ASR, TTS, and spoken language understanding
An end-to-end speech processing toolkit that provides a unified framework for ASR, TTS, speech translation, and enhancement using PyTorch.
leon: an open-source personal AI assistant with agentic execution and local-first privacy
An open-source personal AI assistant that uses tools, memory, and agentic execution to perform tasks while supporting local AI models for privacy.
HunyuanVideo-1.5: a lightweight 8.3B parameter video generation model for high-quality synthesis on consumer GPUs
A lightweight 8.3B parameter video generation model that enables high-quality text-to-video and image-to-video synthesis on consumer-grade GPUs.
OpenMontage: an agentic video production system that orchestrates research, scripting, and editing into a full production pipeline
An open-source, agentic video production system that orchestrates research, scripting, asset generation, and editing to create professional videos from plain-language prompts.