651

comfyui_LLM_party: a comprehensive suite of ComfyUI nodes for building complex LLM and agentic workflows

A comprehensive set of ComfyUI nodes for building LLM workflows, enabling the integration of AI agents, RAG, and VLMs into visual image generation pipelines.

652

OpenManus-RL: an RL-based agent tuning framework for enhancing LLM reasoning and decision-making

An open-source framework for optimizing LLM agents using reinforcement learning to enhance their reasoning, tool use, and decision-making capabilities.

653

llm-beginner: a progressive hands-on curriculum for mastering LLMs and AI agents through from-scratch implementations

A comprehensive, hands-on tutorial series for beginners to learn LLMs and AI agents through six progressive implementation tasks, from building a mini-GPT to creating a coding agent.

654

SkillOpt: an executive strategy for self-evolving agent skills using a deep-learning-style optimization loop

SkillOpt is a framework that treats agent skill documents as trainable states, using a deep-learning-style optimization loop to evolve and improve agent performance without modifying model weights.

655

mrpt: a comprehensive C++ framework for mobile robotics featuring SLAM, navigation, and 3D visualization

A robust, modular C++ framework for mobile robotics providing essential tools for SLAM, navigation, and sensor integration.

656

unrealcv: a bridge between Unreal Engine and AI frameworks for creating synthetic computer vision environments

UnrealCV is a plugin for Unreal Engine that allows computer vision researchers to programmatically interact with virtual worlds to generate synthetic data and PyTorch/Tensorflow integration.

657

paperlib: an open-source academic paper manager with AI-powered summarization and semantic library search

An open-source academic paper management tool that simplifies metadata scraping and organization, featuring LLM-powered summarization and semantic search.

658

bgslibrary: a comprehensive C++ framework for foreground-background separation in video streams

A comprehensive C++ framework for background subtraction in computer vision that provides over 40 algorithms to detect moving objects in video streams.

659

emgucv: a cross-platform .NET wrapper for the OpenCV image-processing library

A cross-platform .NET wrapper for the OpenCV image processing library, enabling .NET compatible languages to use OpenCV functions across multiple operating systems.

660

MaaNTE: an image and audio recognition based automation tool for simplifying repetitive tasks in the game *異環*

An automation tool for the game *異環* (The Ring) that uses image and audio recognition to simulate user interactions to simplify repetitive gameplay.

661

stitching: a Python package for fast and robust image stitching and panorama creation

A Python package for fast and robust image stitching that combines multiple overlapping images into a single panorama using OpenCV.

662

CV-CUDA: a high-throughput GPU-accelerated computer vision library for AI preprocessing pipelines

A GPU-accelerated computer vision library from NVIDIA that provides high-throughput, low-latency image and video processing for AI pipelines.

663

comic-translate: an AI-powered comic translator that leverages LLMs and inpainting to translate text within comic panels

An AI-powered comic translator that uses LLMs, OCR, and inpainting to translate comics from multiple languages while preserving the original artwork.

664

habitat-lab: a modular framework for training and evaluating embodied AI agents in indoor environments

A modular high-level library for end-to-end development in embodied AI, used to train and evaluate agents performing tasks in indoor environments.

665

WebPlotDigitizer: a computer vision assisted tool for extracting numerical data from images of data visualizations

WebPlotDigitizer is a computer vision assisted tool that extracts numerical data from images of data visualizations for researchers and scientists.

666

ComputeLibrary: a collection of low-level machine learning functions optimized for Arm hardware

A collection of low-level machine learning functions optimized for Arm Cortex-A, Neoverse, and Mali GPU architectures to provide high-performance ML inference.

667

LichtFeld-Studio: a modular workstation for training, editing, and automating 3D Gaussian Splatting scenes

A modular workstation for 3D Gaussian Splatting that integrates training, real-time visualization, editing, and export in a single native application.

668

anylabeling: an AI-assisted image annotation tool with auto-labeling via YOLOv8 and Segment Anything

An AI-powered image annotation tool that integrates YOLOv8 and the Segment Anything Model (SAM) family to automate and accelerate data labeling for computer vision.

669

AliceVision: a photogrammetric computer vision framework for 3D reconstruction and camera tracking

A photogrammetric computer vision framework that provides algorithms for 3D reconstruction and camera tracking from photographs or videos.

670

habitat-sim: a high-performance physics-enabled 3D simulator for embodied AI research

A high-performance physics-enabled 3D simulator for embodied AI research, providing fast rendering and rigid-body dynamics for training agents in realistic 3D environments.

671

scenic: a JAX-based research library for prototyping large-scale attention-based computer vision models

A JAX and Flax-based library for computer vision research that provides scalable libraries and project templates for developing attention-based models across images, video, and audio.

672

torchgeo: a PyTorch domain library for deep learning with multispectral geospatial and remote sensing data

A PyTorch domain library that provides datasets, samplers, and pre-trained models specifically designed for geospatial and remote sensing data.

673

VLMEvalKit: a unified evaluation toolkit for large vision-language models with support for 70+ benchmarks

An open-source evaluation toolkit for large vision-language models that enables one-command evaluation across 70+ benchmarks and 200+ models.

674

MaaFramework: a cross-platform image-recognition framework for building low-code black-box automation tools

An image-recognition-based automation framework for black-box testing that enables developers to create visual-driven automation tools with low-code pipelines.

675

obs-backgroundremoval: a virtual green-screen and low-light enhancement plugin for OBS Studio

An OBS Studio plugin that uses neural networks to provide virtual green-screen background removal and low-light enhancement for portrait video and images.

676

webots: an open-source robot simulator for modeling and programming mechanical systems

An open-source robot simulator providing a complete development environment to model, program, and simulate robots, vehicles, and mechanical systems.

677

ceres-solver: a mature C++ library for solving large-scale non-linear least squares and general optimization problems

Ceres Solver is a C++ library for modeling and solving large, complex non-linear least squares and general unconstrained optimization problems.

678

sparrow: an API-first document intelligence platform for structured data extraction and agentic workflows

An API-first platform for enterprise document intelligence that converts invoices, receipts, and statements into structured JSON using Vision LLMs and agentic workflows.

679

Chinese-CLIP: a large-scale Chinese vision-language model for cross-modal retrieval and zero-shot image classification

A Chinese-language version of the CLIP model trained on 200 million image-text pairs for cross-modal retrieval and zero-shot image classification.

680

scikit-image: a comprehensive Python library for scientific image processing and analysis

scikit-image is a Python library for image processing that provides a collection of algorithms for analyzing and manipulating images.

681

Final2x: a cross-platform image super-resolution tool supporting custom models and Nvidia 50 series GPUs

A cross-platform image super-resolution tool that uses AI models to upscale images and improve their resolution.

682

gocv: Go language bindings for the OpenCV 4 computer vision library

GoCV provides Go language bindings for OpenCV 4, enabling Go developers to integrate advanced computer vision and hardware-accelerated image processing into their applications.

683

librealsense: a cross-platform library for streaming depth and color data from RealSense cameras

A cross-platform SDK for RealSense depth cameras that enables depth and color streaming and provides calibration data for computer vision and robotics.

684

rerun: a multimodal data layer and visual debugger for physical AI and robotics

A multimodal data layer and visual debugger for physical AI, enabling developers to log, query, and visualize multi-rate sensor data in real-time.

685

pcl: a large-scale open-source library for 2D and 3D image and point cloud processing

A large-scale open-source library for 2D/3D image and point cloud processing, used widely in robotics and 3D perception.

686

segmentation_models.pytorch: a high-level PyTorch library for image semantic segmentation with over 800 pretrained encoders

A PyTorch library providing a high-level API to easily create and train image semantic segmentation models using a wide variety of pretrained encoders and architectures.

687

colmap: a general-purpose Structure-from-Motion and Multi-View Stereo pipeline for 3D reconstruction

A general-purpose Structure-from-Motion and Multi-View Stereo pipeline for reconstructing 3D structures from image collections.

688

Meshroom: a node-based visual programming framework for 3D reconstruction and computer vision pipelines

A node-based visual programming framework for 3D reconstruction and computer vision, enabling users to build complex data processing pipelines using AI and photogrammetry.

689

open_clip: an open-source framework for training and deploying large-scale contrastive language-image and audio-text models

An open-source implementation of OpenAI's CLIP that enables training and using large-scale contrastive language-image models for zero-shot classification and retrieval.

690

carla: an open-source urban driving simulator for training and validating autonomous driving systems

An open-source simulator for autonomous driving research used to develop, train, and validate autonomous driving systems in simulated urban environments.

691

labelme: a graphical image annotation tool with AI-assisted masking and multi-format dataset export

A graphical image annotation tool that allows users to to label images using various shapes and AI-assisted tools to create datasets for computer vision models.

692

cvat: a professional data annotation platform for building high-quality computer vision datasets

An open-source data annotation platform for building high-quality visual datasets for computer vision, supporting image, video, and 3D annotation.

693

AirSim: a high-fidelity visual and physical simulator for autonomous vehicles and AI research

An open-source simulator for drones and cars built on Unreal Engine, designed as a platform for AI research in deep learning, computer vision, and reinforcement learning.

694

MaaAssistantArknights: an image-recognition based automation assistant for Arknights that automates daily tasks and base management

An automation assistant for Arknights that uses image recognition and deep learning to automate daily tasks, base management, and resource farming.

695

vit-pytorch: a comprehensive collection of Vision Transformer (ViT) and its variants implemented in PyTorch

A comprehensive PyTorch library implementing the original Vision Transformer (ViT) and dozens of its advanced variants for image classification.

696

label-studio: a multi-modal open-source data labeling tool with ML-assisted pre-labeling and active learning

An open-source data labeling tool that allows users to annotate audio, text, images, and video to prepare or improve training data for machine learning models.

697

espnet: a comprehensive end-to-end speech processing toolkit for ASR, TTS, and spoken language understanding

An end-to-end speech processing toolkit that provides a unified framework for ASR, TTS, speech translation, and enhancement using PyTorch.

698

leon: an open-source personal AI assistant with agentic execution and local-first privacy

An open-source personal AI assistant that uses tools, memory, and agentic execution to perform tasks while supporting local AI models for privacy.

699

HunyuanVideo-1.5: a lightweight 8.3B parameter video generation model for high-quality synthesis on consumer GPUs

A lightweight 8.3B parameter video generation model that enables high-quality text-to-video and image-to-video synthesis on consumer-grade GPUs.

700

OpenMontage: an agentic video production system that orchestrates research, scripting, and editing into a full production pipeline

An open-source, agentic video production system that orchestrates research, scripting, asset generation, and editing to create professional videos from plain-language prompts.