StanfordVL/taskonomy

A framework and dataset for disentangling task transfer learning in computer vision, providing a pretrained task bank and tools to analyze task affinities.

qaz812345/TrackNetV3

A computer vision project that enhances shuttlecock tracking in badminton videos using trajectory prediction with background estimation and an inpainting-based rectification module to handle occlusions.

EthanH3514/AL_Yolo

A real-time visual target detection and tracking system based on YOLOv5, optimized for low-latency screen capture and GPU inference in gaming scenarios.

Breakthrough/DVR-Scan

A command-line and GUI application that automatically detects motion events in video files and saves each event as a separate clip.

soruly/trace.moe-telegram-bot

A Telegram bot that identifies anime titles, episodes, and timestamps from screenshots or videos using the trace.moe service.

pupil-labs/pupil

An open-source eye tracking platform that provides the software infrastructure for capturing, playing back, and analyzing eye-tracking data.

ermig1979/Simd

A high-performance C/C++ image processing and machine learning library that uses SIMD CPU extensions to accelerate algorithms across x86, ARM, and Hexagon architectures.

iago-suarez/ELSED

ELSED is a high-speed line segment detector designed for resource-limited devices like drones and smartphones.

Faceplugin-ltd/FaceRecognition-Android

An on-device face recognition and liveness detection SDK for Android that enables private, offline biometric identity verification and attribute analysis.

stereolabs/zed-unity

A Unity plugin that integrates ZED cameras' depth sensing, spatial mapping, and object detection capabilities into the Unity engine for AR/MR development.

deepfates/memery

A natural language image search tool that allows users to find specific images in local folders using text or image queries powered by OpenAI's CLIP.

nachifur/MulimgViewer

A multi-image viewer designed for efficient comparison, filtering, and stitching of large image datasets, featuring parallel zooming and automated figure generation.

VicenteVivan/geo-clip

GeoCLIP is a CLIP-inspired model that aligns images with geographical locations to enable high-accuracy worldwide image geo-localization and GPS-to-vector embeddings.

lartpang/PySODEvalToolkit

A Python toolbox for evaluating grayscale and binary image segmentation models, providing comprehensive metrics and automated curve plotting.

OSU-NLP-Group/UGround

UGround is an open‑source visual‑language model (2 B/7 B/72 B) fine‑tuned to locate UI elements in screenshots from a textual description. The repo ships pretrained weights, a >1 M‑image GUI grounding dataset, evaluation scripts for four benchmarks, and a Hugging Face demo. Inference runs via vLLM with a simple prompt that returns (x, y) coordinates. The 72 B model achieves state‑of‑the‑art scores on the ScreenSpot benchmark and is described in an ICLR 2025 oral paper.

juliansteenbakker/mobile_scanner

A fast, lightweight Flutter plugin for scanning barcodes and QR codes with the device camera, supporting Android, iOS, macOS, and web with real-time detection and customizable camera behavior.

bopen/sarsen

A Python library for cloud-native Synthetic Aperture Radar (SAR) processing that provides geometric and radiometric terrain correction for Sentinel-1 satellite data.

leomariga/pyRANSAC-3D

A Python implementation of the RANSAC method for fitting primitive geometric shapes to 3D point clouds, useful for 3D reconstruction and SLAM.

ouster-lidar/ouster-sdk

A cross-platform C++/Python development toolkit for connecting to, configuring, and visualizing data from Ouster Lidar sensors.

opencv/opencv_extra

A repository of extra data and resources that supplement the core OpenCV computer vision library.

Thinklab-SJTU/ThinkMatch

A research framework for developing and benchmarking deep graph matching algorithms to find node-to-node correspondences between graphs.

liuliu/ccv

A minimalistic and portable C-based computer vision library providing state-of-the-art algorithms for face, object, and text detection.

PaddlePaddle/PaddleClas

A comprehensive image recognition and classification toolkit providing ultra-lightweight models and industrial-grade backbones for high-performance vision applications.

SegmentationBLWX/sssegmentation

A PyTorch-based supervised semantic segmentation toolbox providing a unified framework and extensive model zoo for training and testing segmentation algorithms.

LAION-AI/CLIP_benchmark

A standardized evaluation framework for CLIP-like models to measure performance on zero-shot classification, retrieval, captioning, and linear probing across diverse datasets.

facebookresearch/mvdust3r

A single-stage 3D scene reconstruction tool that creates 3D point clouds and camera poses from sparse RGB views in about 2 seconds without requiring pre-known camera poses.

zsylvester/segmenteverygrain

A Python package for detecting and segmenting grains in images using a hybrid U-Net and SAM 2.1 approach, designed for geomorphology and sedimentary geology research.

phamquiluan/ResidualMaskingNetwork

A facial expression recognition system using a Residual Masking Network to detect human emotions from images and live video feeds.

NVlabs/nvdiffrecmc

A PyTorch implementation for jointly reconstructing 3D topology, materials, and lighting from multi-view images using Monte Carlo rendering and denoising.

adalca/neurite

A neural networks toolbox for medical image analysis that provides specialized tensor operations, prebuilt architectures, and plotting tools for spatial data.

scribeocr/scribe.js

A JavaScript library for performing OCR and text extraction from images and PDFs, capable of creating searchable PDFs with invisible text layers.

hysts/anime-face-detector

A PyTorch-based tool for detecting anime faces and estimating 28 facial landmark points using pretrained Faster R-CNN, YOLOv3, and HRNetV2 models.

mkalten/reacTIVision

A cross-platform computer vision framework for tracking fiducial markers and multi-touch fingers to enable the creation of tangible user interfaces.

roboflow/roboflow-python

The official Python package for Roboflow, enabling developers to interact with models, datasets, and projects to build and deploy computer vision models programmatically.

lessthanoptimal/BoofCV

A real-time computer vision library written in Java that provides tools for low-level image processing, camera calibration, and feature tracking.

githubharald/SimpleHTR

A TensorFlow-based handwritten text recognition system that converts images of single words or text lines into digital text using CNN and LSTM layers.

scribeocr/scribeocr

A browser-based web application for recognizing text from images, proofreading OCR data with high precision, and creating fully digitized, ebook-style documents.

visionworkbench/visionworkbench

A NASA-developed C++ library for general-purpose image processing and computer vision, providing tools for camera models, cartographic projections, and mosaic compositing.

openclimatefix/graph_weather

A PyTorch library implementing various graph neural networks for weather forecasting and data assimilation, including models like GenCast and Aurora.

norlab-ulaval/libpointmatcher

A modular C++ library with Python bindings that implements the Iterative Closest Point (ICP) algorithm for aligning 3D point clouds in robotics and computer vision.

owensgroup/RXMesh

A GPU-accelerated library for processing triangle meshes, featuring integrated matrix infrastructure and automatic differentiation for geometry processing tasks.

PoseLib/PoseLib

A collection of minimal solvers and robust estimators for camera pose estimation, supporting various correspondences and camera models with minimal dependencies.

PozzettiAndrea/ComfyUI-SAM3

A ComfyUI integration for Meta's SAM3 that enables open-vocabulary image and video segmentation using natural language text prompts.

UCSC-VLAA/OpenVision

A family of open-source, cost-effective vision encoders for multimodal learning that optimizes training efficiency and performance across understanding and generation tasks.

Seeed-Studio/OSHW-reCamera-Series

A hardware documentation and design-resource hub for the reCamera series of AI-powered vision modules with integrated NPUs for edge AI.

CellProfiler/CellProfiler

CellProfiler is an open-source software that enables biologists to automatically and quantitatively measure phenotypes from thousands of images without requiring programming or computer vision expertise.

TechyNilesh/DeepImageSearch

A Python library for building AI-powered image search systems supporting text, image, and hybrid search using multimodal embeddings and vector indexing.

MITK/MITK

An open-source C++ toolkit for developing interactive medical image processing software by combining ITK and VTK.