StanfordVL/taskonomy
A framework and dataset for disentangling task transfer learning in computer vision, providing a pretrained task bank and tools to analyze task affinities.
qaz812345/TrackNetV3
A computer vision project that enhances shuttlecock tracking in badminton videos using trajectory prediction with background estimation and an inpainting-based rectification module to handle occlusions.
EthanH3514/AL_Yolo
A real-time visual target detection and tracking system based on YOLOv5, optimized for low-latency screen capture and GPU inference in gaming scenarios.
Breakthrough/DVR-Scan
A command-line and GUI application that automatically detects motion events in video files and saves each event as a separate clip.
soruly/trace.moe-telegram-bot
A Telegram bot that identifies anime titles, episodes, and timestamps from screenshots or videos using the trace.moe service.
pupil-labs/pupil
An open-source eye tracking platform that provides the software infrastructure for capturing, playing back, and analyzing eye-tracking data.
ermig1979/Simd
A high-performance C/C++ image processing and machine learning library that uses SIMD CPU extensions to accelerate algorithms across x86, ARM, and Hexagon architectures.
iago-suarez/ELSED
ELSED is a high-speed line segment detector designed for resource-limited devices like drones and smartphones.
Faceplugin-ltd/FaceRecognition-Android
An on-device face recognition and liveness detection SDK for Android that enables private, offline biometric identity verification and attribute analysis.
stereolabs/zed-unity
A Unity plugin that integrates ZED cameras' depth sensing, spatial mapping, and object detection capabilities into the Unity engine for AR/MR development.
deepfates/memery
A natural language image search tool that allows users to find specific images in local folders using text or image queries powered by OpenAI's CLIP.
nachifur/MulimgViewer
A multi-image viewer designed for efficient comparison, filtering, and stitching of large image datasets, featuring parallel zooming and automated figure generation.
VicenteVivan/geo-clip
GeoCLIP is a CLIP-inspired model that aligns images with geographical locations to enable high-accuracy worldwide image geo-localization and GPS-to-vector embeddings.
lartpang/PySODEvalToolkit
A Python toolbox for evaluating grayscale and binary image segmentation models, providing comprehensive metrics and automated curve plotting.
OSU-NLP-Group/UGround
UGround is an open‑source visual‑language model (2 B/7 B/72 B) fine‑tuned to locate UI elements in screenshots from a textual description. The repo ships pretrained weights, a >1 M‑image GUI grounding dataset, evaluation scripts for four benchmarks, and a Hugging Face demo. Inference runs via vLLM with a simple prompt that returns (x, y) coordinates. The 72 B model achieves state‑of‑the‑art scores on the ScreenSpot benchmark and is described in an ICLR 2025 oral paper.
juliansteenbakker/mobile_scanner
A fast, lightweight Flutter plugin for scanning barcodes and QR codes with the device camera, supporting Android, iOS, macOS, and web with real-time detection and customizable camera behavior.
bopen/sarsen
A Python library for cloud-native Synthetic Aperture Radar (SAR) processing that provides geometric and radiometric terrain correction for Sentinel-1 satellite data.
leomariga/pyRANSAC-3D
A Python implementation of the RANSAC method for fitting primitive geometric shapes to 3D point clouds, useful for 3D reconstruction and SLAM.
ouster-lidar/ouster-sdk
A cross-platform C++/Python development toolkit for connecting to, configuring, and visualizing data from Ouster Lidar sensors.
opencv/opencv_extra
A repository of extra data and resources that supplement the core OpenCV computer vision library.
Thinklab-SJTU/ThinkMatch
A research framework for developing and benchmarking deep graph matching algorithms to find node-to-node correspondences between graphs.
liuliu/ccv
A minimalistic and portable C-based computer vision library providing state-of-the-art algorithms for face, object, and text detection.
PaddlePaddle/PaddleClas
A comprehensive image recognition and classification toolkit providing ultra-lightweight models and industrial-grade backbones for high-performance vision applications.
SegmentationBLWX/sssegmentation
A PyTorch-based supervised semantic segmentation toolbox providing a unified framework and extensive model zoo for training and testing segmentation algorithms.
LAION-AI/CLIP_benchmark
A standardized evaluation framework for CLIP-like models to measure performance on zero-shot classification, retrieval, captioning, and linear probing across diverse datasets.
facebookresearch/mvdust3r
A single-stage 3D scene reconstruction tool that creates 3D point clouds and camera poses from sparse RGB views in about 2 seconds without requiring pre-known camera poses.
zsylvester/segmenteverygrain
A Python package for detecting and segmenting grains in images using a hybrid U-Net and SAM 2.1 approach, designed for geomorphology and sedimentary geology research.
phamquiluan/ResidualMaskingNetwork
A facial expression recognition system using a Residual Masking Network to detect human emotions from images and live video feeds.
NVlabs/nvdiffrecmc
A PyTorch implementation for jointly reconstructing 3D topology, materials, and lighting from multi-view images using Monte Carlo rendering and denoising.
adalca/neurite
A neural networks toolbox for medical image analysis that provides specialized tensor operations, prebuilt architectures, and plotting tools for spatial data.
scribeocr/scribe.js
A JavaScript library for performing OCR and text extraction from images and PDFs, capable of creating searchable PDFs with invisible text layers.
hysts/anime-face-detector
A PyTorch-based tool for detecting anime faces and estimating 28 facial landmark points using pretrained Faster R-CNN, YOLOv3, and HRNetV2 models.
mkalten/reacTIVision
A cross-platform computer vision framework for tracking fiducial markers and multi-touch fingers to enable the creation of tangible user interfaces.
roboflow/roboflow-python
The official Python package for Roboflow, enabling developers to interact with models, datasets, and projects to build and deploy computer vision models programmatically.
lessthanoptimal/BoofCV
A real-time computer vision library written in Java that provides tools for low-level image processing, camera calibration, and feature tracking.
githubharald/SimpleHTR
A TensorFlow-based handwritten text recognition system that converts images of single words or text lines into digital text using CNN and LSTM layers.
scribeocr/scribeocr
A browser-based web application for recognizing text from images, proofreading OCR data with high precision, and creating fully digitized, ebook-style documents.
visionworkbench/visionworkbench
A NASA-developed C++ library for general-purpose image processing and computer vision, providing tools for camera models, cartographic projections, and mosaic compositing.
openclimatefix/graph_weather
A PyTorch library implementing various graph neural networks for weather forecasting and data assimilation, including models like GenCast and Aurora.
norlab-ulaval/libpointmatcher
A modular C++ library with Python bindings that implements the Iterative Closest Point (ICP) algorithm for aligning 3D point clouds in robotics and computer vision.
owensgroup/RXMesh
A GPU-accelerated library for processing triangle meshes, featuring integrated matrix infrastructure and automatic differentiation for geometry processing tasks.
PoseLib/PoseLib
A collection of minimal solvers and robust estimators for camera pose estimation, supporting various correspondences and camera models with minimal dependencies.
PozzettiAndrea/ComfyUI-SAM3
A ComfyUI integration for Meta's SAM3 that enables open-vocabulary image and video segmentation using natural language text prompts.
UCSC-VLAA/OpenVision
A family of open-source, cost-effective vision encoders for multimodal learning that optimizes training efficiency and performance across understanding and generation tasks.
Seeed-Studio/OSHW-reCamera-Series
A hardware documentation and design-resource hub for the reCamera series of AI-powered vision modules with integrated NPUs for edge AI.
CellProfiler/CellProfiler
CellProfiler is an open-source software that enables biologists to automatically and quantitatively measure phenotypes from thousands of images without requiring programming or computer vision expertise.
TechyNilesh/DeepImageSearch
A Python library for building AI-powered image search systems supporting text, image, and hybrid search using multimodal embeddings and vector indexing.
MITK/MITK
An open-source C++ toolkit for developing interactive medical image processing software by combining ITK and VTK.