MITK/MITK
An open-source C++ toolkit for developing interactive medical image processing software by combining ITK and VTK.
NeoGeographyToolkit/StereoPipeline
A NASA-developed suite of open-source automated geodesy and stereogrammetry tools for processing stereo images into 3D terrain models and cartographic products.
savvasdalkitsis/uhuruphotos-android
An open-source Android media gallery that provides a privacy-focused alternative to Google Photos with optional LibrePhotos server integration for AI-powered search and face recognition.
soruly/trace.moe
An anime scene search engine that identifies the specific anime, episode, and timestamp of a screenshot using image vector search.
gowvp/owl
An open-source network video platform that unifies GB28181, ONVIF, and RTSP streams into a single browser-based management system with AI alerting support.
flatironinstitute/CaImAn
A Python toolbox for large-scale calcium and voltage imaging analysis, providing scalable algorithms for motion correction, source extraction, and spike deconvolution.
ros-perception/image_pipeline
A ROS 2 package that processes raw camera images into a usable format for higher-level vision processing in robotics.
nguyenq/tess4j
Tess4J is a Java JNA wrapper for the Tesseract OCR API that allows developers to extract text from various image formats and PDF documents.
aws-samples/amazon-textract-textractor
A Python package that simplifies using Amazon Textract to extract text, tables, forms, and identity data from documents.
yo-WASSUP/Good-GYM
An AI fitness assistant that uses RTMPose for real-time pose detection and automatic exercise repetition counting via a webcam.
TissueImageAnalytics/tiatoolbox
A computational pathology toolbox based on PyTorch that provides an end-to-end API for analyzing tissue images, from data loading to visualization.
LSH9832/edgeyolo
EdgeYOLO is an anchor-free object detector optimized for real-time performance on embedded edge devices like Nvidia Jetson, supporting multiple deployment formats including TensorRT and RKNN.
deltacv/VisionGraph
A visual node editor for creating and prototyping OpenCV algorithms with live pipeline previews.
apple-aiml-research/ml-fastvit
A fast hybrid vision transformer that uses structural reparameterization to achieve low-latency image classification on mobile devices.
shaoshengsong/DeepSORT
A real-time multi-object tracking system written in C++ that implements ByteTrack, OC-SORT, Deep OC-SORT, and DeepSORT using YOLO and ONNX Runtime.
apple-aiml-research/ml-mobileone
MobileOne provides PyTorch implementations of five ultra‑fast CNN backbones (S0‑S4) that achieve 71‑79 % ImageNet Top‑1 accuracy with ≤ 2 ms latency on an iPhone 12 Pro. The repo includes training‑ready and inference‑ready checkpoints, CoreML models for on‑device deployment, and an iOS benchmark app (ModelBench). Re‑parameterization fuses training‑time branches into a single fast network for inference.
apple-aiml-research/ml-matrix3d
Matrix3D is a unified photogrammetry model that performs pose estimation, depth prediction, and novel view synthesis to enable 3D reconstruction from single or few-shot unposed images.
cvignac/DiGress
A discrete denoising diffusion model for generating synthetic graphs and molecular structures using a Graph Transformer architecture.
EasyLive2D/relive2d
A Python wrapper for the Live2D Native SDK that allows direct loading and rendering of Live2D models using OpenGL without needing a web engine.
HZAI-ZJNU/Mamba-YOLO
Mamba YOLO is an object detection baseline that replaces traditional architectures with State Space Models to provide an efficient alternative for real-time detection.
fieldtrip/fieldtrip
A MATLAB toolbox for the advanced analysis of MEG, EEG, and iEEG data, supporting multiple hardware formats and source reconstruction.
sgoldenlab/simba
SimBA is a GUI-based toolkit that lets researchers train supervised machine-learning classifiers to automatically detect and quantify animal behaviors from pose-estimation video data, without needing programming skills. It turns raw tracking data into validated behavioral classifications and rich downstream analyses for behavioral neuroscience.
ecmwf-lab/ai-models
A command-line tool for running AI-based weather forecasting models, providing a unified interface for data retrieval and model execution.
isl-org/Open3D-ML
An extension of Open3D that provides tools, pretrained models, and pipelines for 3D machine learning tasks like semantic segmentation and object detection.
ANTsX/ANTs
A C++ library for high-dimensional medical image registration and segmentation, used to analyze brain structure and function across species and organ systems.
MiniAiLive/Android-FaceRecognition
An Android SDK for facial recognition and 3D passive liveness detection to prevent spoofing in security and fintech applications.
layumi/University1652-Baseline
A multi-view, multi-source benchmark and baseline for drone-based geo-localization, enabling the matching of drone, satellite, and street-view images to locate buildings.
supervisely/supervisely
Supervisely is a web‑based computer‑vision platform that combines data labeling, model training, inference and collaboration tools. It offers a Python SDK and REST API for automation, and an app ecosystem where developers can create head‑less scripts, interactive UI tools, or labeling‑tool extensions—all deployable with a single click.
PozzettiAndrea/ComfyUI-DepthAnythingV3
Custom ComfyUI nodes that integrate Depth Anything V3 for high-quality depth estimation, 3D point cloud reconstruction, and consistent video depth mapping.
sh4den/Montscan
An automated document processor that uses local Vision AI via Ollama to analyze scanned PDFs and automatically rename them with descriptive filenames.
cosanlab/py-feat
A Python toolbox for facial expression research that detects faces and extracts emotional expressions, facial muscle movements, and landmarks from images and videos.
Zyphra/zuna
ZUNA1.1 is an open foundation model for EEG that denoises, reconstructs missing channels, and upsamples sparse electrode layouts using 3D scalp coordinates.
Pixel-Talk/PEAR
PEAR is a real-time framework for expressive 3D human mesh recovery that can predict human mesh parameters at 100 FPS from images or video.
scu-zjz/IMDLBenCo
A comprehensive benchmark and modular codebase for image manipulation detection and localization, providing standardized components and SOTA model implementations.
Mengqi-Lei/count-anything
Count Anything is a research model that counts objects in images based on a natural‑language query. It works across six domains (scenes, satellite, medical, etc.) by combining a sparse region counter and a dense pixel counter, merging their results with a parameter‑free fusion step. The repo provides a pretrained checkpoint, training/evaluation scripts, and instructions for preparing the large CLOC dataset.
Intellindust-AI-Lab/EdgeCrafter
EdgeCrafter provides compact Vision Transformers (ViTs) optimized for edge devices, enabling high-performance object detection, instance segmentation, and pose estimation with low latency.
ultralytics/xview-yolov3
A specialized implementation of YOLOv3 designed to train and run object detection on the xView satellite imagery dataset for remote sensing applications.
mlmed/torchxrayvision
A library for chest X-ray datasets and models, providing pre-trained models and a uniform interface for working with multiple public chest X-ray datasets.
tschnz/Live-Video-Magnification
A real-time video magnification tool that amplifies subtle motion and color changes using Eulerian video magnification techniques.
lightly-ai/lightly-studio
A local-first data curation and annotation tool for computer vision datasets that uses embeddings to help users efficiently manage and label image and video data.
dog-qiuqiu/FastestDet
An ultra-lightweight, anchor-free real-time object detection algorithm optimized for high-speed inference on ARM CPUs and embedded devices.
open-edge-platform/dlstreamer
A media analytics framework based on GStreamer and OpenVINO that enables hardware-accelerated video and audio intelligence pipelines on Intel CPU, GPU, and NPU.
xrdevrob/QuestCameraKit
A collection of Unity samples for Meta Quest 3/3S that enables mixed reality experiences to detect objects, track QR codes, and use multimodal AI to understand the surroundings.
torch-points3d/torch-points3d
A deep learning framework for point cloud analysis that provides a high-level API and a library of state-of-the-art models for 3D classification, segmentation, and detection.
BGU-CS-VIL/WTConv
A library implementing Wavelet Convolutions to provide CNNs with large receptive fields, featuring optimized backends for CUDA, Metal, and Triton.
reall3d-com/Reall3dViewer
A Three.js-based Web renderer for 3D Gaussian Splatting that enables high-performance streaming and adaptive LOD for large-scale 3D scenes.
cre185/InstantSfM
A GPU-native Structure from Motion (SfM) pipeline that accelerates the process of reconstructing 3D structures from 2D images, including integrated support for 3D Gaussian Splatting.
ilastik/ilastik
An interactive machine learning toolkit for segmenting, classifying, tracking, and counting cells and other experimental image data.