MITK/MITK

An open-source C++ toolkit for developing interactive medical image processing software by combining ITK and VTK.

NeoGeographyToolkit/StereoPipeline

A NASA-developed suite of open-source automated geodesy and stereogrammetry tools for processing stereo images into 3D terrain models and cartographic products.

savvasdalkitsis/uhuruphotos-android

An open-source Android media gallery that provides a privacy-focused alternative to Google Photos with optional LibrePhotos server integration for AI-powered search and face recognition.

soruly/trace.moe

An anime scene search engine that identifies the specific anime, episode, and timestamp of a screenshot using image vector search.

gowvp/owl

An open-source network video platform that unifies GB28181, ONVIF, and RTSP streams into a single browser-based management system with AI alerting support.

flatironinstitute/CaImAn

A Python toolbox for large-scale calcium and voltage imaging analysis, providing scalable algorithms for motion correction, source extraction, and spike deconvolution.

ros-perception/image_pipeline

A ROS 2 package that processes raw camera images into a usable format for higher-level vision processing in robotics.

nguyenq/tess4j

Tess4J is a Java JNA wrapper for the Tesseract OCR API that allows developers to extract text from various image formats and PDF documents.

aws-samples/amazon-textract-textractor

A Python package that simplifies using Amazon Textract to extract text, tables, forms, and identity data from documents.

yo-WASSUP/Good-GYM

An AI fitness assistant that uses RTMPose for real-time pose detection and automatic exercise repetition counting via a webcam.

TissueImageAnalytics/tiatoolbox

A computational pathology toolbox based on PyTorch that provides an end-to-end API for analyzing tissue images, from data loading to visualization.

LSH9832/edgeyolo

EdgeYOLO is an anchor-free object detector optimized for real-time performance on embedded edge devices like Nvidia Jetson, supporting multiple deployment formats including TensorRT and RKNN.

deltacv/VisionGraph

A visual node editor for creating and prototyping OpenCV algorithms with live pipeline previews.

apple-aiml-research/ml-fastvit

A fast hybrid vision transformer that uses structural reparameterization to achieve low-latency image classification on mobile devices.

shaoshengsong/DeepSORT

A real-time multi-object tracking system written in C++ that implements ByteTrack, OC-SORT, Deep OC-SORT, and DeepSORT using YOLO and ONNX Runtime.

apple-aiml-research/ml-mobileone

MobileOne provides PyTorch implementations of five ultra‑fast CNN backbones (S0‑S4) that achieve 71‑79 % ImageNet Top‑1 accuracy with ≤ 2 ms latency on an iPhone 12 Pro. The repo includes training‑ready and inference‑ready checkpoints, CoreML models for on‑device deployment, and an iOS benchmark app (ModelBench). Re‑parameterization fuses training‑time branches into a single fast network for inference.

apple-aiml-research/ml-matrix3d

Matrix3D is a unified photogrammetry model that performs pose estimation, depth prediction, and novel view synthesis to enable 3D reconstruction from single or few-shot unposed images.

cvignac/DiGress

A discrete denoising diffusion model for generating synthetic graphs and molecular structures using a Graph Transformer architecture.

EasyLive2D/relive2d

A Python wrapper for the Live2D Native SDK that allows direct loading and rendering of Live2D models using OpenGL without needing a web engine.

HZAI-ZJNU/Mamba-YOLO

Mamba YOLO is an object detection baseline that replaces traditional architectures with State Space Models to provide an efficient alternative for real-time detection.

fieldtrip/fieldtrip

A MATLAB toolbox for the advanced analysis of MEG, EEG, and iEEG data, supporting multiple hardware formats and source reconstruction.

sgoldenlab/simba

SimBA is a GUI-based toolkit that lets researchers train supervised machine-learning classifiers to automatically detect and quantify animal behaviors from pose-estimation video data, without needing programming skills. It turns raw tracking data into validated behavioral classifications and rich downstream analyses for behavioral neuroscience.

ecmwf-lab/ai-models

A command-line tool for running AI-based weather forecasting models, providing a unified interface for data retrieval and model execution.

isl-org/Open3D-ML

An extension of Open3D that provides tools, pretrained models, and pipelines for 3D machine learning tasks like semantic segmentation and object detection.

ANTsX/ANTs

A C++ library for high-dimensional medical image registration and segmentation, used to analyze brain structure and function across species and organ systems.

MiniAiLive/Android-FaceRecognition

An Android SDK for facial recognition and 3D passive liveness detection to prevent spoofing in security and fintech applications.

layumi/University1652-Baseline

A multi-view, multi-source benchmark and baseline for drone-based geo-localization, enabling the matching of drone, satellite, and street-view images to locate buildings.

supervisely/supervisely

Supervisely is a web‑based computer‑vision platform that combines data labeling, model training, inference and collaboration tools. It offers a Python SDK and REST API for automation, and an app ecosystem where developers can create head‑less scripts, interactive UI tools, or labeling‑tool extensions—all deployable with a single click.

PozzettiAndrea/ComfyUI-DepthAnythingV3

Custom ComfyUI nodes that integrate Depth Anything V3 for high-quality depth estimation, 3D point cloud reconstruction, and consistent video depth mapping.

sh4den/Montscan

An automated document processor that uses local Vision AI via Ollama to analyze scanned PDFs and automatically rename them with descriptive filenames.

cosanlab/py-feat

A Python toolbox for facial expression research that detects faces and extracts emotional expressions, facial muscle movements, and landmarks from images and videos.

Zyphra/zuna

ZUNA1.1 is an open foundation model for EEG that denoises, reconstructs missing channels, and upsamples sparse electrode layouts using 3D scalp coordinates.

Pixel-Talk/PEAR

PEAR is a real-time framework for expressive 3D human mesh recovery that can predict human mesh parameters at 100 FPS from images or video.

scu-zjz/IMDLBenCo

A comprehensive benchmark and modular codebase for image manipulation detection and localization, providing standardized components and SOTA model implementations.

Mengqi-Lei/count-anything

Count Anything is a research model that counts objects in images based on a natural‑language query. It works across six domains (scenes, satellite, medical, etc.) by combining a sparse region counter and a dense pixel counter, merging their results with a parameter‑free fusion step. The repo provides a pretrained checkpoint, training/evaluation scripts, and instructions for preparing the large CLOC dataset.

Intellindust-AI-Lab/EdgeCrafter

EdgeCrafter provides compact Vision Transformers (ViTs) optimized for edge devices, enabling high-performance object detection, instance segmentation, and pose estimation with low latency.

ultralytics/xview-yolov3

A specialized implementation of YOLOv3 designed to train and run object detection on the xView satellite imagery dataset for remote sensing applications.

mlmed/torchxrayvision

A library for chest X-ray datasets and models, providing pre-trained models and a uniform interface for working with multiple public chest X-ray datasets.

tschnz/Live-Video-Magnification

A real-time video magnification tool that amplifies subtle motion and color changes using Eulerian video magnification techniques.

lightly-ai/lightly-studio

A local-first data curation and annotation tool for computer vision datasets that uses embeddings to help users efficiently manage and label image and video data.

dog-qiuqiu/FastestDet

An ultra-lightweight, anchor-free real-time object detection algorithm optimized for high-speed inference on ARM CPUs and embedded devices.

open-edge-platform/dlstreamer

A media analytics framework based on GStreamer and OpenVINO that enables hardware-accelerated video and audio intelligence pipelines on Intel CPU, GPU, and NPU.

xrdevrob/QuestCameraKit

A collection of Unity samples for Meta Quest 3/3S that enables mixed reality experiences to detect objects, track QR codes, and use multimodal AI to understand the surroundings.

torch-points3d/torch-points3d

A deep learning framework for point cloud analysis that provides a high-level API and a library of state-of-the-art models for 3D classification, segmentation, and detection.

BGU-CS-VIL/WTConv

A library implementing Wavelet Convolutions to provide CNNs with large receptive fields, featuring optimized backends for CUDA, Metal, and Triton.

reall3d-com/Reall3dViewer

A Three.js-based Web renderer for 3D Gaussian Splatting that enables high-performance streaming and adaptive LOD for large-scale 3D scenes.

cre185/InstantSfM

A GPU-native Structure from Motion (SfM) pipeline that accelerates the process of reconstructing 3D structures from 2D images, including integrated support for 3D Gaussian Splatting.

ilastik/ilastik

An interactive machine learning toolkit for segmenting, classifying, tracking, and counting cells and other experimental image data.