qubvel-org/segmentation_models.pytorch

A PyTorch library providing a high-level API for image semantic segmentation, featuring 12 architectures and over 800 pretrained encoders.

10x-Engineers/Infinite-ISP

A full-stack ISP development platform that provides a complete pipeline from Python-based algorithm design to RTL and FPGA implementation for converting RAW sensor images to RGB.

scier/MetalSplatter

A Swift/Metal library for rendering 3D Gaussian Splats on Apple platforms, enabling real-time visualization of 3D scenes on iOS, macOS, and visionOS.

google-ar/arcore-android-sdk

The ARCore SDK for Android provides APIs for motion tracking, environmental understanding, and light estimation to build augmented reality experiences.

Smorodov/Multitarget-tracker

A comprehensive multi-target tracking framework that combines SOTA deep learning detectors, matching algorithms, and Kalman filters to follow multiple objects in video.

VisionDepth/VisionDepth3D

VisionDepth 3D is a Windows desktop app that converts 2‑D images or videos into stereoscopic 3‑D using AI depth models, depth‑map blending, RIFE frame‑interpolation, and AI super‑resolution. It offers a free tier (limited length, 1080p, watermarked) and a one‑time $59.99 Pro tier (no limits, 4K+, batch processing, advanced controls). Installation is via an Itch.io “Setup Hub” that provides CUDA (NVIDIA) or DirectML (AMD/Intel) builds. The software is proprietary, but the AI models it downloads are open‑source under their own licenses.

SimpleITK/SimpleITK

An image analysis toolkit that provides a simplified interface to the Insight Segmentation and Registration Toolkit (ITK) for image segmentation and registration.

rdk/p2rank

A machine learning-based command-line tool for fast and accurate prediction of ligand-binding sites from protein structures.

illuin-tech/colpali

ColPali is a research library for visual document retrieval. It encodes document images with vision‑language models (e.g., PaliGemma, Qwen‑VL) into multi‑vector embeddings and scores them against query embeddings using a ColBERT‑style late‑interaction method. The original `colpali-engine` package (now deprecated) provides model classes, processors, optional fused kernels, fast‑Plaid indexing, token‑pooling, and interpretability tools. New projects should use Sentence‑Transformers v6’s `MultiVectorEncoder`, which integrates ColPali‑style models.

zju3dv/MatchAnything

MatchAnything is a research project that trains a single neural network to find correspondences between images from any visual modality (e.g., RGB, infrared, sketches). Pre‑trained weights are hosted on Hugging Face, and an online demo lets users try cross‑modality matching instantly. The approach aims to replace many specialized matching pipelines with one universal model, useful for robotics, remote sensing, medical imaging, and AR.

weecology/DeepForest

A Python package for training and predicting ecological objects, such as tree crowns and birds, in airborne RGB imagery using deep learning object detection.

nianticlabs/map-free-reloc

A reference implementation for map-free visual relocalization that enables metric pose estimation relative to a single image, removing the need for large 3D maps.

alephpi/Texo

A minimalist, open-source LaTeX OCR model with only 20M parameters that converts mathematical formula images into LaTeX code.

xg-chu/GAGAvatar

GAGAvatar is a framework for one-shot 3D head avatar reconstruction and real-time reenactment from a single image using 3D Gaussian Splatting.

mgonzs13/yolo_ros

A ROS 2 wrapper for Ultralytics YOLO models that enables object detection, tracking, instance segmentation, and 3D object localization for robots.

microsoft/farmvibes-ai

A multi-modal geospatial machine learning framework for agriculture and sustainability that fuses satellite imagery, weather data, and elevation maps to generate robust agricultural insights.

z1069614715/objectdetection_script

A comprehensive toolkit for improving and compressing object detection models, featuring modified YOLO and RT-DETR architectures, pruning, distillation, and multimodal support.

ucam-eo/geotessera

A Python library for reading, exporting, and visualizing geospatial embeddings from the Tessera Earth foundation model, supporting both cloud-native streaming and local tile downloads.

opencap-org/opencap-core

A pipeline that estimates 3D human movement kinematics and marker positions from multiple smartphone videos, outputting data in OpenSim format.

openmv/openmv

An open-source machine vision platform that enables beginners to implement AI and image processing on microcontrollers using Python3.

DanBloomberg/leptonica

A comprehensive ANSI C image processing and analysis library used by projects like Tesseract and OpenCV to handle document and natural image manipulation.

nipreps/mriqc

An open-source tool that extracts no-reference image quality metrics from structural, functional, and diffusion MRI data for automated quality assessment.

luxonis/oak-examples

A collection of demonstrations and tutorials for Luxonis OAK devices, showcasing AI vision, depth measurement, and neural network implementation using DepthAI.

ANTsX/ANTsPy

A Python wrapper for the ANTs C++ framework that provides high-performance tools for biomedical image registration, segmentation, and processing.

sstary/SSRS

A PyTorch collection of deep learning models for remote sensing semantic segmentation, featuring single-modal, multimodal fusion, and unsupervised domain adaptation techniques.

jianboqi/CSF

A LiDAR filtering method based on cloth simulation that separates ground points from non-ground points in airborne point clouds.

sentinel-hub/eo-learn

A Python library that bridges Earth observation data with the machine learning ecosystem, simplifying the extraction of information from spatio-temporal satellite imagery.

zibo-chen/ocr-rs

A lightweight Rust OCR library based on PaddleOCR models and the MNN inference runtime, providing text detection and recognition across 50+ languages.

apple-aiml-research/ml-cvnets

A computer vision toolkit for training and evaluating standard and novel mobile and non-mobile models for tasks like classification, detection, and segmentation.

vietanhdev/samexporter

A tool to export and run SAM, SAM 2, SAM 3, MobileSAM, and EfficientSAM models in ONNX Runtime for portable and efficient deployment.

NativeSensors/EyeGestures

An open-source eye-tracking library that enables eye-driven interfaces using standard webcams and phone cameras instead of expensive specialized hardware.

Jianghanxiao/PhysTwin

PhysTwin is a research codebase that reconstructs physics‑informed digital twins of deformable objects from RGB‑D video, provides interactive simulation, analysis tools, and example pipelines for robot planning and batched simulation.

facebookresearch/pytorch3d

A PyTorch-based library for 3D Computer Vision research providing differentiable rendering and efficient tools for manipulating 3D meshes and point clouds.

Yuliang-Liu/MultimodalOCR

A collection of benchmarks (MDPBench, OCRBench v2, and OCRBench) designed to evaluate the OCR and document parsing capabilities of Large Multimodal Models across multiple languages and real-world scenarios.

yehonathanlitman/Lift4D

Lift4D reconstructs temporally consistent 4D assets from single-view in-the-wild videos by fusing per-frame 3D estimations into a deformable model.

NVIDIA-RTX/RTXNS

A developer framework for integrating machine learning into graphics applications, enabling the use of neural networks to represent shaders and textures.

ultralytics/yolo-flutter-app

An official Ultralytics plugin for Flutter that enables real-time YOLO model inference for object detection, segmentation, and pose estimation on iOS and Android.

agentmorris/MegaDetector

An AI model that identifies animals, people, and vehicles in camera trap images to help conservation biologists efficiently remove blank images.

alexandremendoncaalvaro/CorridorKey-Runtime

A native AI keying runtime and OFX plugin for DaVinci Resolve and Foundry Nuke that provides automated, ML-accelerated background removal for video editors.

torchgeo/terratorch

A PyTorch domain library for fine-tuning and using pretrained Geospatial Foundation Models for tasks like image segmentation and classification.

kornia/kornia-rs

A low-level computer vision library written in Rust that provides fast, zero-copy image processing on CPU and NVIDIA GPUs for real-time ML pipelines.

raphaelvallat/yasa

A Python-based sleep analysis toolbox for automatic sleep staging and event detection in polysomnography and EEG data.

MrGiovanni/UNetPlusPlus

A nested U-Net architecture for medical image segmentation that improves upon the original U-Net by redesigning skip connections and allowing for flexible network depth.

allenai/olmoearth_pretrain

A family of multi-modal, spatio-temporal foundation models for Earth Observation, providing a scalable framework for planetary intelligence using satellite data.

ultralytics/yolo-ios-app

A native iOS app and Swift SDK that enables real-time, on-device YOLO inference for object detection, segmentation, and pose estimation using Core ML and Apple Neural Engine.

next-state/open-dreamer

Open Dreamer is a JAX/Flax implementation of the Dreamer 4 world‑model pipeline, providing training code for a video tokenizer and an action‑conditioned latent dynamics model on Minecraft gameplay data, plus tools for rollout generation and FVD evaluation. It includes a live in‑browser demo and a separate inference repo for running trained checkpoints.

noahcao/OC_SORT

A pure motion-model-based multi-object tracker that improves robustness in crowded scenes and non-linear motion by rethinking the SORT algorithm.

zju3dv/manhattan_sdf

Manhattan‑SDF is a CVPR‑2022 research code that learns a neural signed‑distance‑function for indoor 3‑D reconstruction, adding a Manhattan‑world (axis‑aligned) regulariser to improve geometry quality. It provides conda‑based installation, scripts for training on ScanNet (or custom data), mesh extraction, and evaluation against many baselines.