Tobias-Fischer/rt_gene
A real-time eye gaze and blink estimation system using PyTorch and ROS 2 for tracking gaze direction and blink probability in natural environments.
TIGER-AI-Lab/Pixel-Reasoner
Pixel Reasoner is a research framework that adds visual‑manipulation operations (zoom‑in, select‑frame, etc.) to a Vision‑Language Model, training it via instruction‑tuning and curiosity‑driven reinforcement learning. The repo supplies installation guides, scripts for both SFT and RL stages, pretrained 7 B checkpoints, datasets, and evaluation pipelines for image and video benchmarks.
jabberjabberjabber/ImageIndexer
A local AI tool that automatically generates captions and keywords for images and writes them directly into image metadata for easy indexing and searching.
yuanze-lin/Olympus
A universal task router for computer vision that converts a single natural-language instruction into a sequence of specialist model calls to produce images, videos, and 3D models.
Jumpat/SegAnyGAussians
SAGA is a framework for segmenting 3D Gaussian Splatting scenes, allowing users to isolate 3D objects using interactive point prompts or open-vocabulary text queries.
jasmcaus/caer
Caer is a Python‑only, GPU‑accelerated computer‑vision library that offers fast image/video utilities (resize, colour conversion, augmentations, etc.) with a clean, type‑checked API. It can replace OpenCV for research prototypes, is installable via `pip`, and is MIT‑licensed.
Pbatch/ClashRoyaleBuildABot
An automation bot for Clash Royale that uses an advanced state generator and object detection for educational and research purposes.
adipandas/multi-object-tracker
A Python library providing easy-to-use implementations of various multi-object tracking algorithms and OpenCV-based object detectors for video analysis.
parkpow/deep-license-plate-recognition
A collection of example clients and operational utilities for integrating Plate Recognizer's license plate recognition API into images and video streams.
VisualComputingInstitute/diffusion-e2e-ft
A framework for fine-tuning image-conditional diffusion models into single-step deterministic estimators for fast and accurate depth and surface normal estimation.
shoumikchow/bbox-visualizer
A Python package for drawing labeled bounding boxes on images, supporting VOC, COCO, and YOLO formats to simplify object detection visualization.
hysts/pytorch_mpiigaze_demo
A demo program for gaze estimation that uses pretrained models to track where a person is looking in images and videos.
hysts/pytorch_mpiigaze
A PyTorch implementation of gaze estimation methods from the MPIIGaze and MPIIFaceGaze papers for training and evaluation of appearance-based gaze tracking.
luxonis/depthai-ros
A ROS integration for Luxonis OAK cameras, enabling the use of hardware-accelerated AI and depth perception in robotic systems.
FluxML/Metalhead.jl
A library of standard machine learning vision models for Flux.jl, providing implementations of popular architectures like ResNet, ViT, and EfficientNet.
Karine-Huang/T2I-CompBench
A comprehensive benchmark and evaluation suite for text-to-image generation models, focusing on their ability to handle complex compositional prompts, attribute binding, and spatial relationships.
frgfm/holocron
A library of high-quality PyTorch implementations of recent deep learning tricks and model architectures for computer vision tasks.
rapidsai/cucim
cuCIM is a GPU-accelerated computer vision and image processing library designed for large, multidimensional images used in biomedical, geospatial, and remote sensing applications.
theICTlab/3DUNDERWORLD-SLS-GPU_CPU
A structured light scanning tool that reconstructs 3D point clouds from images of objects lit by patterned light, featuring both CPU and GPU acceleration.
alicevision/popsift
A CUDA-accelerated implementation of the SIFT algorithm for real-time image feature extraction.
rgeirhos/Stylized-ImageNet
A tool for creating Stylized-ImageNet, a version of ImageNet that distorts textures to encourage CNNs to prioritize object shape over texture for better robustness.
VSLAM-LAB/VSLAM-LAB
A comprehensive framework for Visual SLAM that streamlines the installation, execution, and evaluation of multiple baseline algorithms across various datasets.
l3p-cv/lost
A flexible, web-based framework for collaborative image annotation that supports semi-automatic pipelines to speed up the labeling of machine learning datasets.
genicam/harvesters
A Python library for simplifying image acquisition in computer vision applications by providing a unified interface for GenICam-compliant devices.
basler/pypylon
Official Python bindings for Basler pylon C++ APIs, enabling Python applications to control Basler machine vision cameras and perform image processing tasks.
cyrusbehr/YOLOv8-TensorRT-CPP
A C++ implementation of YOLOv8 using TensorRT for high-performance GPU inference, supporting object detection, semantic segmentation, and pose estimation.
jenissimo/unfake.js
A JavaScript library and browser tool that cleans up AI-generated pixel art and converts raster images into clean, scalable SVGs.
insight-platform/Savant
A high-level framework for building real-time, high-performance computer vision and video analytics pipelines on Nvidia hardware, abstracting the complexity of DeepStream.
luispedro/mahotas
A fast Python computer vision library implemented in C++ that provides over 100 image processing algorithms operating on NumPy arrays.
opendatacam/opendatacam
An open-source computer vision tool that detects, tracks, and counts moving objects in video feeds to quantify real-world movement, commonly used for traffic studies.
Visionary-Laboratory/visionary
A WebGPU-powered platform for real-time rendering of Gaussian Splatting variants and 3D meshes directly in the browser using ONNX Runtime.
open-edge-platform/geti_v2
An end-to-end computer vision platform that uses active learning and smart annotations to build and deploy optimized AI models with minimal data.
Koldim2001/YOLO-Patch-Based-Inference
A Python library that enables patch-based inference for YOLO models to improve the detection of small objects in large images through image tiling and result consolidation.
zhiqwang/yolort
A runtime stack for YOLOv5 that simplifies object detection deployment by embedding pre- and post-processing into the model graph for various hardware accelerators.
VIAME/VIAME
VIAME is an open‑source computer‑vision toolkit that provides detection, tracking, annotation, search and many other image/video processing capabilities through a modular, multi‑language pipeline framework. It ships with desktop and web GUIs, command‑line tools, pre‑built binaries and Docker images, and can be extended via C++, Python or MATLAB plugins.
azavea/raster-vision
A Python library and low-code framework for building computer vision models on satellite, aerial, and drone imagery using PyTorch.
opengeos/segment-geospatial
A Python package that adapts the Segment Anything Model (SAM) for geospatial data, enabling easy segmentation of satellite imagery using text, points, or boxes.
orbbec/OrbbecSDK
A software development kit for Orbbec 3D cameras that enables developers to capture data streams and control hardware across various device series.
layumi/Person_reID_baseline_pytorch
A lightweight and powerful PyTorch baseline for object and person re-identification, providing a comprehensive set of architectures and loss functions for cross-camera object retrieval.
google-research/augmix
AugMix is a data processing technique that mixes augmented images and enforces consistent embeddings to improve the robustness and and uncertainty calibration of image classification models.
spytensor/plants_disease_detection
A PyTorch-based image classification pipeline for detecting crop diseases, featuring a ResNet50 backbone and modern PyTorch 2.x optimizations.
openclimatefix/metnet
A PyTorch implementation of Google Research's MetNet models for short-term and global precipitation forecasting using satellite and atmospheric data.
FORTH-ModelBasedTracker/MocapNET
A real-time 2D-to-3D human pose estimator that converts RGB video feeds into standard BVH motion capture files for 3D animation.
andrewssobral/bgslibrary
A comprehensive C++ framework for background subtraction in computer vision that provides over 40 algorithms to detect moving objects in video streams.
zhanghang1989/ResNeSt
A Split-Attention Network variant of ResNet that improves performance for computer vision tasks like object detection and semantic segmentation.
apple-aiml-research/ml-facelit
FaceLit is a neural 3D rendering project that enables the generation of high-quality facial images that can be realistically relit under varying lighting conditions.
MrGiovanni/SuPreM
A suite of pre-trained 3D models and the AbdomenAtlas 1.1 dataset for supervised pre-training in 3D medical imaging, improving organ and tumor segmentation.
MrGiovanni/ModelsGenesis
A self-supervised learning framework that pre-trains 3D models on unlabeled CT and MRI scans to create foundation models for medical image segmentation and classification.