Tobias-Fischer/rt_gene

A real-time eye gaze and blink estimation system using PyTorch and ROS 2 for tracking gaze direction and blink probability in natural environments.

TIGER-AI-Lab/Pixel-Reasoner

Pixel Reasoner is a research framework that adds visual‑manipulation operations (zoom‑in, select‑frame, etc.) to a Vision‑Language Model, training it via instruction‑tuning and curiosity‑driven reinforcement learning. The repo supplies installation guides, scripts for both SFT and RL stages, pretrained 7 B checkpoints, datasets, and evaluation pipelines for image and video benchmarks.

jabberjabberjabber/ImageIndexer

A local AI tool that automatically generates captions and keywords for images and writes them directly into image metadata for easy indexing and searching.

yuanze-lin/Olympus

A universal task router for computer vision that converts a single natural-language instruction into a sequence of specialist model calls to produce images, videos, and 3D models.

Jumpat/SegAnyGAussians

SAGA is a framework for segmenting 3D Gaussian Splatting scenes, allowing users to isolate 3D objects using interactive point prompts or open-vocabulary text queries.

jasmcaus/caer

Caer is a Python‑only, GPU‑accelerated computer‑vision library that offers fast image/video utilities (resize, colour conversion, augmentations, etc.) with a clean, type‑checked API. It can replace OpenCV for research prototypes, is installable via `pip`, and is MIT‑licensed.

Pbatch/ClashRoyaleBuildABot

An automation bot for Clash Royale that uses an advanced state generator and object detection for educational and research purposes.

adipandas/multi-object-tracker

A Python library providing easy-to-use implementations of various multi-object tracking algorithms and OpenCV-based object detectors for video analysis.

parkpow/deep-license-plate-recognition

A collection of example clients and operational utilities for integrating Plate Recognizer's license plate recognition API into images and video streams.

VisualComputingInstitute/diffusion-e2e-ft

A framework for fine-tuning image-conditional diffusion models into single-step deterministic estimators for fast and accurate depth and surface normal estimation.

shoumikchow/bbox-visualizer

A Python package for drawing labeled bounding boxes on images, supporting VOC, COCO, and YOLO formats to simplify object detection visualization.

hysts/pytorch_mpiigaze_demo

A demo program for gaze estimation that uses pretrained models to track where a person is looking in images and videos.

hysts/pytorch_mpiigaze

A PyTorch implementation of gaze estimation methods from the MPIIGaze and MPIIFaceGaze papers for training and evaluation of appearance-based gaze tracking.

luxonis/depthai-ros

A ROS integration for Luxonis OAK cameras, enabling the use of hardware-accelerated AI and depth perception in robotic systems.

FluxML/Metalhead.jl

A library of standard machine learning vision models for Flux.jl, providing implementations of popular architectures like ResNet, ViT, and EfficientNet.

Karine-Huang/T2I-CompBench

A comprehensive benchmark and evaluation suite for text-to-image generation models, focusing on their ability to handle complex compositional prompts, attribute binding, and spatial relationships.

frgfm/holocron

A library of high-quality PyTorch implementations of recent deep learning tricks and model architectures for computer vision tasks.

rapidsai/cucim

cuCIM is a GPU-accelerated computer vision and image processing library designed for large, multidimensional images used in biomedical, geospatial, and remote sensing applications.

theICTlab/3DUNDERWORLD-SLS-GPU_CPU

A structured light scanning tool that reconstructs 3D point clouds from images of objects lit by patterned light, featuring both CPU and GPU acceleration.

alicevision/popsift

A CUDA-accelerated implementation of the SIFT algorithm for real-time image feature extraction.

rgeirhos/Stylized-ImageNet

A tool for creating Stylized-ImageNet, a version of ImageNet that distorts textures to encourage CNNs to prioritize object shape over texture for better robustness.

VSLAM-LAB/VSLAM-LAB

A comprehensive framework for Visual SLAM that streamlines the installation, execution, and evaluation of multiple baseline algorithms across various datasets.

l3p-cv/lost

A flexible, web-based framework for collaborative image annotation that supports semi-automatic pipelines to speed up the labeling of machine learning datasets.

genicam/harvesters

A Python library for simplifying image acquisition in computer vision applications by providing a unified interface for GenICam-compliant devices.

basler/pypylon

Official Python bindings for Basler pylon C++ APIs, enabling Python applications to control Basler machine vision cameras and perform image processing tasks.

cyrusbehr/YOLOv8-TensorRT-CPP

A C++ implementation of YOLOv8 using TensorRT for high-performance GPU inference, supporting object detection, semantic segmentation, and pose estimation.

jenissimo/unfake.js

A JavaScript library and browser tool that cleans up AI-generated pixel art and converts raster images into clean, scalable SVGs.

insight-platform/Savant

A high-level framework for building real-time, high-performance computer vision and video analytics pipelines on Nvidia hardware, abstracting the complexity of DeepStream.

luispedro/mahotas

A fast Python computer vision library implemented in C++ that provides over 100 image processing algorithms operating on NumPy arrays.

opendatacam/opendatacam

An open-source computer vision tool that detects, tracks, and counts moving objects in video feeds to quantify real-world movement, commonly used for traffic studies.

Visionary-Laboratory/visionary

A WebGPU-powered platform for real-time rendering of Gaussian Splatting variants and 3D meshes directly in the browser using ONNX Runtime.

open-edge-platform/geti_v2

An end-to-end computer vision platform that uses active learning and smart annotations to build and deploy optimized AI models with minimal data.

Koldim2001/YOLO-Patch-Based-Inference

A Python library that enables patch-based inference for YOLO models to improve the detection of small objects in large images through image tiling and result consolidation.

zhiqwang/yolort

A runtime stack for YOLOv5 that simplifies object detection deployment by embedding pre- and post-processing into the model graph for various hardware accelerators.

VIAME/VIAME

VIAME is an open‑source computer‑vision toolkit that provides detection, tracking, annotation, search and many other image/video processing capabilities through a modular, multi‑language pipeline framework. It ships with desktop and web GUIs, command‑line tools, pre‑built binaries and Docker images, and can be extended via C++, Python or MATLAB plugins.

azavea/raster-vision

A Python library and low-code framework for building computer vision models on satellite, aerial, and drone imagery using PyTorch.

opengeos/segment-geospatial

A Python package that adapts the Segment Anything Model (SAM) for geospatial data, enabling easy segmentation of satellite imagery using text, points, or boxes.

orbbec/OrbbecSDK

A software development kit for Orbbec 3D cameras that enables developers to capture data streams and control hardware across various device series.

layumi/Person_reID_baseline_pytorch

A lightweight and powerful PyTorch baseline for object and person re-identification, providing a comprehensive set of architectures and loss functions for cross-camera object retrieval.

google-research/augmix

AugMix is a data processing technique that mixes augmented images and enforces consistent embeddings to improve the robustness and and uncertainty calibration of image classification models.

spytensor/plants_disease_detection

A PyTorch-based image classification pipeline for detecting crop diseases, featuring a ResNet50 backbone and modern PyTorch 2.x optimizations.

openclimatefix/metnet

A PyTorch implementation of Google Research's MetNet models for short-term and global precipitation forecasting using satellite and atmospheric data.

FORTH-ModelBasedTracker/MocapNET

A real-time 2D-to-3D human pose estimator that converts RGB video feeds into standard BVH motion capture files for 3D animation.

andrewssobral/bgslibrary

A comprehensive C++ framework for background subtraction in computer vision that provides over 40 algorithms to detect moving objects in video streams.

zhanghang1989/ResNeSt

A Split-Attention Network variant of ResNet that improves performance for computer vision tasks like object detection and semantic segmentation.

apple-aiml-research/ml-facelit

FaceLit is a neural 3D rendering project that enables the generation of high-quality facial images that can be realistically relit under varying lighting conditions.

MrGiovanni/SuPreM

A suite of pre-trained 3D models and the AbdomenAtlas 1.1 dataset for supervised pre-training in 3D medical imaging, improving organ and tumor segmentation.

MrGiovanni/ModelsGenesis

A self-supervised learning framework that pre-trains 3D models on unlabeled CT and MRI scans to create foundation models for medical image segmentation and classification.