google/cameratrapai

An ensemble of AI models that automates wildlife species classification in camera trap images using an object detector and a species classifier.

mprib/caliscope

A multicamera calibration tool for 3D motion capture that estimates camera intrinsics and spatial positions to enable accurate 3D triangulation.

linwhitehat/ET-BERT

A Transformer-based model for encrypted network traffic classification that learns datagram contextual relationships through pre-training and fine-tuning.

allenv0/AirPosture

An iOS app that uses AirPods' head tracking and local AI to provide real-time posture coaching and nudges to reduce neck strain.

LiteReality/LiteReality

LiteReality converts RGB-D scans from mobile devices into graphics-ready 3D scenes with physically-based rendering materials for high-fidelity reconstruction.

3dem/relion

RELION is a software program for the Maximum A Posteriori refinement of 3D reconstructions and 2D class averages in cryo-electron microscopy.

cvg/glue-factory

A library for training and evaluating deep neural networks that extract and match local visual features, supporting state-of-the-art models like LightGlue and GlueStick.

cvg/limap

A toolbox for holistic 3D mapping and structure from motion that jointly optimizes points, lines, planes, and parametric primitives for more accurate 3D reconstructions.

nvpro-samples/vk_gaussian_splatting

A high-performance Vulkan-based viewer and testbed for 3D Gaussian Splatting models, allowing users to compare rasterization, ray tracing, and hybrid rendering techniques.

OschAI/VisioFirm

An AI-powered image and video annotation tool that uses models like SAM2 and YOLO to provide semi-automated pre-labeling for computer vision datasets.

PRBonn/RAP

A 3D point cloud registration framework that uses flow matching to align multiple unposed scans into a common frame by treating registration as a conditional generation task.

UB-Mannheim/zotero-ocr

A Zotero plugin that uses Tesseract OCR to add a searchable text layer to scanned PDFs within a Zotero library.

Linketic/CityGaussian

A framework for high-quality, large-scale 3D scene reconstruction using Gaussian Splatting, featuring multi-GPU support and geometrically accurate rendering.

drivelineresearch/openbiomechanics

A public research dataset providing cleaned motion-capture files, time-series biomechanics, and physical assessment metrics for baseball pitching and hitting.

ScrollPrize/villa

A collection of machine learning and computer vision tools for the Vesuvius Challenge to virtually unwrap and read ancient Herculaneum scrolls from CT scans.

dstl/Stone-Soup

Stone Soup is a framework for the development and testing of target tracking and state estimation algorithms.

SubhamTyagi/android-ocr

A privacy-focused Android application that uses Tesseract 5 to perform offline optical character recognition (OCR) in over 120 languages.

hmorimitsu/ptlflow

A unified PyTorch Lightning framework for training and evaluating a wide collection of state-of-the-art optical flow estimation models.

oliverbravery/PrintGuard

A local vision-based monitoring system that detects 3D print failures in real-time to automatically pause printers and alert users, avoiding wasted filament.

inspatio/querysplat

QuerySplat is a 3D Gaussian Splatting prediction framework that decouples geometry and appearance representations to improve the quality of 3D scene reconstructions from images.

bytedeco/javacv

A Java wrapper for popular computer vision and video processing libraries like OpenCV and FFmpeg, enabling high-performance image and video manipulation on Java and Android.

tomas-gajarsky/facetorch

A Python library for facial detection and analysis that curates open-source models and packages them as portable torch.export models for efficient inference.

ch-sa/labelCloud

A lightweight 3D bounding box labeling tool for point clouds, used to generate training data for 3D object detection and 6D pose estimation.

databricks-industry-solutions/pixels

A medical imaging solution accelerator that ingests and indexes DICOM files to enable SQL-based metadata analysis and AI-powered image segmentation using MONAI.

fastvideo/gpu-camera-sample

A high-performance GPU-accelerated camera application for real-time raw image processing and ISP pipelines, supporting a wide range of industrial cameras.

jamjamjon/usls

A cross-platform Rust library powered by ONNX Runtime for efficient inference of vision and vision-language models under 1B parameters.

pymatting/pymatting

A Python library for alpha matting that estimates the transparency of foreground objects using trimaps to enable precise background removal and composition.

apple-aiml-research/ml-hypersim

Hypersim is a large synthetic indoor dataset (≈77 k HDR images, 1.9 TB) with per‑pixel geometry, semantics, and material channels, plus a V‑Ray‑based toolkit for generating and editing similar data. It’s designed for training and evaluating scene‑understanding models (depth, normals, segmentation, intrinsic image decomposition) and includes a ready‑made train/val/test split.

apple-aiml-research/ml-hugs

HUGS is a reference implementation for reconstructing an animatable human and their surrounding background scene from a single video using Gaussian Splatting.

jasongzy/Make-It-Animatable

An efficient framework that automates the creation of animation-ready 3D characters by predicting joint positions and skinning weights from 3D meshes.

nmwsharp/polyscope

A lightweight C++/Python 3D viewer and UI for quickly visualizing meshes and point clouds with associated scalar and vector data.

Anttwo/Surflo

Surflo is a 3D surface flow model that turns a handful of unposed RGB views into detailed 3D meshes using a global state and flow matching.

amebalabs/TRex

A macOS utility that uses OCR to extract non-selectable text from any screen area and copy it directly to the clipboard.

rtr46/meikipop

A universal Japanese OCR popup dictionary that enables instant word lookups from any on-screen content, including games, manga, and videos.

deepdoctection/deepdoctection

A Python library for document understanding that orchestrates layout analysis, OCR, and token classification to extract structured data from PDFs and scans.

nutonomy/nuscenes-devkit

A software development kit for the nuScenes and nuImages datasets, providing tools to load, analyze, and evaluate multimodal sensor data for autonomous driving research.

espressif/esp-who

An image processing development platform for Espressif chips that provides examples for face and pedestrian detection and QR code recognition.

facebookresearch/OrienterNet

A deep neural network for visual localization that determines an image's position and orientation by matching it to 2D public maps like OpenStreetMap.

mjkwon2021/CAT-Net

A network for detecting and localizing image manipulations by analyzing JPEG compression artifacts, providing tools for image forensics research.

Project-MONAI/MONAILabel

An intelligent open-source ecosystem for AI-assisted medical image annotation that uses a server-client architecture to enable interactive and automated labeling for radiology, pathology, and endoscopy.

yihong1120/Construction-Hazard-Detection

An AI-driven construction site safety monitoring system that uses YOLO to detect PPE violations and proximity hazards from live camera feeds.

zae-bayern/elpv-dataset

A benchmark dataset of 2,624 annotated electroluminescence images of solar cells used for the visual identification of defects that reduce power efficiency.

Climate-Vision/ClimateVision

An open-source machine learning platform that uses deep learning and satellite imagery to automatically detect deforestation, arctic ice melting, and flooding.

cuevhv/mamma

MAMMA is a markerless multi-person 3D motion acquisition system that reconstructs accurate human poses from multi-view video footage.

emgucv/emgucv

A cross-platform .NET wrapper for the OpenCV library that enables image-processing functions to be used in .NET compatible languages.

tianrun-chen/SAM-Adapter-PyTorch

A framework for adapting the Segment Anything Model (SAM) to improve performance in challenging scenes like camouflage, shadows, and medical imaging.

cameraui/camera.ui

A self-hosted security camera platform providing live viewing, recording, and on-device AI detection while keeping footage on user-owned hardware.

facebookresearch/ocean

Ocean is a C++ framework for building Computer Vision and Augmented Reality applications across various platforms including mobile, desktop, and Meta Quest.