google/cameratrapai
An ensemble of AI models that automates wildlife species classification in camera trap images using an object detector and a species classifier.
mprib/caliscope
A multicamera calibration tool for 3D motion capture that estimates camera intrinsics and spatial positions to enable accurate 3D triangulation.
linwhitehat/ET-BERT
A Transformer-based model for encrypted network traffic classification that learns datagram contextual relationships through pre-training and fine-tuning.
allenv0/AirPosture
An iOS app that uses AirPods' head tracking and local AI to provide real-time posture coaching and nudges to reduce neck strain.
LiteReality/LiteReality
LiteReality converts RGB-D scans from mobile devices into graphics-ready 3D scenes with physically-based rendering materials for high-fidelity reconstruction.
3dem/relion
RELION is a software program for the Maximum A Posteriori refinement of 3D reconstructions and 2D class averages in cryo-electron microscopy.
cvg/glue-factory
A library for training and evaluating deep neural networks that extract and match local visual features, supporting state-of-the-art models like LightGlue and GlueStick.
cvg/limap
A toolbox for holistic 3D mapping and structure from motion that jointly optimizes points, lines, planes, and parametric primitives for more accurate 3D reconstructions.
nvpro-samples/vk_gaussian_splatting
A high-performance Vulkan-based viewer and testbed for 3D Gaussian Splatting models, allowing users to compare rasterization, ray tracing, and hybrid rendering techniques.
OschAI/VisioFirm
An AI-powered image and video annotation tool that uses models like SAM2 and YOLO to provide semi-automated pre-labeling for computer vision datasets.
PRBonn/RAP
A 3D point cloud registration framework that uses flow matching to align multiple unposed scans into a common frame by treating registration as a conditional generation task.
UB-Mannheim/zotero-ocr
A Zotero plugin that uses Tesseract OCR to add a searchable text layer to scanned PDFs within a Zotero library.
Linketic/CityGaussian
A framework for high-quality, large-scale 3D scene reconstruction using Gaussian Splatting, featuring multi-GPU support and geometrically accurate rendering.
drivelineresearch/openbiomechanics
A public research dataset providing cleaned motion-capture files, time-series biomechanics, and physical assessment metrics for baseball pitching and hitting.
ScrollPrize/villa
A collection of machine learning and computer vision tools for the Vesuvius Challenge to virtually unwrap and read ancient Herculaneum scrolls from CT scans.
dstl/Stone-Soup
Stone Soup is a framework for the development and testing of target tracking and state estimation algorithms.
SubhamTyagi/android-ocr
A privacy-focused Android application that uses Tesseract 5 to perform offline optical character recognition (OCR) in over 120 languages.
hmorimitsu/ptlflow
A unified PyTorch Lightning framework for training and evaluating a wide collection of state-of-the-art optical flow estimation models.
oliverbravery/PrintGuard
A local vision-based monitoring system that detects 3D print failures in real-time to automatically pause printers and alert users, avoiding wasted filament.
inspatio/querysplat
QuerySplat is a 3D Gaussian Splatting prediction framework that decouples geometry and appearance representations to improve the quality of 3D scene reconstructions from images.
bytedeco/javacv
A Java wrapper for popular computer vision and video processing libraries like OpenCV and FFmpeg, enabling high-performance image and video manipulation on Java and Android.
tomas-gajarsky/facetorch
A Python library for facial detection and analysis that curates open-source models and packages them as portable torch.export models for efficient inference.
ch-sa/labelCloud
A lightweight 3D bounding box labeling tool for point clouds, used to generate training data for 3D object detection and 6D pose estimation.
databricks-industry-solutions/pixels
A medical imaging solution accelerator that ingests and indexes DICOM files to enable SQL-based metadata analysis and AI-powered image segmentation using MONAI.
fastvideo/gpu-camera-sample
A high-performance GPU-accelerated camera application for real-time raw image processing and ISP pipelines, supporting a wide range of industrial cameras.
jamjamjon/usls
A cross-platform Rust library powered by ONNX Runtime for efficient inference of vision and vision-language models under 1B parameters.
pymatting/pymatting
A Python library for alpha matting that estimates the transparency of foreground objects using trimaps to enable precise background removal and composition.
apple-aiml-research/ml-hypersim
Hypersim is a large synthetic indoor dataset (≈77 k HDR images, 1.9 TB) with per‑pixel geometry, semantics, and material channels, plus a V‑Ray‑based toolkit for generating and editing similar data. It’s designed for training and evaluating scene‑understanding models (depth, normals, segmentation, intrinsic image decomposition) and includes a ready‑made train/val/test split.
apple-aiml-research/ml-hugs
HUGS is a reference implementation for reconstructing an animatable human and their surrounding background scene from a single video using Gaussian Splatting.
jasongzy/Make-It-Animatable
An efficient framework that automates the creation of animation-ready 3D characters by predicting joint positions and skinning weights from 3D meshes.
nmwsharp/polyscope
A lightweight C++/Python 3D viewer and UI for quickly visualizing meshes and point clouds with associated scalar and vector data.
Anttwo/Surflo
Surflo is a 3D surface flow model that turns a handful of unposed RGB views into detailed 3D meshes using a global state and flow matching.
amebalabs/TRex
A macOS utility that uses OCR to extract non-selectable text from any screen area and copy it directly to the clipboard.
rtr46/meikipop
A universal Japanese OCR popup dictionary that enables instant word lookups from any on-screen content, including games, manga, and videos.
deepdoctection/deepdoctection
A Python library for document understanding that orchestrates layout analysis, OCR, and token classification to extract structured data from PDFs and scans.
nutonomy/nuscenes-devkit
A software development kit for the nuScenes and nuImages datasets, providing tools to load, analyze, and evaluate multimodal sensor data for autonomous driving research.
espressif/esp-who
An image processing development platform for Espressif chips that provides examples for face and pedestrian detection and QR code recognition.
facebookresearch/OrienterNet
A deep neural network for visual localization that determines an image's position and orientation by matching it to 2D public maps like OpenStreetMap.
mjkwon2021/CAT-Net
A network for detecting and localizing image manipulations by analyzing JPEG compression artifacts, providing tools for image forensics research.
Project-MONAI/MONAILabel
An intelligent open-source ecosystem for AI-assisted medical image annotation that uses a server-client architecture to enable interactive and automated labeling for radiology, pathology, and endoscopy.
yihong1120/Construction-Hazard-Detection
An AI-driven construction site safety monitoring system that uses YOLO to detect PPE violations and proximity hazards from live camera feeds.
zae-bayern/elpv-dataset
A benchmark dataset of 2,624 annotated electroluminescence images of solar cells used for the visual identification of defects that reduce power efficiency.
Climate-Vision/ClimateVision
An open-source machine learning platform that uses deep learning and satellite imagery to automatically detect deforestation, arctic ice melting, and flooding.
cuevhv/mamma
MAMMA is a markerless multi-person 3D motion acquisition system that reconstructs accurate human poses from multi-view video footage.
emgucv/emgucv
A cross-platform .NET wrapper for the OpenCV library that enables image-processing functions to be used in .NET compatible languages.
tianrun-chen/SAM-Adapter-PyTorch
A framework for adapting the Segment Anything Model (SAM) to improve performance in challenging scenes like camouflage, shadows, and medical imaging.
cameraui/camera.ui
A self-hosted security camera platform providing live viewing, recording, and on-device AI detection while keeping footage on user-owned hardware.
facebookresearch/ocean
Ocean is a C++ framework for building Computer Vision and Augmented Reality applications across various platforms including mobile, desktop, and Meta Quest.