liebharc/homr
An Optical Music Recognition (OMR) software that transforms photos or PDFs of sheet music into machine-readable MusicXML files.
cvg/vidmap
An offline Structure-from-Motion system for video that leverages temporal tracks, loop closures, and metric depth to estimate camera poses and generate sparse 3D maps.
nv-tlabs/NKSR
Neural Kernel Surface Reconstruction (NKSR) is a method for creating high-quality 3D implicit surfaces from large-scale, sparse, and noisy point clouds, capable of handling millions of points in seconds.
ultralytics/yolov3
Ultralytics YOLOv3 is a PyTorch implementation of the YOLOv3 object‑detection model, providing three variants (full, SPP, tiny) with scripts for training, inference, validation, and export to many deployment formats. It includes automatic COCO‑pretrained weights, supports PyTorch Hub loading, and integrates with popular ML‑ops tools and cloud notebooks. Licensed under AGPL‑3.0 with a commercial Enterprise option.
lightly-ai/lightly
A computer vision framework for self-supervised learning that provides modular building blocks and implementations of numerous state-of-the-art SSL models.
google-ai-edge/model-explorer
Model Explorer is a Google‑AI‑Edge open‑source visualizer that renders TensorFlow, TFLite, TF‑JS, MLIR, and PyTorch ExportedProgram models as interactive hierarchical graphs. It offers GPU‑accelerated rendering, searchable nodes, I/O highlighting, metadata overlays, and an extensible adapter system for additional formats.
ndl-lab/ndlocr-lite
A lightweight OCR tool designed for home computers to extract text from digitized books and magazines without requiring a GPU.
albumentations-team/AlbumentationsX
A high-performance Python library for image augmentation that creates synthetic training samples to improve the quality and robustness of computer vision models.
bytedance/Protenix
An open-source biomolecular structure prediction framework that provides high-accuracy 3D modeling of proteins and RNA, designed to be an accessible alternative to AlphaFold3.
facebookresearch/MHR
A high-fidelity parametric 3D human body model that uses identity, pose, and facial expression parameters to generate realistic, differentiable human meshes.
triple-mu/YOLOv8-TensorRT
A high-performance TensorRT acceleration wrapper for YOLOv8, enabling fast inference for detection, segmentation, pose, OBB, and classification in both Python and C++.
cvg/GeoCalib
GeoCalib is a research‑grade Python library that predicts camera intrinsics and scene gravity from a single image using a deep network plus geometric optimisation. It supports multiple lens models, partial priors, batch and multi‑camera rig calibration, and ships with easy‑install inference code, a Gradio web demo, a Colab notebook, and a live webcam demo. Pre‑trained models are trained on the OpenPano dataset (≈37 k perspective crops from HDR panoramas) and achieve state‑of‑the‑art results on benchmarks such as LaMAR, MegaDepth, TartanAir, and Stanford2D3D.
SunOner/sunone_aimbot_2
A C++ based AI aimbot that uses computer vision models to detect targets and automate mouse movement for gaming.
orbbec/pyorbbecsdk
Python bindings for the Orbbec SDK v2.x, enabling developers to interface with RGB-D cameras for depth streaming, 3D point cloud generation, and computer vision tasks.
luicfrr/react-native-vision-camera-face-detector
A React Native library that integrates Google MLKit with Vision Camera to provide real-time and static image face detection.
facebookresearch/detectron2
Detectron2 is a high-performance computer vision library from Meta AI Research providing state-of-the-art object detection and segmentation algorithms.
LazoVelko/neverclick
A desktop application that uses computer vision to let users perform mouse actions via their keyboard, reducing the need for a physical mouse.
dweep-desai/FaceGate-Mac
A native macOS app-locker that uses on-device face recognition, Touch ID, or passwords to restrict access to specific applications.
mne-tools/mne-python
An open-source Python package for exploring, visualizing, and analyzing human neurophysiological data like MEG and EEG.
egdels/makeacopy
An offline, privacy-focused Android document scanner that uses on-device ML and OCR to create searchable PDFs for self-hosted workflows.
krupkat/xpano
A tool for panorama stitching that automates image group detection and provides projection adjustments for export of full-resolution panoramas.
sceneview/sceneview
A multi-platform 3D and AR framework that allows developers to build immersive experiences using declarative UI frameworks like Jetpack Compose and SwiftUI.
blakeblackshear/frigate-hass-integration
A Home Assistant integration that connects Frigate's AI-powered object detection and NVR capabilities to a smart home hub for automated security monitoring.
perfanalytics/pose2sim
A markerless 3D kinematics workflow that converts 2D video from low-cost cameras into research-grade 3D joint angles and skeletal analysis using OpenSim.
playcanvas/supersplat-viewer
A high-performance web viewer for Gaussian Splatting scenes, allowing developers to embed interactive 3D captures into websites using WebGPU or WebGL.
microsoft/aurora
Aurora is a foundation model for the Earth system that predicts atmospheric variables like temperature, air pollution, and ocean waves by adapting a general-purpose pre-trained model to specialized forecasting tasks.
yatengLG/ISAT_with_segment_anything
An interactive semi-automatic image segmentation annotation tool that uses the Segment Anything Model (SAM) to accelerate the creation of labeled datasets.
secluso/core
A private, DIY home security camera system for Raspberry Pi that uses end-to-end encryption to avoid cloud surveillance.
mahmoodlab/TRIDENT
A toolkit for large-scale whole-slide image processing in AI pathology, providing an end-to-end pipeline for tissue segmentation and feature extraction using numerous foundation models.
CERN/TIGRE
A GPU-accelerated toolbox for fast and accurate 3D tomographic reconstruction using Python and MATLAB, supporting a wide range of iterative algorithms and flexible CT geometries.
InsightSoftwareConsortium/ITK
An open-source, cross-platform toolkit for N-dimensional scientific image processing, segmentation, and registration, primarily used for medical image analysis.
mtalcott/google-photos-deduper
A Chrome extension that uses local image embeddings to find and remove duplicate photos from a Google Photos library without requiring OAuth setup.
Geekgineer/YOLOs-CPP
A production-ready C++ inference engine that provides a unified API for the entire YOLO family, supporting tasks from object detection to metric depth estimation.
alicevision/AliceVision
A photogrammetric computer vision framework that provides algorithms for 3D reconstruction and camera tracking from unordered photographs or videos.
chaofengc/IQA-PyTorch
A PyTorch-based toolbox for Image Quality Assessment (IQA) that provides GPU-accelerated implementations of numerous full-reference and no-reference quality metrics.
Babyhamsta/Aimmy
An AI-based aim alignment tool that uses YOLOv8 and DirectML to help gamers with physical or visual impairments improve their accuracy in FPS games.
pur1fying/blue_archive_auto_script
An automation script for Blue Archive that uses YOLOv8 for character detection to automate combat, daily tasks, and resource management.
withoutbg/withoutbg-python
A Python library for removing image backgrounds using either a free local ONNX model or a professional Cloud API.
apple-aiml-research/ARKitScenes
A large-scale RGB-D dataset of indoor scenes captured via mobile LiDAR, providing ground truth for 3D object detection and depth upsampling.
lukasHoel/video_to_world
A 3D reconstruction pipeline that resolves geometric inconsistencies in videos generated by video diffusion models to create stable 3D worlds.
Slicer/Slicer
3D Slicer is a free, open-source software package for the visualization and analysis of 3D medical imaging data.
qupath/qupath
QuPath is open-source software for bioimage analysis that provides tools for annotating and analyzing whole slide and microscopy images using machine learning and specialized algorithms.
mudler/locate-anything.cpp
A C++/ggml inference port of NVIDIA's LocateAnything-3B that enables fast, open-vocabulary object detection on CPU and GPU without a Python runtime.
deepinv/deepinv
DeepInverse is an open-source PyTorch-based library designed to solve imaging inverse problems using deep learning through a modular framework of operators and algorithms.
Junjue-Wang/LoveDA
A high-resolution remote sensing land-cover dataset for improving semantic segmentation and domain adaptation in urban and rural environments.
chenwei-zhao/captcha-recognizer
A deep learning-based library for identifying the gap coordinates and confidence levels in slider CAPTCHAs.
privatenumber/mac-ocr
A macOS command-line tool and Node.js API for local, private OCR that extracts text from images and PDFs and creates searchable PDFs using Apple's Vision framework.
nsfw-filter/nsfw-filter
A privacy-focused browser extension that uses on-device AI to block NSFW images without uploading data to any server.