serengil/deepface
A lightweight Python framework for face recognition and facial attribute analysis that wraps multiple state-of-the-art models into a single, easy-to-use API.
qubvel-org/segmentation_models.pytorch
A PyTorch library providing a high-level API for image semantic segmentation, featuring 12 architectures and over 800 pretrained encoders.
ultralytics/yolov5
YOLOv5 is a PyTorch‑based, open‑source object detection/segmentation/classification model suite from Ultralytics, offering multiple model sizes, simple Python/CLI inference, one‑line training, and export to many deployment formats.
lyuwenyu/RT-DETR
A real-time object detection framework based on the Detection Transformer (DETR) architecture that aims to outperform YOLO models in speed and accuracy.
royshil/obs-backgroundremoval
An OBS Studio plugin that uses neural networks to remove portrait backgrounds and enhance low-light video in real-time without requiring a physical green screen.
margelo/react-native-vision-camera
A high-performance camera library for React Native that enables professional photo/video capture and real-time AI frame processing.
MrNeRF/LichtFeld-Studio
A modular workstation for 3D Gaussian Splatting that integrates training, real-time visualization, editing, and export in a single native application.
deepfakes/faceswap
A deep learning tool for recognizing and swapping faces in pictures and videos through an extract-train-convert pipeline.
google-deepmind/alphafold3
An implementation of the AlphaFold 3 inference pipeline for the accurate 3D structure prediction of biomolecular interactions.
microsoft/MoGe
MoGe is a model for recovering accurate 3D geometry, including metric depth and point maps, from single open-domain images.
facebookresearch/map-anything
An open-source research framework for universal metric 3D reconstruction that uses a single transformer model to handle multiple 3D tasks from various input combinations.
ArthurBrussee/brush
A cross-platform 3D reconstruction engine using Gaussian splatting that runs on WebGPU and the Burn framework to support macOS, Windows, Linux, Android, and browsers.
davnords/LoMa
LoMa is a family of fast and accurate local feature matchers for images, providing a robust drop-in replacement for existing visual localization and 3D reconstruction pipelines.
datalab-to/surya
A 650M parameter OCR model for document intelligence that provides high-accuracy text recognition, layout analysis, and table extraction across 91 languages.
Project-MONAI/MONAI
A PyTorch-based open-source framework for deep learning in healthcare imaging that provides standardized workflows and domain-specific tools for medical data.
manycoretech/aholo-viewer
A high-performance 3D Gaussian Splatting and Mesh renderer that uses Chunked Streaming LOD to handle massive 3D datasets.
roboflow/sports
A collection of computer vision tools and datasets for sports analytics, focusing on challenges like ball tracking, player re-identification, and jersey number OCR.
LiteReality/LiteReality-Agent
LiteReality‑Agent is an open‑source toolkit that converts a phone‑captured RGB‑D scan (via the LiteReality iOS app) into a fully‑articulated, graphics‑ready indoor 3D scene. An LLM‑driven agent coordinates vision models (TRELLIS, GroundingDINO) and Blender to generate, place, and polish objects, producing GLB/BLEND assets usable in game engines, AR/VR, or robotics. It runs on macOS or Linux, can use Modal’s cloud GPUs or a local ≥24 GB GPU, and requires API keys for image generation and the LLM.
Peterande/D-FINE
D‑FINE is a real‑time object‑detection framework that redefines DETR’s box regression as a fine‑grained distribution refinement and adds a self‑distillation module. It offers several model sizes (‑N to ‑X) that achieve state‑of‑the‑art COCO/AP scores (42.8 %–55.8 %) while running at 70‑470 FPS on a T4 GPU, with no extra inference cost. The repo provides pretrained checkpoints, config files, and full training/evaluation scripts for COCO, Objects365, and custom COCO‑style datasets.
torinmb/mediapipe-touchdesigner
A GPU-accelerated MediaPipe plugin for TouchDesigner that enables real-time computer vision tasks like pose and hand tracking without requiring external installations.
egdels/makeacopy
An offline, privacy-focused Android document scanner that uses on-device ML and OCR to create searchable PDFs for self-hosted workflows.
microsoft/OmniParser
A screen parsing tool that converts GUI screenshots into structured elements, enabling vision-based AI agents to accurately identify and interact with interface components.
valkia/aramgg_client
A Windows desktop assistant for League of Legends ARAM that provides champion selection advice and uses OCR to recommend builds based on in-game Hextech Augments.
open-edge-platform/anomalib
A deep learning library for benchmarking, developing, and deploying visual anomaly detection algorithms to identify and locate defects in images and videos.
Tencent-Hunyuan/HunyuanOCR
A lightweight, end-to-end OCR vision-language model that unifies document parsing and text extraction with speculative decoding for faster inference.
pixpark/gpupixel
A high-performance, cross-platform C++ library for image and video beauty filters using OpenGL/ES.
realsenseai/librealsense
A cross-platform SDK for RealSense depth cameras that enables depth and color streaming and provides calibration data for computer vision and robotics.
TickLabVN/biopass
A multi-modal biometric login system for Linux that enables face and fingerprint authentication with AI-powered anti-spoofing and a GUI manager.
wasserth/TotalSegmentator
An AI-powered tool for the automatic segmentation of most major anatomical structures in CT and MR images, based on the nnU-Net framework.
google/GNM
A family of parametric statistical 3D human models, starting with GNM Head, for high-fidelity representation of human geometry and appearance.
Audiveris/audiveris
An open-source Optical Music Recognition (OMR) application that transcribes music score images into symbolic digital formats for playback and editing.
dweep-desai/FaceGate-Mac
A native macOS app-locker that uses on-device face recognition, Touch ID, or passwords to restrict access to specific applications.
wkentaro/labelme
A graphical image annotation tool for creating datasets for computer vision tasks like segmentation and object detection using manual and AI-assisted labeling.
pq-yang/MatAnyone2
A human video matting framework that preserves fine details and avoids coarse boundaries to extract people from videos with high robustness in real-world conditions.
GunduLabs/gaze
A facial authentication system for Linux that provides on-device face recognition and PAM integration for secure, passwordless login and sudo access.
huggingface/pytorch-image-models
A comprehensive library of state-of-the-art PyTorch image models, layers, and training utilities designed to provide a standardized way to access and reproduce vision model results.
alicevision/Meshroom
A node-based visual programming framework for 3D reconstruction and computer vision, enabling users to build complex data processing pipelines using AI and photogrammetry.
antvis/chart-visualization-skills
AI‑enhanced AntV skill library that lets LLMs generate accurate, ready‑to‑run charts (G2, G6, X6, infographics, narrative text, etc.) via a searchable skill catalogue, CLI, HTTP service, and npm package.
suzuran0y/CCTV-Smartphone-AI-Monitoring
Sentinel is an open‑source LAN‑only system that turns Android phones into real‑time camera nodes and a Python/Flask PC server into a live preview, segmented recorder, and AI‑triggered monitoring dashboard. It uses motion detection to call external vision models, outputs structured JSON events, and keeps all data local for privacy‑focused monitoring or data‑collection research.
WebODM/WebODM
A commercial-grade software for processing drone images into georeferenced maps, point clouds, 3D models, and Gaussian splats.
opengeos/geoai
A Python package that integrates AI frameworks with geospatial data analysis, providing tools for satellite imagery processing, model training, and inference.
facebookresearch/MHR
A high-fidelity parametric 3D human body model that uses identity, pose, and facial expression parameters to generate realistic, differentiable human meshes.
MeshInspector/MeshLib
A high-performance 3D mesh processing SDK that provides tools for repairing, optimizing, and manipulating 3D data across multiple programming languages.
AprilRobotics/apriltag
A visual fiducial system for robotics that detects unique markers in images to provide precise identification and 3D pose estimation.
Intellindust-AI-Lab/DEIMv2
DEIMv2 is a real-time object detection framework leveraging DINOv3 features to provide scalable, state-of-the-art performance across various model sizes for edge and server deployment.
SunOner/sunone_aimbot_2
A C++ based AI aimbot that uses computer vision models to detect targets and automate mouse movement for gaming.
isl-org/Open3D
Open3D is a modern library for 3D data processing that provides optimized tools for point cloud and mesh manipulation, visualization, and 3D machine learning.
SharpAI/DeepCamera
An open-source AI camera platform that enables local deployment of VLM scene analysis and object detection skills with autonomous, hardware-aware installation.