baidu/Unlimited-OCR
A multimodal OCR model for one-shot long-horizon parsing of complex documents, PDFs, and multi-page images.
ruvnet/RuView
A WiFi sensing platform that turns radio signals into spatial intelligence to detect people, track vitals, and estimate pose without cameras or wearables.
immich-app/immich
A high-performance self-hosted photo and video management solution featuring AI-powered facial recognition and CLIP-based search.
PaddlePaddle/PaddleOCR
PaddleOCR is an open‑source OCR and document‑AI toolkit that converts images/PDFs into structured, LLM‑ready JSON or Markdown. It offers multilingual scene‑text models (PP‑OCRv6), a lightweight vision‑language parser (PaddleOCR‑VL‑1.6) for tables, formulas, and ancient scripts, and a structure‑aware pipeline (PP‑StructureV3). Models run on CPU, GPU, XPU, NPU or via ONNX/TensorRT, and integrate with RAG/agent platforms like Dify and RAGFlow.
Faceplugin-ltd/Open-Source-Face-Recognition-SDK
An open-source face recognition SDK for Windows and Linux that provides on-premise face detection, landmark detection, and similarity comparison using deep learning.
bhimrazy/receipt-ocr
An OCR engine that extracts either raw text or structured JSON data from receipt images using Tesseract and LLMs.
hacksider/Deep-Live-Cam
A real-time face-swapping and deepfake tool that allows users to replace faces in live webcam streams or videos using a single source image.
amap-cvlab/ABot-Recon
ABot-Recon is a streaming 3D reconstruction system that uses a fixed 12-frame local context to efficiently build global trajectories and point clouds from long video streams without needing persistent long-range memory.
ultralytics/ultralytics
A high-performance computer vision framework implementing YOLO models for real-time object detection, segmentation, classification, and pose estimation.
ant-research/4DAnyone
A tool that converts a single monocular video of a person into multi-view videos to enable high-quality 4D Gaussian Splatting reconstruction.
roryclear/clearcam
An AI-powered NVR system that adds object detection, tracking, and AI-generated notification summaries to any RTSP security camera.
MaaAssistantArknights/MaaAssistantArknights
An automation assistant for Arknights that uses image recognition and OCR to automate daily tasks, resource farming, and base management.
mkturkcan/DART
A training-free framework that converts SAM3 into a real-time multi-class open-vocabulary detector using TensorRT optimization and distilled backbones.
ruoyuw22-byte/NeuroMorph-Assessment
An automated structural MRI morphometry system that integrates CAT12 processing to measure brain tissue volumes and generate quantitative reports.
om-ai-lab/VLX-Seek
VLX‑Seek is an open‑source vision‑language model that performs fine‑grained perception for edge‑device agents. It generates region proposals, encodes each region as a special token, and lets the language model answer by selecting those tokens—producing compact, parsable outputs for detection, referring‑expression comprehension, OCR, counting, and VQA. The repo ships inference code, a Python API, and the 10 B checkpoint (weights hosted on Hugging Face).
julyx10/lap
An open-source, local-first photo manager that uses local AI for search and face recognition to provide a private alternative to cloud photo services.
roboflow/rf-detr
RF-DETR is a real-time transformer architecture for object detection, instance segmentation, and keypoint detection built on a DINOv2 backbone.
blakeblackshear/frigate
A local NVR with real-time AI object detection for IP cameras, designed for tight integration with Home Assistant.
RapidAI/RapidOCR
An open-source OCR tool that converts PaddleOCR models to ONNX format for fast, low-resource offline deployment across multiple platforms.
Robbyant/lingbot-map
A feed-forward 3D foundation model for streaming 3D reconstruction that enables real-time generation of point clouds and camera trajectories from video.
facebookresearch/vggt-omega
VGGT-Omega is a computer vision model that estimates camera poses and depth from images to reconstruct 3D scenes.
roboflow/supervision
A computer vision toolkit that provides model-agnostic building blocks for data loading, visualization, and dataset management to accelerate the development of vision applications.
tesseract-ocr/tesseract
Tesseract is an open-source OCR engine and command-line tool that converts images of text into machine-readable text across more than 100 languages.
thiagotigaz/ocr-it
A Chrome and Firefox extension that uses local OCR to extract text from paginated web documents and viewers that prevent text selection.
harry7557558/spirula-studio
A self-contained 3D Gaussian Splatting pipeline that transforms photos and videos into 3D splats and meshes without requiring Python or PyTorch.
opencv/opencv
OpenCV is an open-source computer vision library that provides tools and algorithms for AI and visual perception development.
ssrajadh/sentrysearch
A semantic video search tool that lets you find specific events in footage using natural language or images and automatically trims the matching clips.
microsoft/MoGe
MoGe is a model for recovering accurate 3D geometry, including metric depth and point maps, from single open-domain images.
ace-trump-tech/DeltaForce-OBS-Locker
A real-time object detection and target locking system for Delta Force using YOLOv14 to identify game characters while filtering out environmental false positives.
duy-phamduc68/TrafficLab-3D
An end-to-end traffic analysis suite that creates 3D digital twins from CCTV footage and satellite maps using computer vision and a custom calibration pipeline.
facebookresearch/sam3
SAM 3 (Segment Anything with Concepts) is Meta’s 848 M‑parameter foundation model that lets you prompt an image or video with free‑form text (or visual exemplars) and receive masks, boxes, and scores for *all* matching objects. It combines a DETR‑style detector and a SAM 2‑style tracker, introduces a presence token for fine‑grained prompt discrimination, and is trained on >4 M auto‑annotated concepts. The repo provides installation steps, example notebooks, and a new SA‑CO benchmark (270 K concepts) for evaluation.
ByteDance-Seed/Depth-Anything-3
A model that predicts spatially consistent 3D geometry, depth maps, and camera poses from single or multiple images using a unified depth-ray representation.
colmap/colmap
A general-purpose Structure-from-Motion and Multi-View Stereo pipeline for reconstructing 3D structures from image collections.
ocrmypdf/OCRmyPDF
A command-line tool that adds an OCR text layer to scanned PDFs, making them searchable and copy-pasteable while maintaining image quality.
BigBodyCobain/Shadowbroker
A decentralized geospatial intelligence platform that aggregates 60+ real-time OSINT feeds into a single map interface for global threat monitoring.
MaaEnd/MaaEnd
An automation assistant for the game Endfield that uses screen recognition to automate combat, resource farming, and daily management tasks.
Yuliang-Liu/MonkeyOCRv2
MonkeyOCRv2 is an open‑source visual‑text foundation model for Document AI. It ships a ViT‑based vision encoder plus multilingual parsing and understanding heads (≈0.6‑1.8 B parameters) that achieve state‑of‑the‑art scores on OCR, layout parsing, VQA and formula recognition across 17 languages. Models are available on Hugging Face/ModelScope, can be run with Transformers or vLLM (CPU support and DFlash acceleration optional), and are accompanied by the large MonkeyDoc v2 pre‑training dataset.
ZhengPeng7/BiRefNet
BiRefNet is a high-resolution dichotomous image segmentation model that provides state-of-the-art background removal and object segmentation for images of varying resolutions.
oncologylab/histopia
A computational research tool for serial-section histology and proteomic image analysis, providing 3D reconstruction, inter-section alignment, and spatial topology profiling.
deepinsight/insightface
An open-source 2D and 3D face analysis toolbox providing state-of-the-art algorithms for face detection, recognition, and alignment.
MaaXYZ/MaaFramework
A cross-platform, image-recognition-based automation framework for black-box testing and creating software assistants.
ultralytics/yolov5
YOLOv5 is a PyTorch‑based, open‑source object detection/segmentation/classification model suite from Ultralytics, offering multiple model sizes, simple Python/CLI inference, one‑line training, and export to many deployment formats.
yakhyo/uniface
UniFace is a Python library that unifies 15 face‑analysis tasks—detection, landmarks, 3‑D mesh, parsing, matting, gaze, head pose, demographics, emotion, quality, anti‑spoofing, recognition, anonymisation, and FAISS‑based similarity—into a single, easy‑to‑install package (CPU, Apple Silicon, or CUDA). It provides a consistent API, auto‑downloads verified model weights, and includes extensive docs, notebooks, and a community Discord.
roflcoopter/viseron
A self-hosted, local-only NVR and AI computer vision software for home and office monitoring with object and face recognition.
MrNeRF/LichtFeld-Studio
A modular workstation for 3D Gaussian Splatting that integrates training, real-time visualization, and editing into a single native application.
serengil/deepface
A lightweight Python framework for face recognition and facial attribute analysis that wraps multiple state-of-the-art models into a single, easy-to-use API.
freemocap/freemocap
A free and open-source motion capture platform that provides research-grade tracking without requiring expensive, proprietary hardware.
sparkjsdev/spark
An advanced 3D Gaussian Splatting renderer for THREE.js that enables high-performance, photorealistic 3D scenes on the web with broad device support.