baidu/Unlimited-OCR

A multimodal OCR model for one-shot long-horizon parsing of complex documents, PDFs, and multi-page images.

ruvnet/RuView

A WiFi sensing platform that turns radio signals into spatial intelligence to detect people, track vitals, and estimate pose without cameras or wearables.

immich-app/immich

A high-performance self-hosted photo and video management solution featuring AI-powered facial recognition and CLIP-based search.

PaddlePaddle/PaddleOCR

PaddleOCR is an open‑source OCR and document‑AI toolkit that converts images/PDFs into structured, LLM‑ready JSON or Markdown. It offers multilingual scene‑text models (PP‑OCRv6), a lightweight vision‑language parser (PaddleOCR‑VL‑1.6) for tables, formulas, and ancient scripts, and a structure‑aware pipeline (PP‑StructureV3). Models run on CPU, GPU, XPU, NPU or via ONNX/TensorRT, and integrate with RAG/agent platforms like Dify and RAGFlow.

Faceplugin-ltd/Open-Source-Face-Recognition-SDK

An open-source face recognition SDK for Windows and Linux that provides on-premise face detection, landmark detection, and similarity comparison using deep learning.

bhimrazy/receipt-ocr

An OCR engine that extracts either raw text or structured JSON data from receipt images using Tesseract and LLMs.

hacksider/Deep-Live-Cam

A real-time face-swapping and deepfake tool that allows users to replace faces in live webcam streams or videos using a single source image.

amap-cvlab/ABot-Recon

ABot-Recon is a streaming 3D reconstruction system that uses a fixed 12-frame local context to efficiently build global trajectories and point clouds from long video streams without needing persistent long-range memory.

ultralytics/ultralytics

A high-performance computer vision framework implementing YOLO models for real-time object detection, segmentation, classification, and pose estimation.

ant-research/4DAnyone

A tool that converts a single monocular video of a person into multi-view videos to enable high-quality 4D Gaussian Splatting reconstruction.

roryclear/clearcam

An AI-powered NVR system that adds object detection, tracking, and AI-generated notification summaries to any RTSP security camera.

MaaAssistantArknights/MaaAssistantArknights

An automation assistant for Arknights that uses image recognition and OCR to automate daily tasks, resource farming, and base management.

mkturkcan/DART

A training-free framework that converts SAM3 into a real-time multi-class open-vocabulary detector using TensorRT optimization and distilled backbones.

ruoyuw22-byte/NeuroMorph-Assessment

An automated structural MRI morphometry system that integrates CAT12 processing to measure brain tissue volumes and generate quantitative reports.

om-ai-lab/VLX-Seek

VLX‑Seek is an open‑source vision‑language model that performs fine‑grained perception for edge‑device agents. It generates region proposals, encodes each region as a special token, and lets the language model answer by selecting those tokens—producing compact, parsable outputs for detection, referring‑expression comprehension, OCR, counting, and VQA. The repo ships inference code, a Python API, and the 10 B checkpoint (weights hosted on Hugging Face).

julyx10/lap

An open-source, local-first photo manager that uses local AI for search and face recognition to provide a private alternative to cloud photo services.

roboflow/rf-detr

RF-DETR is a real-time transformer architecture for object detection, instance segmentation, and keypoint detection built on a DINOv2 backbone.

blakeblackshear/frigate

A local NVR with real-time AI object detection for IP cameras, designed for tight integration with Home Assistant.

RapidAI/RapidOCR

An open-source OCR tool that converts PaddleOCR models to ONNX format for fast, low-resource offline deployment across multiple platforms.

Robbyant/lingbot-map

A feed-forward 3D foundation model for streaming 3D reconstruction that enables real-time generation of point clouds and camera trajectories from video.

facebookresearch/vggt-omega

VGGT-Omega is a computer vision model that estimates camera poses and depth from images to reconstruct 3D scenes.

roboflow/supervision

A computer vision toolkit that provides model-agnostic building blocks for data loading, visualization, and dataset management to accelerate the development of vision applications.

tesseract-ocr/tesseract

Tesseract is an open-source OCR engine and command-line tool that converts images of text into machine-readable text across more than 100 languages.

thiagotigaz/ocr-it

A Chrome and Firefox extension that uses local OCR to extract text from paginated web documents and viewers that prevent text selection.

harry7557558/spirula-studio

A self-contained 3D Gaussian Splatting pipeline that transforms photos and videos into 3D splats and meshes without requiring Python or PyTorch.

opencv/opencv

OpenCV is an open-source computer vision library that provides tools and algorithms for AI and visual perception development.

ssrajadh/sentrysearch

A semantic video search tool that lets you find specific events in footage using natural language or images and automatically trims the matching clips.

microsoft/MoGe

MoGe is a model for recovering accurate 3D geometry, including metric depth and point maps, from single open-domain images.

ace-trump-tech/DeltaForce-OBS-Locker

A real-time object detection and target locking system for Delta Force using YOLOv14 to identify game characters while filtering out environmental false positives.

duy-phamduc68/TrafficLab-3D

An end-to-end traffic analysis suite that creates 3D digital twins from CCTV footage and satellite maps using computer vision and a custom calibration pipeline.

facebookresearch/sam3

SAM 3 (Segment Anything with Concepts) is Meta’s 848 M‑parameter foundation model that lets you prompt an image or video with free‑form text (or visual exemplars) and receive masks, boxes, and scores for *all* matching objects. It combines a DETR‑style detector and a SAM 2‑style tracker, introduces a presence token for fine‑grained prompt discrimination, and is trained on >4 M auto‑annotated concepts. The repo provides installation steps, example notebooks, and a new SA‑CO benchmark (270 K concepts) for evaluation.

ByteDance-Seed/Depth-Anything-3

A model that predicts spatially consistent 3D geometry, depth maps, and camera poses from single or multiple images using a unified depth-ray representation.

colmap/colmap

A general-purpose Structure-from-Motion and Multi-View Stereo pipeline for reconstructing 3D structures from image collections.

ocrmypdf/OCRmyPDF

A command-line tool that adds an OCR text layer to scanned PDFs, making them searchable and copy-pasteable while maintaining image quality.

BigBodyCobain/Shadowbroker

A decentralized geospatial intelligence platform that aggregates 60+ real-time OSINT feeds into a single map interface for global threat monitoring.

MaaEnd/MaaEnd

An automation assistant for the game Endfield that uses screen recognition to automate combat, resource farming, and daily management tasks.

Yuliang-Liu/MonkeyOCRv2

MonkeyOCRv2 is an open‑source visual‑text foundation model for Document AI. It ships a ViT‑based vision encoder plus multilingual parsing and understanding heads (≈0.6‑1.8 B parameters) that achieve state‑of‑the‑art scores on OCR, layout parsing, VQA and formula recognition across 17 languages. Models are available on Hugging Face/ModelScope, can be run with Transformers or vLLM (CPU support and DFlash acceleration optional), and are accompanied by the large MonkeyDoc v2 pre‑training dataset.

ZhengPeng7/BiRefNet

BiRefNet is a high-resolution dichotomous image segmentation model that provides state-of-the-art background removal and object segmentation for images of varying resolutions.

oncologylab/histopia

A computational research tool for serial-section histology and proteomic image analysis, providing 3D reconstruction, inter-section alignment, and spatial topology profiling.

deepinsight/insightface

An open-source 2D and 3D face analysis toolbox providing state-of-the-art algorithms for face detection, recognition, and alignment.

MaaXYZ/MaaFramework

A cross-platform, image-recognition-based automation framework for black-box testing and creating software assistants.

ultralytics/yolov5

YOLOv5 is a PyTorch‑based, open‑source object detection/segmentation/classification model suite from Ultralytics, offering multiple model sizes, simple Python/CLI inference, one‑line training, and export to many deployment formats.

yakhyo/uniface

UniFace is a Python library that unifies 15 face‑analysis tasks—detection, landmarks, 3‑D mesh, parsing, matting, gaze, head pose, demographics, emotion, quality, anti‑spoofing, recognition, anonymisation, and FAISS‑based similarity—into a single, easy‑to‑install package (CPU, Apple Silicon, or CUDA). It provides a consistent API, auto‑downloads verified model weights, and includes extensive docs, notebooks, and a community Discord.

roflcoopter/viseron

A self-hosted, local-only NVR and AI computer vision software for home and office monitoring with object and face recognition.

MrNeRF/LichtFeld-Studio

A modular workstation for 3D Gaussian Splatting that integrates training, real-time visualization, and editing into a single native application.

serengil/deepface

A lightweight Python framework for face recognition and facial attribute analysis that wraps multiple state-of-the-art models into a single, easy-to-use API.

freemocap/freemocap

A free and open-source motion capture platform that provides research-grade tracking without requiring expensive, proprietary hardware.

sparkjsdev/spark

An advanced 3D Gaussian Splatting renderer for THREE.js that enables high-performance, photorealistic 3D scenes on the web with broad device support.