ruvnet/RuView
A WiFi sensing platform that turns radio signals into spatial intelligence for camera-free presence, vital signs, and pose estimation.
roboflow/supervision
A computer vision toolkit that provides model-agnostic building blocks for data loading, visualization, and dataset management to accelerate the development of vision applications.
immich-app/immich
A high-performance self-hosted photo and video management solution featuring AI-powered facial recognition and CLIP-based search.
img2threejs/img2threejs
img2threejs is a Python‑backed skill that lets an LLM turn a single reference image into a procedural Three.js model. It builds a spec, runs a gated vision‑review loop, and emits TypeScript code that recreates the object from primitives and shaders, ready for animation. The system is extensible via domain plugins (e.g., CS2 weapon skins) and includes a live demo gallery.
baidu/Unlimited-OCR
A multimodal OCR model for one-shot long-horizon parsing of complex documents, PDFs, and multi-page images.
PaddlePaddle/PaddleOCR
PaddleOCR is an open‑source OCR and document‑AI toolkit that converts images or PDFs into structured, LLM‑ready text, tables and markdown/JSON. It offers multilingual scene‑text models (PP‑OCRv6) and lightweight vision‑language models (PaddleOCR‑VL, PP‑StructureV3) with state‑of‑the‑art accuracy, fast CPU/GPU inference, and deployment options ranging from Docker to browser SDKs. Integrated with RAG platforms like Dify and RAGFlow, it serves as a data‑engine for building AI agents and retrieval‑augmented generation pipelines.
om-ai-lab/VLX-Seek
VLX‑Seek is an open‑source vision‑language model that performs fine‑grained perception for edge‑device agents. It generates region proposals, encodes each region as a special token, and lets the language model answer by selecting those tokens—producing compact, parsable outputs for detection, referring‑expression comprehension, OCR, counting, and VQA. The repo ships inference code, a Python API, and the 10 B checkpoint (weights hosted on Hugging Face).
huawei-bayerlab/marigold-v2
Marigold V2 is a research codebase that repurposes a pretrained diffusion transformer (Qwen‑Image‑Edit‑2509) into fast, single‑step predictors for monocular depth, see‑through depth, surface normals, and albedo. It offers cheap fine‑tuning (≈5 days on a single 32 GB GPU), state‑of‑the‑art benchmark results, ready‑to‑use inference/evaluation scripts, and a modular YAML‑driven training framework that can be extended to new dense‑vision tasks.
playcanvas/supersplat
A browser-based open source tool for inspecting, editing, and optimizing 3D Gaussian Splats.
ultralytics/ultralytics
A comprehensive computer vision framework featuring YOLO models for real-time object detection, segmentation, classification, and pose estimation.
blakeblackshear/frigate
A local NVR with real-time AI object detection for IP cameras, designed for tight integration with Home Assistant.
julyx10/lap
A private, local-first photo manager for macOS, Windows, and Linux that uses local AI for search and face clustering without requiring cloud uploads.
viettranx/3dviz-pro-max
3Dviz Pro Max is an AI‑agent skill that lets large language models turn a natural‑language description into a complete Three.js scene. It provides a ten‑step workflow, a catalog of 223 recipes, 22 proven Three.js kit blueprints, style defaults, and Python helpers that render the scene in headless Chromium, capture PNG frames, and feed the observations back to the model for refinement. Install by linking the `skills/3dviz-pro-max/` folder into Claude Code or Codex, then run the bundled examples (37 runnable demos) with Vite. The project is MIT‑licensed and focuses on reproducible, evidence‑backed 3D visualisation rather than claiming that every recipe is fully rendered.
wilsjo2/OptiScaler-DLSSNR-PreSR-Multipass
A game mod that uses NVIDIA AI to enhance lighting, detail, and color in supported games, allowing for customizable neural rendering effects.
ooolabdev/ooosplat
A local desktop application that converts surround-view videos or image sequences into 3D Gaussian Splatting models through an automated local pipeline.
amap-cvlab/ABot-Recon
ABot-Recon is a streaming 3D reconstruction system that uses a fixed 12-frame local context to build global trajectories and point clouds from long video streams without needing persistent long-range memory.
MaaAssistantArknights/MaaAssistantArknights
An automation assistant for Arknights that uses image recognition and OCR to automate daily tasks, resource farming, and base management.
gazijarin/itsgiving
A real-time webcam application that detects facial expressions and hand gestures to trigger matching meme overlays, featuring a virtual camera for use in video calls.
facebookresearch/vggt-omega
VGGT-Omega is a computer vision model that predicts camera poses and depth from images to enable 3D scene reconstruction.
ika-rwth-aachen/ros2-depth-anything-v3-trt
A ROS 2 node for real-time metric depth estimation and point cloud generation using Depth Anything V3 and TensorRT acceleration.
tesseract-ocr/tesseract
Tesseract is an open-source OCR engine and command-line tool that converts images of text into machine-readable text across more than 100 languages.
StarTrail-org/PixelRAG
PixelRAG renders web pages, PDFs, and images into screenshot tiles, embeds them with a vision‑language model, builds a FAISS/Qdrant vector index, and provides a searchable API (or Claude plugin) so LLMs can retrieve information based on visual layout rather than plain text.
harry7557558/spirula-studio
A self-contained 3D Gaussian Splatting pipeline that converts photos and videos into 3D models without requiring Python, PyTorch, or COLMAP.
hacksider/Deep-Live-Cam
A real-time face-swapping and deepfake tool that allows users to replace faces in videos or live webcam feeds using a single source image.
agamrossen/VolAnti
An open-source acoustic drone detection system that identifies multirotor drones by the harmonic sound of their propellers rather than radio signals.
ace-trump-tech/DeltaForce-OBS-Locker
A real-time object detection and target locking system for Delta Force using YOLOv14 to identify game characters while filtering out environmental false positives.
ocrmypdf/OCRmyPDF
A command-line tool that adds an OCR text layer to scanned PDFs, making them searchable and copy-pasteable while maintaining image quality.
colmap/colmap
COLMAP is an open‑source Structure‑from‑Motion and Multi‑View Stereo pipeline that automatically turns collections of photos into 3‑D models. It offers a GUI, command‑line tools, Python bindings, GPU‑accelerated feature extraction, and cross‑platform binaries, making it a go‑to foundation for research and applications in computer vision, robotics, AR/VR, and any workflow that needs camera poses and dense scene geometry.
Robbyant/lingbot-map
LingBot‑Map is a feed‑forward 3‑D reconstruction model that streams video frames through a Geometric Context Transformer, using a paged KV‑cache (FlashInfer) to keep memory low. It runs at ~20 fps on modest GPUs, supports very long sequences via keyframe‑interval or windowed inference, and ships with ready‑to‑run demos, an offline batch renderer, and pretrained checkpoints on HuggingFace/ModelScope.
opencv/opencv
OpenCV is an open-source computer vision library that provides tools and algorithms for AI and visual perception development.
RapidAI/RapidOCR
An open-source OCR tool that converts PaddleOCR models to ONNX format for fast, low-resource offline deployment across multiple platforms and languages.
roboflow/rf-detr
RF‑DETR is a real‑time transformer‑based vision model suite (detection, segmentation, keypoint) built on a DINOv2 backbone. It ships as the `rfdetr` Python package with several size variants (Nano‑2XL), provides Apache‑2.0 (core) and PML 1.0 (XL/2XL) licenses, and includes benchmark tables showing state‑of‑the‑art accuracy‑latency trade‑offs on COCO and RF100‑VL. Installation is via `pip install rfdetr`; usage is a single‑line `model.predict(...)` followed by optional visualisation with the `supervision` library. The project also offers a NAS pipeline on the Roboflow platform for custom architecture search.
facebookresearch/sam3
SAM 3 (Segment Anything with Concepts) is Meta’s 848 M‑parameter foundation model that lets you prompt an image or video with free‑form text (or visual exemplars) and receive masks, boxes, and scores for *all* matching objects. It combines a DETR‑style detector and a SAM 2‑style tracker, introduces a presence token for fine‑grained prompt discrimination, and is trained on >4 M auto‑annotated concepts. The repo provides installation steps, example notebooks, and a new SA‑CO benchmark (270 K concepts) for evaluation.
cvat-ai/cvat
An open-source data annotation platform for building high-quality visual datasets for computer vision, supporting image, video, and 3D annotation.
aoguai/LiYing
An automated ID photo processing tool that uses AI to handle face detection, background replacement, and layout arrangement for professional-grade ID photos.
jeremyipark/vision-demos
vision‑demos is a set of four concrete Python demos that apply state‑of‑the‑art pose‑estimation (`vitpose-plus-large`) and segmentation (`sam3.1`) models to everyday activities—dance synchronization, chin‑up counting, bouldering route analysis, and running cadence measurement. Each demo includes a README, setup instructions, and visual examples, and the whole repo is released under Apache‑2.0.
deepinsight/insightface
An open-source 2D and 3D deep face analysis toolbox providing state-of-the-art algorithms for face recognition, detection, and alignment.
ultralytics/yolov5
YOLOv5 is a PyTorch‑based, open‑source object detection/segmentation/classification model suite from Ultralytics, offering multiple model sizes, simple Python/CLI inference, one‑line training, and export to many deployment formats.
yakhyo/uniface
UniFace is a Python library that unifies 15 face‑analysis tasks—detection, landmarks, 3‑D mesh, parsing, matting, gaze, head pose, demographics, emotion, quality, anti‑spoofing, recognition, anonymisation, and FAISS‑based similarity—into a single, easy‑to‑install package (CPU, Apple Silicon, or CUDA). It provides a consistent API, auto‑downloads verified model weights, and includes extensive docs, notebooks, and a community Discord.
oncologylab/histopia
A computational research tool for serial-section histology and proteomic image analysis, enabling 3D reconstruction and spatial topology profiling across tissue sections.
MrNeRF/LichtFeld-Studio
A modular workstation for 3D Gaussian Splatting that integrates training, real-time visualization, editing, and export in a single native application.
jayin92/Skyfall-GS
Skyfall-GS is a framework that synthesizes large-scale, immersive 3D urban scenes by combining satellite imagery for geometry and diffusion models for high-quality textures.
ArthurBrussee/brush
A cross-platform 3D reconstruction engine using Gaussian splatting that runs on WebGPU and the Burn framework to support macOS, Windows, Linux, Android, and browsers.
serengil/deepface
A lightweight Python framework for face recognition and facial attribute analysis that wraps multiple state-of-the-art models for identity verification and demographic prediction.
pq-yang/MatAnyone2
A human video matting framework that preserves fine details and avoids coarse boundaries to extract people from videos with high robustness in real-world conditions.
facebookresearch/boxer
Boxer is a system that lifts open-world 2D object detections into 3D oriented bounding boxes for indoor scenes, utilizing posed images and semi-dense point clouds.
ByteDance-Seed/Depth-Anything-3
A model that predicts spatially consistent 3D geometry, depth maps, and camera poses from single or multiple images using a unified depth-ray representation.
microsoft/OmniParser
A screen parsing tool that converts GUI screenshots into structured elements, enabling vision-based AI agents to accurately identify and interact with interface components.