liustack/modlens
ModLens is a plug‑in for DeepSeek Harness (and other local AI assistants) that adds vision capability to text‑only LLMs. Users paste an image directly into the chat; the plugin forwards it to a configurable vision engine (Gemini, OpenAI‑compatible, Anthropic, Antigravity CLI, Claude/Kimi CLIs, or local agents) and returns a structured transcription for the model to cite. Installation is a single `npx` command or via `skills.sh`, with zero‑config defaults and automatic key rotation/fail‑over. The project is MIT‑licensed and focuses on lightweight integration rather than heavy wrappers.
royshil/obs-backgroundremoval
An OBS Studio plugin that uses neural networks to remove portrait backgrounds and enhance low-light video in real-time without requiring a physical green screen.
lyuwenyu/RT-DETR
A real-time object detection framework based on the Detection Transformer (DETR) architecture that aims to outperform YOLO models in speed and accuracy.
MaaEnd/MaaEnd
An automation assistant for the game Endfield that uses screen recognition to automate combat, resource gathering, and daily management tasks.
microsoft/MoGe
MoGe is a model for recovering accurate 3D geometry, including metric depth and point maps, from single open-domain images.
sparkjsdev/spark
An advanced 3D Gaussian Splatting renderer for THREE.js that enables high-performance, cross-device rendering of splat-based 3D scenes.
MaaXYZ/MaaFramework
An image-recognition-based automation framework for black-box testing that allows developers to create software assistants and automation tools using a low-code pipeline protocol.
beixiaocai/rebucca
A multi-channel video analysis platform that combines YOLO detection and LLM verification to provide structured alarms for security and monitoring.
davidakpele/wifi-densepose
A privacy-preserving human pose estimation system that uses WiFi Channel State Information (CSI) and machine learning to detect poses without cameras.
nerfstudio-project/gsplat
A CUDA-accelerated library for the rasterization of 3D Gaussians, providing a faster and more memory-efficient way to render radiance fields compared to the official implementation.
aloshdenny/reverse-SynthID
A reverse-engineering project that detects and removes Google's invisible SynthID watermarks from Gemini-generated images using spectral analysis and a multi-stage adversarial pipeline.
wasserth/TotalSegmentator
An automated tool for segmenting major anatomical structures in CT and MR images, providing detailed anatomical masks and body statistics.
valkia/aramgg_client
A Windows desktop assistant for League of Legends ARAM that provides champion selection advice and uses OCR to recommend builds based on in-game Hextech Augments.
google-deepmind/alphafold3
An implementation of the AlphaFold 3 inference pipeline for the accurate 3D structure prediction of biomolecular interactions.
isl-org/Open3D
Open3D is a modern library for 3D data processing that provides optimized tools for point cloud and mesh manipulation, visualization, and 3D machine learning.
open-edge-platform/anomalib
A deep learning library for benchmarking and deploying anomaly detection algorithms, with a strong focus on visual anomaly detection in images and videos.
pynicolas/FairScan
An open-source Android document scanner that uses local machine learning for automatic document detection and perspective correction to create private PDFs.
freemocap/freemocap
A free and open-source motion capture platform that provides research-grade tracking without requiring expensive, proprietary hardware.
lightly-ai/lightly-train
A computer vision framework for pretraining, fine-tuning, and distilling SOTA models like DINOv3 and LTDETRv2 for tasks such as object detection and segmentation.
huggingface/pytorch-image-models
A comprehensive PyTorch library providing a vast collection of state-of-the-art computer vision backbones and pretrained weights for easy integration into AI projects.
WebODM/WebODM
A commercial-grade software for processing drone images into georeferenced maps, point clouds, 3D models, and Gaussian splats.
ZhengPeng7/BiRefNet
BiRefNet is a high-resolution dichotomous image segmentation model that provides state-of-the-art background removal and object segmentation for images of varying resolutions.
vllm-project/llm-compressor
LLM Compressor is a Python library that quantizes and prunes large language models into the `compressed‑tensors` format, enabling memory‑efficient deployment with vLLM. It supports many low‑precision formats (NVFP4, FP8, INT4, etc.), several PTQ/GPTQ algorithms, DDP and disk‑offloading for huge models, and ships pre‑quantized checkpoints for popular LLMs.
NVIDIA/DLSS
The NVIDIA DLSS SDK provides tools for high-performance image scaling and sharpening for games, supporting integration with DX11, DX12, and Vulkan.
opengeos/geoai
A Python package that integrates AI frameworks with geospatial data analysis, providing tools for satellite imagery processing, model training, and inference.
realsenseai/librealsense
A cross-platform SDK for RealSense depth cameras that enables depth and color streaming and provides calibration data for computer vision and robotics.
voxel51/fiftyone
An open-source tool for building high-quality computer vision datasets and models by providing visual data exploration, labeling, and model evaluation.
NVlabs/Fast-FoundationStereo
A family of real-time stereo matching architectures that achieve strong zero-shot generalization by using knowledge distillation, neural architecture search, and structured pruning.
opendatalab/OmniDocBench
A comprehensive benchmark for evaluating document parsing in real-world scenarios, featuring rich annotations for layout detection, OCR, table, and formula recognition.
apple-aiml-research/ml-sharp
SHARP is a fast monocular view synthesis tool that uses a neural network to generate a metric 3D Gaussian representation from a single image in under a second.
Pointcept/Pointcept
A powerful and flexible codebase for point cloud perception research, integrating various 3D backbones and self-supervised pre-training frameworks.
prs-eth/Marigold
A family of conditional generative models that repurpose pretrained diffusion models like Stable Diffusion for high-resolution monocular depth, surface normal, and intrinsic image analysis.
roboflow/trackers
A plug-and-play Python library providing benchmarked implementations of multiple object tracking algorithms for any detection model.
alicevision/Meshroom
A node-based visual programming framework for 3D reconstruction and computer vision, enabling users to build complex data processing pipelines using AI and photogrammetry.
roryclear/clearcam
An AI-powered NVR system that adds object detection, tracking, and AI-generated event summaries to any RTSP security camera.
Tencent/YOLO-Master
A YOLO-style framework for real-time object detection that uses Mixture-of-Experts (MoE) to adaptively allocate computation based on scene complexity for higher precision and lower latency.
vietanhdev/anylabeling
An AI-powered image annotation tool that integrates YOLOv8 and Segment Anything to provide automatic labeling for computer vision datasets.
facebookresearch/map-anything
An open-source research framework for universal metric 3D reconstruction that uses a single transformer model to handle multiple 3D tasks from various input combinations.
Breakthrough/PySceneDetect
A video cut detection and analysis tool that automatically identifies scene changes to split videos into individual shots or extract frames.
google/GNM
A family of parametric statistical 3D human models, starting with GNM Head, for high-fidelity representation of human geometry and appearance.
kornia/kornia
A differentiable computer vision library for PyTorch that integrates classical image processing and geometric vision algorithms into deep learning pipelines.
Audiveris/audiveris
An open-source Optical Music Recognition (OMR) application that transcribes music score images into symbolic digital formats for playback and editing.
datalab-to/surya
A 650M parameter OCR model for document intelligence that provides high-accuracy text recognition, layout analysis, and table extraction across 91 languages.
torinmb/mediapipe-touchdesigner
A GPU-accelerated MediaPipe plugin for TouchDesigner that enables real-time computer vision tasks like pose and hand tracking without requiring external installations.
Faceplugin-ltd/Open-Source-Face-Recognition-SDK
An open-source face recognition SDK for Windows and Linux that provides on-premise face detection, landmark detection, and similarity comparison using deep learning.
1bananachicken/MaaNTE
An automation tool for the game *異環* (The Ring) that uses image and audio recognition to automate repetitive gameplay and mini-games.
pixpark/gpupixel
A high-performance, cross-platform C++ library for image and video beauty filters using OpenGL/ES.
ucam-eo/tessera
TESSERA is an open geospatial foundation model that converts cloud-corrupted satellite time series into compact 128-dimensional embeddings for global Earth representation and analysis.