liustack/modlens

ModLens is a plug‑in for DeepSeek Harness (and other local AI assistants) that adds vision capability to text‑only LLMs. Users paste an image directly into the chat; the plugin forwards it to a configurable vision engine (Gemini, OpenAI‑compatible, Anthropic, Antigravity CLI, Claude/Kimi CLIs, or local agents) and returns a structured transcription for the model to cite. Installation is a single `npx` command or via `skills.sh`, with zero‑config defaults and automatic key rotation/fail‑over. The project is MIT‑licensed and focuses on lightweight integration rather than heavy wrappers.

royshil/obs-backgroundremoval

An OBS Studio plugin that uses neural networks to remove portrait backgrounds and enhance low-light video in real-time without requiring a physical green screen.

lyuwenyu/RT-DETR

A real-time object detection framework based on the Detection Transformer (DETR) architecture that aims to outperform YOLO models in speed and accuracy.

MaaEnd/MaaEnd

An automation assistant for the game Endfield that uses screen recognition to automate combat, resource gathering, and daily management tasks.

microsoft/MoGe

MoGe is a model for recovering accurate 3D geometry, including metric depth and point maps, from single open-domain images.

sparkjsdev/spark

An advanced 3D Gaussian Splatting renderer for THREE.js that enables high-performance, cross-device rendering of splat-based 3D scenes.

MaaXYZ/MaaFramework

An image-recognition-based automation framework for black-box testing that allows developers to create software assistants and automation tools using a low-code pipeline protocol.

beixiaocai/rebucca

A multi-channel video analysis platform that combines YOLO detection and LLM verification to provide structured alarms for security and monitoring.

davidakpele/wifi-densepose

A privacy-preserving human pose estimation system that uses WiFi Channel State Information (CSI) and machine learning to detect poses without cameras.

nerfstudio-project/gsplat

A CUDA-accelerated library for the rasterization of 3D Gaussians, providing a faster and more memory-efficient way to render radiance fields compared to the official implementation.

aloshdenny/reverse-SynthID

A reverse-engineering project that detects and removes Google's invisible SynthID watermarks from Gemini-generated images using spectral analysis and a multi-stage adversarial pipeline.

wasserth/TotalSegmentator

An automated tool for segmenting major anatomical structures in CT and MR images, providing detailed anatomical masks and body statistics.

valkia/aramgg_client

A Windows desktop assistant for League of Legends ARAM that provides champion selection advice and uses OCR to recommend builds based on in-game Hextech Augments.

google-deepmind/alphafold3

An implementation of the AlphaFold 3 inference pipeline for the accurate 3D structure prediction of biomolecular interactions.

isl-org/Open3D

Open3D is a modern library for 3D data processing that provides optimized tools for point cloud and mesh manipulation, visualization, and 3D machine learning.

open-edge-platform/anomalib

A deep learning library for benchmarking and deploying anomaly detection algorithms, with a strong focus on visual anomaly detection in images and videos.

pynicolas/FairScan

An open-source Android document scanner that uses local machine learning for automatic document detection and perspective correction to create private PDFs.

freemocap/freemocap

A free and open-source motion capture platform that provides research-grade tracking without requiring expensive, proprietary hardware.

lightly-ai/lightly-train

A computer vision framework for pretraining, fine-tuning, and distilling SOTA models like DINOv3 and LTDETRv2 for tasks such as object detection and segmentation.

huggingface/pytorch-image-models

A comprehensive PyTorch library providing a vast collection of state-of-the-art computer vision backbones and pretrained weights for easy integration into AI projects.

WebODM/WebODM

A commercial-grade software for processing drone images into georeferenced maps, point clouds, 3D models, and Gaussian splats.

ZhengPeng7/BiRefNet

BiRefNet is a high-resolution dichotomous image segmentation model that provides state-of-the-art background removal and object segmentation for images of varying resolutions.

vllm-project/llm-compressor

LLM Compressor is a Python library that quantizes and prunes large language models into the `compressed‑tensors` format, enabling memory‑efficient deployment with vLLM. It supports many low‑precision formats (NVFP4, FP8, INT4, etc.), several PTQ/GPTQ algorithms, DDP and disk‑offloading for huge models, and ships pre‑quantized checkpoints for popular LLMs.

NVIDIA/DLSS

The NVIDIA DLSS SDK provides tools for high-performance image scaling and sharpening for games, supporting integration with DX11, DX12, and Vulkan.

opengeos/geoai

A Python package that integrates AI frameworks with geospatial data analysis, providing tools for satellite imagery processing, model training, and inference.

realsenseai/librealsense

A cross-platform SDK for RealSense depth cameras that enables depth and color streaming and provides calibration data for computer vision and robotics.

voxel51/fiftyone

An open-source tool for building high-quality computer vision datasets and models by providing visual data exploration, labeling, and model evaluation.

NVlabs/Fast-FoundationStereo

A family of real-time stereo matching architectures that achieve strong zero-shot generalization by using knowledge distillation, neural architecture search, and structured pruning.

opendatalab/OmniDocBench

A comprehensive benchmark for evaluating document parsing in real-world scenarios, featuring rich annotations for layout detection, OCR, table, and formula recognition.

apple-aiml-research/ml-sharp

SHARP is a fast monocular view synthesis tool that uses a neural network to generate a metric 3D Gaussian representation from a single image in under a second.

Pointcept/Pointcept

A powerful and flexible codebase for point cloud perception research, integrating various 3D backbones and self-supervised pre-training frameworks.

prs-eth/Marigold

A family of conditional generative models that repurpose pretrained diffusion models like Stable Diffusion for high-resolution monocular depth, surface normal, and intrinsic image analysis.

roboflow/trackers

A plug-and-play Python library providing benchmarked implementations of multiple object tracking algorithms for any detection model.

alicevision/Meshroom

A node-based visual programming framework for 3D reconstruction and computer vision, enabling users to build complex data processing pipelines using AI and photogrammetry.

roryclear/clearcam

An AI-powered NVR system that adds object detection, tracking, and AI-generated event summaries to any RTSP security camera.

Tencent/YOLO-Master

A YOLO-style framework for real-time object detection that uses Mixture-of-Experts (MoE) to adaptively allocate computation based on scene complexity for higher precision and lower latency.

vietanhdev/anylabeling

An AI-powered image annotation tool that integrates YOLOv8 and Segment Anything to provide automatic labeling for computer vision datasets.

facebookresearch/map-anything

An open-source research framework for universal metric 3D reconstruction that uses a single transformer model to handle multiple 3D tasks from various input combinations.

Breakthrough/PySceneDetect

A video cut detection and analysis tool that automatically identifies scene changes to split videos into individual shots or extract frames.

google/GNM

A family of parametric statistical 3D human models, starting with GNM Head, for high-fidelity representation of human geometry and appearance.

kornia/kornia

A differentiable computer vision library for PyTorch that integrates classical image processing and geometric vision algorithms into deep learning pipelines.

Audiveris/audiveris

An open-source Optical Music Recognition (OMR) application that transcribes music score images into symbolic digital formats for playback and editing.

datalab-to/surya

A 650M parameter OCR model for document intelligence that provides high-accuracy text recognition, layout analysis, and table extraction across 91 languages.

torinmb/mediapipe-touchdesigner

A GPU-accelerated MediaPipe plugin for TouchDesigner that enables real-time computer vision tasks like pose and hand tracking without requiring external installations.

Faceplugin-ltd/Open-Source-Face-Recognition-SDK

An open-source face recognition SDK for Windows and Linux that provides on-premise face detection, landmark detection, and similarity comparison using deep learning.

1bananachicken/MaaNTE

An automation tool for the game *異環* (The Ring) that uses image and audio recognition to automate repetitive gameplay and mini-games.

pixpark/gpupixel

A high-performance, cross-platform C++ library for image and video beauty filters using OpenGL/ES.

ucam-eo/tessera

TESSERA is an open geospatial foundation model that converts cloud-corrupted satellite time series into compact 128-dimensional embeddings for global Earth representation and analysis.