MCG-NJU/MOTIP
MOTIP is a multiple object tracking framework that treats tracking as an in-context ID prediction problem to directly decode ID labels for current detections.
xiaobiaodu/Mobile-GS
Mobile-GS is a real-time Gaussian Splatting framework designed to enable high-quality 3D scene rendering on mobile devices.
allenai/olmoearth_pretrain
A family of multi-modal, spatio-temporal foundation models for Earth Observation, providing a scalable framework for planetary intelligence using satellite data.
Tsinghua-MARS-Lab/SLAM-Former
A Transformer-based SLAM system that unifies localization and mapping into a single model for 3D reconstruction and camera tracking from image sequences.
alicevision/CCTag
A computer vision library for the detection and accurate localization of circular fiducial markers made of concentric circles.
81NewArk/AntiCAP-WebApi
A web API service for solving CAPTCHAs using model inference, supporting local and public network deployment.
bhky/opennsfw2
A Keras implementation of the Yahoo Open-NSFW model for detecting pornographic content in images and videos.
RenderKit/oidn
A high-performance library of deep learning-based denoising filters that reduces rendering times for ray tracing applications by removing Monte Carlo noise.
myBoris/wzry_ai
An open-source AI model project designed to automate gameplay for Honor of Kings using machine learning and custom screen coordinate mapping.
LuisaGroup/LuisaRender
A high-performance cross-platform Monte-Carlo renderer for stream architectures that supports multiple hardware backends like CUDA, DirectX, and Metal.
henrywoo/kazam
A Linux screen recorder and broadcaster that includes AI-powered OCR for extracting text from screen captures.
wenbowen123/BundleTrack
BundleTrack is a real-time 6D pose tracking framework for novel objects that operates without requiring 3D CAD models of the object or its category.
OpenStitching/stitching
A Python package for fast and robust image stitching that creates panoramas from overlapping images using OpenCV.
NativeSensors/EyeGestures
An open-source eye-tracking library that enables eye-driven interfaces using standard webcams and phone cameras instead of expensive specialized hardware.
facebookresearch/projectaria_tools
A suite of C++/Python utilities for processing and analyzing sensor and machine perception data from Project Aria AR glasses to support AI and machine perception research.
FaceAISDK/FaceAISDK_Android
An offline on-device face recognition SDK for Android that provides face detection, recognition, and liveness detection without requiring cloud connectivity.
facebookresearch/Ego4d
A toolkit and dataset infrastructure for Ego4D and Ego-Exo4D, providing tools to download, process, and train models on massive first-person and multi-view video datasets.
yuxumin/PoinTr
PoinTr is a transformer-based model for point cloud completion that reconstructs full 3D objects from partial point clouds using a geometry-aware encoder-decoder architecture.
MeshInspector/MeshLib
A high-performance 3D mesh processing SDK that provides tools for repairing, optimizing, and manipulating 3D data across multiple programming languages.
hbb1/2d-gaussian-splatting
A 2D Gaussian Splatting implementation that uses oriented disks (surfels) to create geometrically accurate radiance fields and high-quality 3D meshes.
nipreps/fmriprep
A robust preprocessing pipeline for fMRI data that automates motion correction, normalization, and brain extraction using a combination of state-of-the-art neuroimaging tools.
lolishinshi/imsearch
A large-scale similar image search tool that uses feature point matching to find full images using a small cropped screenshot.
cdcseacave/openMVS
OpenMVS is a multi-view stereo reconstruction library that transforms sparse point-clouds and camera poses into detailed, textured 3D meshes.
microsoft/BiomedParse
A foundation model for joint segmentation, detection, and recognition of biomedical objects across nine imaging modalities, supporting both 2D and 3D volumetric inference.
thiagotigaz/ocr-it
A Chrome and Firefox extension that uses local OCR to extract text from paginated web documents and viewers that prevent text selection.
ohirose/bcpd
BCPD is a MATLAB‑compatible, command‑line suite for Bayesian non‑rigid registration of point clouds and functions (DET), with accelerated “++” modes, multiple kernel options, and demos for 3‑D reconstruction, shape transfer, and spatial‑omics data.
facebook/ThreatExchange
A collection of tools and APIs for content moderation and security threat intelligence, featuring image and video hashing algorithms for similarity matching.
Linfeng-Tang/Image-Fusion
A comprehensive collection of research papers and code implementations for image fusion, covering multi-modal, digital photography, and remote sensing image combination techniques.
freesurfer/freesurfer
FreeSurfer is an open-source software suite for processing, analyzing, and visualizing human brain MRI data from structural and functional studies.
dipy/dipy
DIPY is a Python library designed for the analysis of MR diffusion imaging, providing researchers with tools to process medical imaging data.
NVIDIA-AI-IOT/Lidar_AI_Solution
A collection of highly optimized CUDA and TensorRT implementations for 3D Lidar perception, speeding up sparse convolutions and sensor fusion for self-driving systems.
the-database/mpv-AnimeJaNai
A custom mpv video player build and a set of Real-ESRGAN models that enable real-time 4K upscaling of anime content using TensorRT or DirectML.
naver/anny
Anny is an open‑source, PyTorch‑based differentiable human‑body mesh model covering all ages. It offers multiple rigs (compact 104‑bone default, full MakeHuman), several mesh topologies (including SMPL‑X and SOMA), and fine‑grained shape, local‑change, and facial‑action parameters. Install via pip, run a quick Python example to output a .ply mesh, or launch an interactive Gradio demo. The library is Apache‑2.0 licensed, builds on MakeHuman assets, and is intended for research/production tasks such as pose estimation, animation, and cross‑age body modeling.
arcships/light-ocr
A fast, offline OCR library for Node.js and C++ that provides local text recognition for images and PDFs with built-in hardware acceleration.
kha-white/manga-ocr
An OCR tool specifically optimized for Japanese manga, capable of recognizing multi-line text, furigana, and both vertical and horizontal layouts.
open-edge-platform/geti
An end-to-end platform for building computer vision AI models, providing tools for annotation, training, and optimization for deployment on edge hardware.
voxelmorph/voxelmorph
A learning-based image registration library that uses neural networks to align images and model deformations, featuring contrast-agnostic training via SynthMorph.
JEOresearch/EyeTracker
A lightweight Python-based 3D eye tracking algorithm that detects the pupil ellipse from video or image data.
hank-ai/darknet
An open-source neural network framework written in C/C++ and CUDA that powers the YOLO real-time object detection system.
hku-mars/GS-SDF
A LiDAR-augmented system that combines Gaussian Splatting and neural Signed Distance Fields (SDF) for geometrically consistent photorealistic rendering and 3D surface reconstruction.
stereolabs/zed-sdk
A cross-platform spatial perception SDK for ZED cameras that provides real-time depth sensing, object detection, and positional tracking.
shimat/opencvsharp
A cross-platform .NET wrapper for OpenCV that brings comprehensive computer vision and image processing functionality to C# developers.
leblancfg/autocrop
A Python-based tool that uses the YuNet neural network to automatically crop images around the largest detected face, ideal for profile pictures and ID cards.
imagej/ImageJ
ImageJ is public domain software for processing and analyzing scientific images, designed to run across multiple platforms using Java.
kha-white/mokuro
A tool that performs text detection and OCR on Japanese manga to create selectable text overlays in a browser, facilitating the use of pop-up dictionaries for language learners.
ngageoint/hootenanny
An open-source map data conflation tool that uses machine learning to merge multiple geospatial datasets into a single seamless map.
facebookresearch/pytorch3d
A PyTorch-based library for 3D Computer Vision research providing differentiable rendering and efficient tools for manipulating 3D meshes and point clouds.
luxonis/depthai-core
A C++ and Python library for interfacing with Luxonis DepthAI hardware to build spatial AI and computer vision applications.