Ma-Zhuang/OmniNWM
OmniNWM is a research codebase that builds a panoramic, multi‑modal world model for autonomous‑driving simulation. It jointly generates RGB, semantic, depth, and 3‑D occupancy maps from a vehicle trajectory, uses normalized Plücker ray‑maps for precise control, and provides occupancy‑based dense rewards for closed‑loop policy evaluation. The repo includes installation steps (including a manual patch to `transformers`), data preparation instructions for nuScenes, pretrained checkpoint download links, and scripts for inference, out‑of‑distribution testing, and staged training.
imagej/imagej2
ImageJ2 is a scientific imaging framework that extends the original ImageJ to support multidimensional image data and headless processing across multiple programming languages.
unrealcv/unrealcv
A plugin for Unreal Engine that enables computer vision researchers to build virtual worlds and interact with them using external AI frameworks like PyTorch and TensorFlow.
microsoft/Cognitive-Samples-VideoFrameAnalysis
A C# library and sample applications for analyzing webcam video frames in near-real-time using Microsoft Cognitive Services Vision APIs.
zhanghang1989/PyTorch-Encoding
A PyTorch library providing advanced encoding modules and a model zoo for image classification and semantic segmentation.
apple-aiml-research/ml-aim
A family of large-scale autoregressive vision encoders that provide high-performance backbones for multimodal understanding and image recognition.
hzxie/GaussianCity
GaussianCity is a generative 3D city synthesis tool that uses Gaussian Splatting to create unbounded, large-scale urban environments.
HarborYuan/ovsam
An open-vocabulary extension of the Segment Anything Model (SAM) that enables simultaneous interactive segmentation and recognition of thousands of object classes.
hzxie/CityDreamer
CityDreamer is a research‑grade PyTorch framework that generates unlimited, photorealistic 3‑D cityscapes by combining a layout generator, a background‑stuff generator, and a building instance generator. It includes code for training on large OSM/Google‑Earth datasets, pre‑trained checkpoints, a web demo, and a CLI for rendering videos.
cvg/DeepLSD
DeepLSD is a high-precision line segment detector that combines deep learning with image gradients to extract and refine line segments from real-world images.
VCIP-RGBD/DFormer
DFormer is a research codebase for RGB‑D semantic segmentation, implementing DFormer, DFormerv2 (geometry‑guided attention) and DFormer++ (efficient high‑accuracy variants). It includes pre‑training scripts, dataset helpers, training/evaluation pipelines, and pretrained weights for NYU‑Depth v2 and SUN‑RGBD. The repo is a genuine AI/ML project.
ria-com/nomeroff-net
An open-source Python framework for automatic number plate recognition using YOLOv8 and RNN-based OCR to detect and read license plates across multiple countries.
google-ar/arcore-unity-extensions
A set of extensions for Unity's AR Foundation that provides native API access to Google ARCore's augmented reality features.
QuinnDamerell/OctoPrint-OctoEverywhere
A remote access and monitoring service for 3D printers that includes AI-powered print failure detection to automatically stop failed prints.
half-potato/radiance_meshes
A training framework for Radiance Meshes that enables volumetric reconstruction of 3D scenes from images, with options for optimizing models for mobile web viewing.
tensorflow/graphics
A library of differentiable graphics and geometry layers for TensorFlow that enables self-supervised training of 3D vision models through analysis by synthesis.
HasnainRaz/Fast-SRGAN
A real-time super-resolution implementation based on SR-GAN that uses pixel shuffle to efficiently upscale low-resolution videos to high-resolution.
levyflux/AViD
AViD is a framework for fine-tuning Grounding DINO using LoRA and EMA, enabling efficient adaptation of open-vocabulary object detectors to custom datasets.
lartpang/SAMs-CDConcepts-Eval
A comprehensive evaluation framework for testing SAM and SAM 2's ability to segment context-dependent concepts like medical lesions and industrial defects across images and videos.
Luo-Yihao/FaithC
A near-lossless 3D voxel representation system that replaces SDF and Marching Cubes with Faithful Contour Tokens (FCTs) to preserve sharp edges and internal structures.
google-ar/arcore-ios-sdk
A software development kit that brings Google's cross-platform ARCore features, including Cloud Anchors and Geospatial APIs, to iOS applications.
arthurflor23/handwritten-text-recognition
A TensorFlow-based framework for handwritten text recognition and synthesis, featuring extensive data augmentation and support for multiple standard HTR datasets.
openMVG/openMVG
A C++ framework for 3D reconstruction from images (photogrammetry) that provides libraries and pipelines for solving the Structure from Motion problem.
argoverse/argoverse-api
A Python API for interacting with the Argoverse dataset, providing tools to load HD maps, 3D tracking data, and motion forecasting sequences for autonomous driving research.
ShenhanQian/VHAP
VHAP is a photometric optimization pipeline for human head alignment that uses adaptive appearance priors to track regions without landmarks, such as hair and ears, in monocular and multi-view videos.
lartpang/PySODMetrics
A lightweight Python library for calculating Salient Object Detection (SOD) metrics, providing a fast and verified alternative to Matlab-based evaluation toolboxes.
naruya/gaussian-vrm
A three.js implementation of Instant Skinned Gaussian Avatars that allows realistic, animatable 3D characters to be rendered in web, mobile, and VR applications.
worldcoin/open-iris
An advanced iris recognition pipeline for secure biometric verification and large-scale uniqueness confirmation using computer vision and machine learning.
MrGiovanni/SyntheticTumors
A framework for generating realistic synthetic tumors to train AI segmentation models, reducing the dependency on manually labeled medical imaging data.
MrGiovanni/AbdomenAtlas
A large-scale multi-organ CT dataset and toolset for efficient abdominal segmentation, reducing annotation time from decades to weeks using human-in-the-loop AI.
foo123/FILTER.js
A pure JavaScript library for image and video processing and computer vision that supports hardware acceleration via WebGL, WebAssembly, and Web Workers.
Unity-Technologies/arfoundation-samples
A collection of official Unity sample scenes and code demonstrating how to implement multi-platform augmented reality features using AR Foundation.
pyushkevich/itksnap
ITK-SNAP is an open-source software application for user-guided 3D active contour segmentation of anatomical structures in medical images.
tudelft-iv/view-of-delft-dataset
A multi-sensor automotive dataset containing synchronized LiDAR, camera, and 3+1D radar data with 3D bounding box annotations for road user detection in urban traffic.
dstndstn/astrometry.net
Astrometry.net is a tool for the automatic recognition of astronomical images, providing celestial coordinates and calibration metadata for images with unknown sky locations.
MouseLand/suite2p
A processing pipeline for two-photon calcium imaging data that extracts large-scale neural activity through registration, ROI detection, and spike detection.
ducha-aiki/pydegensac
A Python wrapper for LO-RANSAC and DEGENSAC used to estimate homography and fundamental matrices from sparse image correspondences.
mapillary/OpenSfM
A Structure from Motion library in Python that reconstructs 3D scenes and camera poses from multiple images, integrating sensor data for geographical alignment.
bcmi/Image-Harmonization-Dataset-iHarmony4
A large-scale image harmonization dataset (iHarmony4) and a PyTorch implementation of DoveNet to help models make inserted objects look natural in composite images.
Pointcept/Concerto
A joint 2D-3D self-supervised pre-trained Point Transformer V3 model for extracting high-quality spatial representations from 3D point clouds.
hanna-xu/U2Fusion
A unified unsupervised image fusion network that merges multi-modal, multi-exposure, and multi-focus images into a single high-quality image without requiring labeled training data.
IvanDrokin/torch-conv-kan
TorchConv‑KAN is a PyTorch library that implements convolutional layers where each kernel is a set of learnable univariate functions (KAN, Fast‑KAN, Cheby‑KAN, etc.). It ships with many layer variants, CNN‑style model families (ResKANet, DenseKANet, VGG‑KAN, U‑Net‑KAN), pre‑trained Imagenet‑1k checkpoints, and training scripts that use Accelerate, Hydra, WandB, and Ray‑Tune. The project is under active development and accompanies a research paper on Kolmogorov‑Arnold convolutions.
MIC-DKFZ/nnDetection
A self-configuring framework for medical object detection that automates the configuration process to localize and categorize objects in medical images.
Mark12Ding/SAM2Long
SAM2Long is a training-free enhancement for SAM 2 that uses a memory tree to prevent error accumulation and improve object segmentation in long videos.
Altaheri/EEG-ATCNet
A TensorFlow implementation of the Attention Temporal Convolutional Network (ATCNet) for classifying motor imagery EEG signals to improve brain-computer interfaces.
Sompote/DINOV3-YOLOV12
A hybrid object detection framework combining YOLOv12 and DINOv3 to provide superior accuracy and faster convergence, especially on small or complex datasets.
schappim/macOCR
A command-line OCR tool for macOS that extracts text, QR codes, and barcodes from the screen, images, or PDFs using Apple's Vision framework.
EyeTrackVR/EyeTrackVR
An open-source, affordable VR eye tracking platform that enables eye tracking for VRChat via OSC and UDP protocols.