LightningRAG/LightningRAG
A full-stack RAG framework with a visual agent orchestration canvas and webhook connectors for deploying AI agents to third-party messaging platforms.
DataScienceUIBK/Rankify
A comprehensive Python toolkit for retrieval, re-ranking, and RAG, providing a unified interface to experiment with and deploy state-of-the-art models.
graniet/llm
A Rust library providing a unified API to interact with multiple LLM backends, enabling multi-step chains, agentic workflows, and a standardized REST API.
Farzad-R/LLM-Zero-to-Hundred
A collection of LLM projects and tutorials covering RAG, agentic workflows, multimodal chatbots, and fine-tuning pipelines to move from basic to advanced AI application development.
maze-agent/Maze
A distributed framework for LLM agents that transforms agent programs into observable workflows with heterogeneous resource scheduling and automated model serving.
suyoumo/ClawProBench
A transparent live-first benchmark harness for evaluating AI agent capabilities and performance within the OpenClaw runtime using deterministic grading.
OpenManus/OpenManus-RL
An open-source framework for optimizing LLM agents using reinforcement learning to enhance their reasoning, tool use, and decision-making capabilities.
mkalten/reacTIVision
A cross-platform computer vision framework for tracking fiducial markers and multi-touch fingers to enable the creation of tangible user interfaces.
frgfm/holocron
A library of high-quality PyTorch implementations of recent deep learning tricks and model architectures for computer vision tasks.
rapidsai/cucim
cuCIM is a GPU-accelerated computer vision and image processing library designed for large, multidimensional images used in biomedical, geospatial, and remote sensing applications.
commaai/rednose
A Kalman filter framework for sensor fusion, visual odometry, and SLAM that automates Jacobian computation and supports advanced filters like ESKF and MSCKF.
theICTlab/3DUNDERWORLD-SLS-GPU_CPU
A structured light scanning tool that reconstructs 3D point clouds from images of objects lit by patterned light, featuring both CPU and GPU acceleration.
alicevision/popsift
A CUDA-accelerated implementation of the SIFT algorithm for real-time image feature extraction.
thp/psmoveapi
An open-source library for Linux, macOS, and Windows that enables PC access to Sony Move Motion Controllers for 3D tracking and sensor fusion in AR/VR applications.
rgeirhos/Stylized-ImageNet
A tool for creating Stylized-ImageNet, a version of ImageNet that distorts textures to encourage CNNs to prioritize object shape over texture for better robustness.
Fabric-Project/Fabric
A visual node-based creative coding environment for rapid prototyping of interactive 3D graphics, image/video processing, and AI-enhanced visuals on macOS.
VSLAM-LAB/VSLAM-LAB
A comprehensive framework for simplifying the development, evaluation, and application of Visual SLAM systems through unified configuration and benchmarking.
askui/python-sdk
A Python SDK that enables AI agents to control desktop and mobile devices using vision-based automation instead of brittle UI selectors.
l3p-cv/lost
A flexible, web-based framework for collaborative image annotation that supports semi-automatic pipelines to speed up the labeling of machine learning datasets.
roboflow/roboflow-python
The official Python package for Roboflow, enabling developers to interact with models, datasets, and projects to build and deploy computer vision models programmatically.
genicam/harvesters
A Python library for simplifying image acquisition in computer vision applications by providing a unified interface for GenICam-compliant devices.
zju3dv/Diffuman4D
Diffuman4D is a spatio-temporal diffusion model for high-fidelity 4D consistent human view synthesis, enabling free-viewpoint rendering from sparse-view videos.
open-edge-platform/datumaro
A dataset management framework and CLI tool used to build, transform, and analyze AI datasets, specifically focusing on converting between various annotation formats.
basler/pypylon
Official Python bindings for Basler pylon C++ APIs, enabling Python applications to control Basler machine vision cameras and perform image processing tasks.
10up/classifai
A WordPress plugin that integrates cloud-based and local AI services to automate content generation, image processing, and site management tasks.
cyrusbehr/YOLOv8-TensorRT-CPP
A C++ implementation of YOLOv8 using TensorRT for high-performance GPU inference, supporting object detection, semantic segmentation, and pose estimation.
insight-platform/Savant
A high-level framework for building real-time, high-performance computer vision and video analytics pipelines on Nvidia hardware, abstracting the complexity of DeepStream.
luispedro/mahotas
A fast Python computer vision library implemented in C++ that provides over 100 image processing algorithms operating on NumPy arrays.
FORTH-ModelBasedTracker/MocapNET
A real-time 2D-to-3D human pose estimator that converts RGB video feeds into standard BVH motion capture files for 3D animation.
imagej/imagej2
ImageJ2 is a scientific imaging framework that extends the original ImageJ to support multidimensional image data and headless processing across multiple programming languages.
IliasHad/edit-mind
A local video knowledge base that indexes videos via transcription and visual analysis to enable natural language semantic search of video content.
opendatacam/opendatacam
An open-source computer vision tool that detects, tracks, and counts moving objects in video feeds to quantify real-world movement, commonly used for traffic studies.
MRPT/mrpt
A robust, modular C++ framework for mobile robotics providing essential tools for SLAM, navigation, mapping, and sensor integration.
ARM-software/ComputeLibrary
A collection of low-level machine learning functions optimized for Arm Cortex-A, Neoverse, and Mali GPU architectures to provide high-performance ML inference.
shimat/opencvsharp
A cross-platform .NET wrapper for OpenCV that brings comprehensive computer vision and image processing functionality to C# developers.
deepgram/deepgram-python-sdk
The official Python SDK for Deepgram, providing easy integration of automated speech recognition, text-to-speech, and language understanding APIs.
vargHQ/sdk
An open-source TypeScript SDK that allows developers and AI agents to create AI-generated videos using a declarative JSX syntax and a unified API for multiple AI providers.
sdv-dev/SDGym
A benchmarking framework for modeling and generating synthetic data, allowing users to compare performance, memory usage, and quality across different ML techniques.
digital-go-jp/genai-ai-api
A collection of generative AI microservices and cloud templates developed by the Digital Agency of Japan to help government officials deploy task-specific AI applications.
inseq-team/inseq
A PyTorch-based toolkit for post-hoc interpretability of sequence generation models, providing various feature attribution methods to explain model outputs.
aws-samples/bedrock-engineer
An autonomous software development agent app powered by Amazon Bedrock that can create files, execute commands, and manage projects through a customizable agent interface.
aws-samples/well-architected-iac-analyzer
An AI-powered tool that analyzes Infrastructure as Code (IaC) and architecture diagrams against AWS Well-Architected best practices to provide prioritized remediation suggestions.
eidolon-ai/eidolon
An open-source SDK for building and deploying AI agents as modular, scalable services with built-in HTTP servers and dynamic agent-to-agent communication.
GoogleCloudPlatform/genai-for-marketing
A Google Cloud-based solution that automates marketing content generation, trend analysis, and data querying using Vertex AI and Google Workspace integration.
digital-go-jp/genai-web
A generative AI utilization platform developed by the Digital Agency of Japan that allows government officials to quickly and safely deploy specialized AI applications.
nv-tlabs/XCube
XCube is a generative model for high-resolution sparse 3D voxel grids that uses hierarchical latent diffusion and VDB data structures to create detailed 3D objects and large-scale scenes.
alexiglad/EBT
A framework for Energy-Based Transformers (EBTs) that enables scalable reasoning and System 2 Thinking across text, image, and video modalities.
FoundationVision/Liquid
Liquid is a unified autoregressive multimodal generator that integrates visual comprehension and image generation into a single LLM without needing external visual embeddings.