LightningRAG/LightningRAG

A full-stack RAG framework with a visual agent orchestration canvas and webhook connectors for deploying AI agents to third-party messaging platforms.

DataScienceUIBK/Rankify

A comprehensive Python toolkit for retrieval, re-ranking, and RAG, providing a unified interface to experiment with and deploy state-of-the-art models.

graniet/llm

A Rust library providing a unified API to interact with multiple LLM backends, enabling multi-step chains, agentic workflows, and a standardized REST API.

Farzad-R/LLM-Zero-to-Hundred

A collection of LLM projects and tutorials covering RAG, agentic workflows, multimodal chatbots, and fine-tuning pipelines to move from basic to advanced AI application development.

maze-agent/Maze

A distributed framework for LLM agents that transforms agent programs into observable workflows with heterogeneous resource scheduling and automated model serving.

suyoumo/ClawProBench

A transparent live-first benchmark harness for evaluating AI agent capabilities and performance within the OpenClaw runtime using deterministic grading.

OpenManus/OpenManus-RL

An open-source framework for optimizing LLM agents using reinforcement learning to enhance their reasoning, tool use, and decision-making capabilities.

mkalten/reacTIVision

A cross-platform computer vision framework for tracking fiducial markers and multi-touch fingers to enable the creation of tangible user interfaces.

frgfm/holocron

A library of high-quality PyTorch implementations of recent deep learning tricks and model architectures for computer vision tasks.

rapidsai/cucim

cuCIM is a GPU-accelerated computer vision and image processing library designed for large, multidimensional images used in biomedical, geospatial, and remote sensing applications.

commaai/rednose

A Kalman filter framework for sensor fusion, visual odometry, and SLAM that automates Jacobian computation and supports advanced filters like ESKF and MSCKF.

theICTlab/3DUNDERWORLD-SLS-GPU_CPU

A structured light scanning tool that reconstructs 3D point clouds from images of objects lit by patterned light, featuring both CPU and GPU acceleration.

alicevision/popsift

A CUDA-accelerated implementation of the SIFT algorithm for real-time image feature extraction.

thp/psmoveapi

An open-source library for Linux, macOS, and Windows that enables PC access to Sony Move Motion Controllers for 3D tracking and sensor fusion in AR/VR applications.

rgeirhos/Stylized-ImageNet

A tool for creating Stylized-ImageNet, a version of ImageNet that distorts textures to encourage CNNs to prioritize object shape over texture for better robustness.

Fabric-Project/Fabric

A visual node-based creative coding environment for rapid prototyping of interactive 3D graphics, image/video processing, and AI-enhanced visuals on macOS.

VSLAM-LAB/VSLAM-LAB

A comprehensive framework for simplifying the development, evaluation, and application of Visual SLAM systems through unified configuration and benchmarking.

askui/python-sdk

A Python SDK that enables AI agents to control desktop and mobile devices using vision-based automation instead of brittle UI selectors.

l3p-cv/lost

A flexible, web-based framework for collaborative image annotation that supports semi-automatic pipelines to speed up the labeling of machine learning datasets.

roboflow/roboflow-python

The official Python package for Roboflow, enabling developers to interact with models, datasets, and projects to build and deploy computer vision models programmatically.

genicam/harvesters

A Python library for simplifying image acquisition in computer vision applications by providing a unified interface for GenICam-compliant devices.

zju3dv/Diffuman4D

Diffuman4D is a spatio-temporal diffusion model for high-fidelity 4D consistent human view synthesis, enabling free-viewpoint rendering from sparse-view videos.

open-edge-platform/datumaro

A dataset management framework and CLI tool used to build, transform, and analyze AI datasets, specifically focusing on converting between various annotation formats.

basler/pypylon

Official Python bindings for Basler pylon C++ APIs, enabling Python applications to control Basler machine vision cameras and perform image processing tasks.

10up/classifai

A WordPress plugin that integrates cloud-based and local AI services to automate content generation, image processing, and site management tasks.

cyrusbehr/YOLOv8-TensorRT-CPP

A C++ implementation of YOLOv8 using TensorRT for high-performance GPU inference, supporting object detection, semantic segmentation, and pose estimation.

insight-platform/Savant

A high-level framework for building real-time, high-performance computer vision and video analytics pipelines on Nvidia hardware, abstracting the complexity of DeepStream.

luispedro/mahotas

A fast Python computer vision library implemented in C++ that provides over 100 image processing algorithms operating on NumPy arrays.

FORTH-ModelBasedTracker/MocapNET

A real-time 2D-to-3D human pose estimator that converts RGB video feeds into standard BVH motion capture files for 3D animation.

imagej/imagej2

ImageJ2 is a scientific imaging framework that extends the original ImageJ to support multidimensional image data and headless processing across multiple programming languages.

IliasHad/edit-mind

A local video knowledge base that indexes videos via transcription and visual analysis to enable natural language semantic search of video content.

opendatacam/opendatacam

An open-source computer vision tool that detects, tracks, and counts moving objects in video feeds to quantify real-world movement, commonly used for traffic studies.

MRPT/mrpt

A robust, modular C++ framework for mobile robotics providing essential tools for SLAM, navigation, mapping, and sensor integration.

ARM-software/ComputeLibrary

A collection of low-level machine learning functions optimized for Arm Cortex-A, Neoverse, and Mali GPU architectures to provide high-performance ML inference.

shimat/opencvsharp

A cross-platform .NET wrapper for OpenCV that brings comprehensive computer vision and image processing functionality to C# developers.

deepgram/deepgram-python-sdk

The official Python SDK for Deepgram, providing easy integration of automated speech recognition, text-to-speech, and language understanding APIs.

vargHQ/sdk

An open-source TypeScript SDK that allows developers and AI agents to create AI-generated videos using a declarative JSX syntax and a unified API for multiple AI providers.

sdv-dev/SDGym

A benchmarking framework for modeling and generating synthetic data, allowing users to compare performance, memory usage, and quality across different ML techniques.

digital-go-jp/genai-ai-api

A collection of generative AI microservices and cloud templates developed by the Digital Agency of Japan to help government officials deploy task-specific AI applications.

inseq-team/inseq

A PyTorch-based toolkit for post-hoc interpretability of sequence generation models, providing various feature attribution methods to explain model outputs.

aws-samples/bedrock-engineer

An autonomous software development agent app powered by Amazon Bedrock that can create files, execute commands, and manage projects through a customizable agent interface.

aws-samples/well-architected-iac-analyzer

An AI-powered tool that analyzes Infrastructure as Code (IaC) and architecture diagrams against AWS Well-Architected best practices to provide prioritized remediation suggestions.

eidolon-ai/eidolon

An open-source SDK for building and deploying AI agents as modular, scalable services with built-in HTTP servers and dynamic agent-to-agent communication.

GoogleCloudPlatform/genai-for-marketing

A Google Cloud-based solution that automates marketing content generation, trend analysis, and data querying using Vertex AI and Google Workspace integration.

digital-go-jp/genai-web

A generative AI utilization platform developed by the Digital Agency of Japan that allows government officials to quickly and safely deploy specialized AI applications.

nv-tlabs/XCube

XCube is a generative model for high-resolution sparse 3D voxel grids that uses hierarchical latent diffusion and VDB data structures to create detailed 3D objects and large-scale scenes.

alexiglad/EBT

A framework for Energy-Based Transformers (EBTs) that enables scalable reasoning and System 2 Thinking across text, image, and video modalities.

FoundationVision/Liquid

Liquid is a unified autoregressive multimodal generator that integrates visual comprehension and image generation into a single LLM without needing external visual embeddings.