flagos-ai/FlagAttention

A collection of memory-efficient attention operators implemented in Triton, allowing for easier customization of attention mechanisms compared to CUDA-based implementations.

redis/agent-memory-server

A managed memory layer for AI agents that provides persistent storage for facts and preferences across sessions using a two-tier session and long-term memory model.

guoqingbao/xinfer

A blazing-fast LLM inference engine written in pure Rust that eliminates Python/PyTorch dependencies to enable high-throughput, production-ready model serving on consumer and professional GPUs.

dagger/container-use

An MCP server that provides isolated, containerized environments for coding agents, allowing multiple agents to work in parallel safely and independently.

BorisPolonsky/dify-helm

Helm chart that deploys the open‑source Dify LLM chatbot (API, web UI, workers, sandbox, plugins, vector DB, storage, etc.) on Kubernetes, with configurable external databases, caches, and cloud object stores.

nicobailon/surf-cli

A CLI tool that allows AI agents to control Chrome via simple shell commands, enabling browser automation without complex configuration or API keys.

rryam/VecturaKit

VecturaKit is a Swift library that provides an on‑device vector database with hybrid vector + BM25 search. It supports multiple embedding back‑ends (Apple NaturalLanguage, OpenAI‑compatible APIs, swift‑embeddings models, and Apple MLX), custom storage/search plug‑ins, and CLI tools, enabling privacy‑first semantic search in iOS/macOS/watchOS/tvOS/visionOS apps.

ome-projects/ome

A Kubernetes operator for enterprise-grade LLM serving that automates model management, runtime selection, and GPU resource optimization.

iflytek/astron-rpa

An enterprise-grade open-source RPA desktop application that enables low-code automation of desktop software and web pages, featuring deep integration with AI agents.

mir-group/nequip

An open-source framework for building E(3)-equivariant interatomic potentials using graph neural networks to accurately predict atomic energies and forces.

tdrussell/diffusion-pipe

A pipeline parallel training script that allows for the training of large-scale image and video diffusion models across multiple GPUs.

HUST-AI-HYZ/MemoryAgentBench

A benchmark for evaluating the memory capabilities of LLM agents through incremental multi-turn interactions, focusing on retrieval, learning, and conflict resolution.

hailo-ai/hailo_model_zoo

A collection of pre-trained deep learning models and tools to optimize, compile, and deploy them onto Hailo AI hardware.

Tavris1/AI-Toolkit-Easy-Install

A one-click portable installer for AI-Toolkit by Ostris that simplifies the setup of image and video diffusion model training on Windows NVIDIA GPUs.

OpenCSGs/csghub

An open-source, on-premise alternative to Hugging Face for managing, storing, and distributing LLM assets, datasets, and code.

sktime/pytorch-forecasting

A PyTorch-based library for time series forecasting that provides a high-level API and state-of-the-art deep learning architectures like Temporal Fusion Transformers and N-BEATS.

alexlenail/NN-SVG

A parametric tool for creating publication-ready neural network architecture diagrams in SVG format, eliminating the need to draw them manually.

SmythOS/sre

SmythOS SRE is an open-source runtime and SDK that acts as an operating system for AI agents, providing unified abstractions for LLMs, vector databases, and storage to simplify production deployment.

dais-polymtl/flock

A DuckDB extension that integrates LLMs and RAG pipelines into SQL, enabling multimodal semantic analysis and analytics directly within the database.

agent-network-protocol/anp

A multi-language SDK for the Agent Network Protocol (ANP) that enables AI agents to discover, authenticate, and communicate with each other using standardized interfaces and end-to-end encryption.

Project-HAMi/HAMi-core

An in-container GPU resource controller that intercepts CUDA calls to enforce per-container memory and compute utilization limits without modifying applications or drivers.

dynamiqs/dynamiqs

A GPU-accelerated and differentiable Python library for high-performance simulation of quantum systems, designed for large-scale problems and gradient-based parameter estimation.

awslabs/awsome-distributed-ai

A collection of reference architectures and examples for deploying and operating distributed AI training and inference on AWS infrastructure.

tensorflow/gnn

A TensorFlow library for building and scaling Graph Neural Networks, featuring tools for heterogeneous graph representation and distributed graph sampling.

LuisaGroup/LuisaCompute

A high-performance cross-platform computing framework that allows developers to write computation kernels in C++ and deploy them across CUDA, DirectX, Metal, and CPU backends.

realfishsam/agent-notch

A macOS utility that displays visual status indicators for Claude Code and Codex agents next to the MacBook notch, monitoring activity via system processes and transcript files.

ibrahimqureshae/mdflux

A local-first desktop application that converts documents, scanned PDFs, and audio into clean, structured Markdown to reduce LLM token costs and maintain privacy.

ww-w-ai/bkit-claude-code

A Claude Code plugin that turns AI-native development into a structured system with context-budgeted sprints, multi-agent orchestration, and automated quality verification.

modelscope/easydistill

A config-driven knowledge distillation toolkit that turns black-box teacher models into high-quality SFT and DPO training data across text, image, and video modalities.

hkr04/cpp-mcp

A C++ implementation of the Model Context Protocol (MCP) that allows AI models and agents to interact with tools, resources, and services through a standardized interface.

postmanlabs/postman-mcp-server

An MCP server that connects AI agents and coding assistants to Postman workspaces, collections, and specifications to automate API testing and client code generation.

FrontierCS/Frontier-CS

A benchmark for evaluating AI on unsolved, open-ended, and diverse computer science research and algorithmic problems to move beyond saturated textbook benchmarks.

ianarawjo/ChainForge

A visual data-flow environment for battle-testing and evaluating LLM prompts across multiple models and parameter permutations.

predibase/lorax

A multi-LoRA inference server that allows users to serve thousands of fine-tuned LLMs on a single GPU to reduce serving costs.

reyamira/models

A TUI and CLI tool for browsing AI models, comparing performance benchmarks, tracking AI coding agents, and monitoring provider service statuses.

RobertTLange/evosax

A high-performance JAX library implementing over 30 Evolution Strategies for efficient, vectorized neuroevolution on hardware accelerators.

numba/llvmlite

A lightweight Python binding for LLVM that simplifies the creation of JIT compilers by providing a stable interface to LLVM's IR builder, optimizer, and JIT compiler APIs.

margelo/react-native-fast-tflite

A high-performance TensorFlow Lite library for React Native that enables efficient on-device AI inference using zero-copy ArrayBuffers and GPU acceleration.

Qengineering/Jetson-Nano-Ubuntu-20-image

A pre-configured Ubuntu 20.04 OS image for NVIDIA Jetson Nano that comes pre-installed with OpenCV, TensorFlow, PyTorch, and TensorRT.

YingfanWang/PaCMAP

A dimensionality reduction method for visualization that preserves both the local and global structure of high-dimensional data.

KernelTuner/kernel_tuner

An auto-tuning tool for GPU kernels that benchmarks and optimizes parameters across multiple GPU programming languages to improve performance and and energy efficiency.

spences10/mcp-sequentialthinking-tools

An MCP server that provides a structured scratchpad for AI models to record sequential reasoning steps, validate tool plans, and manage reasoning history.

NVlabs/timeloop

An infrastructure for modeling, mapping, and code-generation for dense and sparse tensor algebra workloads on accelerator architectures.

QuantumKitHub/TensorKit.jl

A Julia package for large-scale tensor computations with symmetries, optimized for building tensor network algorithms to simulate quantum many-body systems.

OpenMined/PySyft

A privacy-preserving framework that lets data scientists run computations on private data without theit data leaving the owner's machine, using existing cloud storage for transport.

lucidrains/vector-quantize-pytorch

A PyTorch library implementing various vector quantization techniques, including VQ, Residual VQ, and FSQ, to map continuous data into discrete representations for generative modeling.

castorini/pyserini

Pyserini is a Python toolkit that wraps the Anserini Lucene search engine and FAISS to provide easy, reproducible first‑stage retrieval (sparse, dense, or hybrid) for information‑retrieval research and RAG pipelines. It includes pre‑built indexes, evaluation scripts, and a REST server, and can be installed via `pip`.

rapidsai/cugraph

A collection of GPU-accelerated graph analytics packages that enable the creation and manipulation of graphs and the execution of scalable graph algorithms.