kubeflow/arena

Arena is a CLI tool that allows data scientists to run and monitor machine learning training jobs on GPU clusters without needing deep Kubernetes knowledge.

tonl-dev/tonl

A token-optimized notation language and data platform that reduces JSON size and LLM token costs by up to 45% while remaining human-readable.

openml/OpenML

An online machine learning platform for sharing and organizing data, algorithms, and experiments to foster open science and collaborative research.

pyRiemann/pyRiemann

A machine learning package for processing and classifying multivariate data using the Riemannian geometry of symmetric positive definite matrices, primarily for biosignals and remote sensing.

kubernetes-sigs/dra-driver-nvidia-gpu

A Kubernetes DRA driver for NVIDIA GPUs that enables flexible device allocation and secure multi-node NVLink connectivity for AI workloads.

svg-project/flash-kmeans

An IO-aware, batched K-Means clustering implementation using Triton GPU kernels that enables fast, memory-efficient clustering of massive and high-dimensional datasets.

litexlang/golitex

Litex is a set-theory-based formal language designed for readable, machine-checked mathematics that allows users to write facts in a natural order and optionally compiles to Lean.

retrage/gpt-macro

A Rust procedural macro that uses ChatGPT to generate and fill in code implementations and tests at compile-time.

Nativu5/Gemini-FastAPI

An OpenAI-compatible API wrapper for web-based Gemini models that allows free access using session cookies instead of an API key.

benbrandt/text-splitter

A Rust-based text splitting library that breaks long documents into semantically sensible chunks for LLM context windows, supporting plain text, Markdown, and source code.

facebookresearch/optimizers

A collection of advanced PyTorch optimization algorithms, including Distributed Shampoo and GPA-AdamW, designed for efficient large-scale model training.

jpmml/jpmml-sklearn

A Java library and command-line tool that converts Scikit-Learn machine learning pipelines into the PMML standard for portable model deployment.

socprime/detectflow-main

A real-time cyberattack detection system that uses AI and Flink-based pipelines to process streaming events and apply Sigma rules in-flight, drastically reducing detection latency.

Samge0/ragflow-upload

A batch document upload and parsing tool for RAGFlow that automates the import of large local datasets into knowledge bases via API.

kunzmi/managedCuda

A .NET wrapper for NVIDIA's CUDA Driver API and libraries, enabling C# and Visual Basic developers to integrate GPU acceleration into their applications.

cdt15/lingam

A Python library for causal discovery that estimates linear non-Gaussian acyclic models to determine the causal relationships between variables.

modal-labs/modal-client

Official SDKs for the Modal platform, enabling developers to deploy and high-performance serverless applications and interact with platform resources.

GACWR/OpenUBA

An open-source User and Entity Behavior Analytics (UEBA) framework for security analytics that provides transparent, white-box security models and a visual rule builder.

NPC-Worldwide/npcsh

A composable multi-agent shell that integrates natural language and bash commands into a single interface for seamless AI-driven command-line experiences.

SciML/BlackBoxOptim.jl

A global optimization package for Julia that uses stochastic and meta-heuristic algorithms to optimize non-differentiable functions.

Avdpro/ai2apps

A local-first AI application and agent platform that turns model runtimes into durable apps and versioned services, featuring optimized local inference for large MoE models.

microprediction/timemachines

A streaming anomaly detection library that uses calibrated forecasters to turn data streams into surprise signals with controlled false-alarm rates.

tweag/monad-bayes

A Haskell library for probabilistic programming that enables the modular composition of Bayesian inference algorithms.

daac-tools/vibrato

A fast Rust-based tokenizer and morphological analyzer that optimizes the Viterbi algorithm for high-speed tokenization, especially for large dictionaries.

OnlyTerp/UltraCode-Shim

A loopback proxy that enables Claude Code's UltraCode mode to work with any LLM backend, featuring smart routing and orchestrator/worker splitting.

facebookresearch/spdl

SPDL is a library for scalable and performant data loading, providing pipeline abstractions and operations for processing array data efficiently.

ROCm/rocBLAS

rocBLAS is a Basic Linear Algebra Subprograms library implemented in HIP and optimized for AMD GPUs to provide high-performance linear algebra operations.

mljs/matrix

A matrix manipulation and computation library for JavaScript that provides essential linear algebra operations and decompositions.

vlang/vsl

A high-performance scientific computing library for the V language providing linear algebra, numerical methods, and GPU acceleration for AI and scientific research.

cornell-zhang/allo

A Python-embedded, MLIR-based language and compiler for the modular design and programming of high-performance machine learning accelerators.

SciML/ComponentArrays.jl

A Julia package that provides structured, named access to flat vectors, allowing complex scientific models to be composed without manual index tracking.

tournesol-app/tournesol

Tournesol is a collaborative platform for identifying public-interest videos to build an open database for AI ethics and recommendation system research.

microprediction/microprediction

A collection of quantitative finance and statistical prediction tools focusing on online time series analysis, covariance estimation, and portfolio optimization.

antoinersx/clawhost

An open-source cloud hosting platform that enables one-click deployment and management of dedicated VPS instances for AI agent runtimes like OpenClaw and Hermes.

LeapLabTHU/limit-of-RLVR

A research project and evaluation framework that analyzes whether Reinforcement Learning with Verifiable Rewards (RLVR) expands an LLM's reasoning capacity or merely optimizes its sampling efficiency.

GoogleCloudPlatform/cluster-toolkit

An open-source toolkit by Google Cloud for deploying and managing AI/ML and high performance computing (HPC) environments using customizable blueprints.

lovstudio/Ataru

A high-performance local memory retrieval system that indexes AI session transcripts from various agents to make historical conversations searchable via GUI, CLI, or API.

HomebrewML/HeavyBall

A high-performance PyTorch optimizer library that uses compiled, composable building blocks to fuse operations into Triton kernels for faster training and lower memory usage.

haroldsultan/MCTS

A Python implementation of Monte Carlo Tree Search (MCTS) used to experiment with decision-making in a state-based game environment.

vijaymasand/PyChem-Pro

A pure-Python desktop application and library for chemistry and cheminformatics that implements molecular visualization, SMILES parsing, and MMFF94 optimization from scratch without external C++ dependencies.

NVIDIA/nvidia-resiliency-ext

A resiliency framework for PyTorch-based AI training that maximizes goodput by providing automatic fault detection, in-process restarting, and efficient checkpointing.

openucx/ucc

A unified collective communication library designed for high-performance AI/ML and HPC workloads across CPU, GPU, and DPU hardware.

sgaofen/cli-in-wechat

A bridge service that lets you run and control AI programming CLI tools like Claude Code and Gemini CLI remotely via WeChat using the official ClawBot API.

InfuseAI/ArtiVC

A command-line tool for data versioning on cloud storage that allows users to snapshot and switch between data versions without requiring a separate server.

cncf/k8s-ai-conformance

A certification program that defines standardized capabilities for Kubernetes platforms to ensure AI and machine learning workloads are portable and reliable across different clusters.

PreferredAI/cornac

A comparative framework for multimodal recommender systems that simplifies the integration of auxiliary data and the benchmarking of various recommendation algorithms.

lmcinnes/pynndescent

A Python library for fast approximate nearest neighbor search using the Nearest Neighbor Descent algorithm, supporting a wide range of distance metrics.

steipete/macos-automator-mcp

An MCP server that enables AI agents to control macOS applications and system functions by executing AppleScript and JavaScript for Automation (JXA) scripts.