google-deepmind/kfac-jax
A JAX-based library for second-order optimization of neural networks using the K-FAC optimizer and scalable curvature approximations.
Stability-AI/stable-audio-metrics
A collection of evaluation metrics for music and audio generative models, providing tools to calculate Fréchet Distance, KL divergence, and CLAP scores for long-form stereo audio.
mixa3607/ML-gfx906
A set of Ubuntu packages and Docker images that bring current ROCm, PyTorch, llama.cpp, ComfyUI and related GPU‑tuning tools to the legacy AMD GFX906 (Vega 20/MI50) GPU.
haoyi-duan/WorldScore
WorldScore is an open‑source benchmark for evaluating world‑generation models (3‑D, 4‑D, video). It provides a dataset, generation adapters, a suite of spatial‑temporal metrics, and a public leaderboard. Users register their model via a small YAML config and a Python class, generate videos, run the evaluation script, and submit the resulting JSON score.
sveltejs/ai-tools
The official Svelte MCP server that allows AI models to interact with Svelte-related data and tools using the Model Context Protocol.
CodeSoul-co/Hypha
A TypeScript framework for building governed, durable AI agents that separates reasoning from execution via a Production Harness and DomainPacks.
UditAkhourii/neuroarxiv
A research-to-decision tool that fetches arXiv papers and uses an isolate-then-converge process to provide a single, grounded architectural recommendation instead of a list of search results.
Entrpi/ds4-on-spark
A high-performance serving stack for DeepSeek-V4-Flash on NVIDIA DGX Spark, featuring continuous batching, speculative decoding, and a memory governor for massive context windows.
modelstudioai/cli
Aliyun Model Studio CLI (bailian‑cli) is a terminal tool for Alibaba Cloud’s AI platform, providing commands for multimodal generation (text, image, video, speech), asset understanding, managed‑agent orchestration, dataset validation, fine‑tuning, deployment, and account management. Install via npm or a one‑line script, authenticate with an API key or OAuth, then issue `bl …` commands or let an AI Agent translate natural‑language prompts into CLI calls.
messkan/prompt-cache
A lightweight, self-hosted semantic cache for GenAI workloads that reduces LLM costs and latency by serving cached responses for similar prompts.
MahmoudAshraf97/ctc-forced-aligner
A Python package for efficient forced alignment of audio and text using Hugging Face CTC models, supporting over 1,100 languages.
dbccccccc/ttsfm
An OpenAI-compatible text-to-speech API service and Python SDK that converts text to natural speech using the openai.fm backend.
ttktjmt/mjswan
A framework for creating real-time, interactive MuJoCo simulations with AI policy control that run entirely in the browser as static sites.
visose/Robots
A plugin for Rhino 8 and Grasshopper that allows users to create, simulate, and generate manufacturer-specific code for various industrial robots.
OpenLegged/URDF-Studio
URDF Studio is a browser‑based robot authoring workstation that lets you edit URDF/MJCF/USD models, assemble multiple robots, export to many formats, and use an AI assistant for generation and inspection. Built with React, Three.js, TypeScript and Zustand, it also ships a reusable `@urdf‑studio/react‑robot‑canvas` component. Apache‑2.0 licensed.
KavrakiLab/vamp
A hardware-accelerated motion planning library that uses CPU SIMD instructions to perform collision checking and forward kinematics in microseconds.
livekit/rust-sdks
A Rust client SDK for LiveKit, enabling real‑time video/audio/data streaming, room management, and hardware‑accelerated encoding across desktop and mobile platforms.
real-stanford/umi-on-legs
A framework for combining human demonstrations with simulation-trained whole-body controllers to enable mobile manipulation skills on quadruped robots with arms.
Simple-Robotics/proxsuite
A collection of numerically robust and efficient numerical solvers for Linear and Quadratic Programs, designed for high-performance robotics and optimization tasks.
Stable-Baselines-Team/stable-baselines3-contrib
A contribution package for Stable-Baselines3 containing experimental reinforcement learning algorithms and niche tools for researchers and developers.
rsasaki0109/lidar_slam_ros2
A ROS 2 LiDAR SLAM system that converts sensor bags into Autoware-compatible map bundles, including point clouds and auto-generated lanelet2 files.
zai-org/GLM-V
A series of vision-language models (VLMs) that enhance multimodal reasoning, supporting native function calling, visual grounding, and complex document understanding.
misyaguziya/VRCT
VRCT is a translation and transcription tool designed to help VRChat users communicate across different languages via real-time audio-to-text and chat integration.
wenet-e2e/wenet
A production-oriented end-to-end speech recognition toolkit that provides full-stack solutions for streaming and non-streaming speech-to-text.
lukaszliniewicz/Pandrator
A unified workspace for creating audiobooks, subtitles, and voiceovers by coordinating local and cloud-based speech and language models.
ssitu/ComfyUI_UltimateSDUpscale
A set of ComfyUI nodes that perform image-to-image diffusion on large images using a tiled approach to improve detail and reduce hardware requirements.
KohakuBlueleaf/LyCORIS
A library implementing various parameter-efficient fine-tuning algorithms like LoHa and LoKr for Stable Diffusion, enabling high-quality model customization with minimal storage and compute.
wooyeolbaek/attention-map-diffusers
attention‑map‑diffusers is a Python library that captures and visualises the attention tensors of Hugging Face Diffusers image and video generation pipelines, letting users see how prompts, image patches, or video frames influence each other.
mybigday/llama.rn
User posted a long README of llama.rn covering many features. Assistant offers to help with specific questions and provides a concise TL;DR of key usage points.
Paritok-official/paritok-4b-v1
Paritok is an open‑source proxy that sits between coding agents (Claude Code, Cursor, Codex, OpenHands, etc.) and LLM APIs. It reduces input‑token usage by (1) filtering irrelevant tool schemas, (2) compressing file reads/tool outputs/history with a 4 B model, and (3) summarising old turns. The gateway is a Python package (`paritok[proxy]`) that can run locally (via Ollama or vLLM) or via a hosted GPU service. Reported savings are ~25 % on the first turn and >60 % in longer sessions, translating to noticeable cost reductions. The project is Apache‑2.0 licensed and includes a dashboard and VS Code extension.
Cocolalilal/LastChat
A feature-rich AI assistant app for Android that offers multi-provider support, RAG-based memory, and built-in code execution engines.
lhotse-speech/lhotse
A Python library for flexible multimodal data preparation that optimizes the loading and manipulation of speech, audio, video, and image data for ML training.
NucleoidAI/Nucleoid
A declarative logic programming language and runtime for building world models in Neuro-Symbolic AI, designed to reduce LLM hallucinations through structured reasoning.
holoviz/lumen
An open-source agent-based framework for chatting with data and RAG, enabling users to generate data pipelines and visualizations via natural language.
christopherkarani/Swarm
A Swift-native agent runtime for building type-safe AI agents with support for on-device Apple Foundation Models and composable workflows.
Project-MONAI/MONAILabel
An intelligent open-source ecosystem for AI-assisted medical image annotation that uses a server-client architecture to enable interactive and automated labeling for radiology, pathology, and endoscopy.
barun-saha/slide-deck-ai
An AI assistant that generates professional PowerPoint presentations from topic descriptions or PDF files using various LLMs.
Deodat-Lawson/LaunchStack
LaunchStack is a TypeScript‑first, port‑based engine for AI‑native applications, providing ingestion, OCR, RAG, LLM chat, and background‑job infrastructure. A Next.js reference app demonstrates wiring the engine to PostgreSQL + pgvector, S3, Inngest, and Google Gemini (or any OpenAI‑compatible endpoint). Packages are currently unpublished but ready for release; the repo can be run locally or via Docker.
yuga-hashimoto/openclaw-assistant
A native Android voice client for OpenClaw and Hermes Agent that enables self-hosted AI assistants to control device hardware and system functions.
yihong1120/Construction-Hazard-Detection
An AI-driven construction site safety monitoring system that uses YOLO to detect PPE violations and proximity hazards from live camera feeds.
zae-bayern/elpv-dataset
A benchmark dataset of 2,624 annotated electroluminescence images of solar cells used for the visual identification of defects that reduce power efficiency.
Climate-Vision/ClimateVision
An open-source machine learning platform that uses deep learning and satellite imagery to automatically detect deforestation, arctic ice melting, and flooding.
a-agmon/rs-graph-llm
A stateful graph workflow framework for AI agents in Rust, providing a type-safe way to build complex, resumable, and interactive agentic workflows.
ModelEngine-Group/unified-cache-management
A cache management framework for LLMs that reduces inference latency and GPU memory usage by persisting KV caches and implementing sparse attention retrieval.
gobii-ai/gobii-platform
An AI employee platform for running durable, always-on autonomous agents that can be contacted via email or SMS and perform production browser automation.
av/harbor
Harbor is a CLI‑driven Docker‑Compose orchestrator that lets you spin up a complete local LLM stack—including model back‑ends (Ollama, llama.cpp, vLLM, MLX, etc.), chat front‑ends (Open WebUI, LibreChat, AnythingLLM, …), web‑search, voice, image generation, and Boost‑based agentic workflows—with a single `harbor up` command.
alondmnt/joplin-plugin-jarvis
An AI note-taking assistant for Joplin that enables chatting with your notes, semantic search, and automated literature reviews using online or offline LLMs.
ChaitanyaEswarRajeshJakki/gemini-youtube-automation
An autonomous AI bot that uses Gemini 2.5 Flash to write, produce, and upload daily educational YouTube videos and Shorts without human intervention.