GitHpriyanshu23/Smart-Plant-Doctor

An AI and IoT plant health platform that combines real-time environmental sensing via ESP32 with image-based disease detection and a Gemma 4 powered care assistant.

worm128/AI-YinMei

A one-stop AI live-streaming platform that integrates LLMs, streaming TTS, and virtual avatars to create interactive AI bots for multiple streaming platforms.

talesofai/neta-skills

A collection of AI agent skills and CLI tools for the Neta Art API, enabling agents to generate multimedia, manage characters, and run interactive story adventures.

facebookresearch/tuna-2

Tuna-2 is a unified multimodal model that replaces complex vision encoders and VAEs with direct pixel embeddings to improve both image understanding and generation.

aimclub/FEDOT

FEDOT is an open‑source Python AutoML framework that automatically designs and optimises machine‑learning pipelines (including preprocessing, feature engineering and models) using evolutionary algorithms. It supports classification, regression, clustering, and time‑series forecasting, integrates with popular ML libraries, offers a simple high‑level API, and provides reproducible pipeline export.

LucasAlegre/morl-baselines

MORL‑Baselines is a PyTorch library of reliable multi‑objective reinforcement‑learning algorithms (single‑ and multi‑policy, SER/ESR). It follows the MO‑Gymnasium API, provides utilities, automatic W&B logging, and integrates with the Open RL Benchmark for reproducible experiments.

zrt-ai-lab/ViNote

An AI-powered tool that converts videos from YouTube, Bilibili, and local files into structured notes, mind maps, and searchable knowledge assets.

BlockRunAI/blockrun-mcp

BlockRun MCP is an open‑source MCP server that gives AI agents (Claude, Codex, etc.) 19 real‑time tools—market data, web search, media generation, on‑chain queries, and the ability to place real USDC bets on Polymarket. Agents pay per call either with a self‑custody USDC wallet (via the x402 micropayment protocol) or with a prepaid API key. The server is installed via a single `npx` command, works with many MCP‑compatible clients, and supports human‑in‑the‑loop spend approval, making it a practical bridge for autonomous agents to access live data and execute real actions.

guochengqian/Magic123

A system for generating high-quality 3D objects from a single 2D image by combining 2D and 3D diffusion priors in a coarse-to-fine pipeline.

pykale/pykale

A PyTorch-based library for multimodal and transfer learning that provides a standardized, pipeline-based API to create sustainable and reusable AI workflows.

rockbenben/img-prompt

A visual, multilingual prompt builder for AI image and video models that helps users assemble high-quality English prompts using a curated library of 5,000+ bilingual tags with preview images.

huggingface/finetrainers

A library for the accessible training and fine-tuning of diffusion models, focusing on memory-efficient text-to-video and text-to-image generation.

laminlabs/lamindb

LaminDB is an open‑source, Git‑like data‑management system for AI/ML, especially multimodal life‑science data. It version‑controls files, tables, and array formats, tracks code and environment provenance, supports branching/merging, and offers ACID‑safe metadata stored in SQLite/Postgres. Integrated with Python/R tools (Polars, DuckDB), bio‑ontologies, and workflow managers, it enables traceable, FAIR‑compliant datasets for research and biotech.

Woolverine94/biniou

A self-hosted web interface for running a wide variety of generative AI models locally, supporting text, image, audio, video, and 3D generation on hardware as basic as 8GB RAM.

hoanganh8389/bizcity-twin-ai

Bizcity Twin AI is an open‑source WordPress plugin that turns a site into a unified AI “brain” for a business. It normalises messages from many channels, builds a single knowledge‑graph with provenance, and exposes the data to LLMs (Claude, ChatGPT) via a managed control plane. Core services (identity, KG, reasoning, automation, logging, permissions) are immutable; developers add domain‑specific capabilities through lightweight plugins (act, channel, view). The project ships a CLI scaffolding tool, diagnostics, and live demos, and is released under GPL‑2.0‑or‑later.

Minidoracat/mcp-feedback-enhanced

MCP Feedback Enhanced is a security‑hardened MCP server that adds a web‑based (and optional Tauri desktop) UI for human‑in‑the‑loop feedback during long AI tasks. It works locally, over SSH, and in WSL, tracks sessions, supports images, markdown, timers, and multilingual UI, and is maintained as a fork of the original interactive‑feedback‑mcp project.

cyx2333hhh/talk-type

A macOS voice input tool that combines real-time speech recognition, local Whisper fallback, and AI-driven text cleanup for seamless bilingual input across applications.

Blacktrupersist/hailuo-ai-pulse

Hailuo AI Pulse is a streamlined AI workspace for Windows that provides content generation, analysis, and creative automation capabilities.

backblaze-labs/genblaze

An AI pipeline SDK for orchestrating generative video, audio, and image workflows with built-in, hash-verified provenance manifests.

AvalancheAlbatross/midjourney-lab

Midjourney Lab is a Windows-based machine learning platform that supports multi-modal generation and processing of text, images, audio, and video.

matterantmill/kling-ai-pulse

A lightweight Windows AI workspace for content generation, analysis, and creative automation using various AI models.

v-modal/vmodal_sdk_android

An Android SDK for adding multimodal semantic search to video and image libraries, allowing users to find specific moments by meaning.

GangTailorUpgrade/undress-service

A self-hosted AI fashion platform that digitizes your wardrobe and uses generative AI to recommend and visualize outfits based on weather, occasion, and style.

ling-kong-ran/pisper

Pisper is a cross‑platform multi‑agent chat and workflow tool that lets you branch, parallelize, and automate LLM conversations. Built on the open‑source Pi Coding Agent, it offers desktop, terminal, and mobile clients, on‑demand plugins, secure local‑only data storage, and optional LAN/P2P remote linking.

AutoArk/TinyEngram

An open research project exploring Engram-based memory injection for LLMs and Stable Diffusion to enable parameter-efficient knowledge updates without catastrophic forgetting.

AutoArk/EVA-OS

An AIOS for real-time multimodal applications and smart hardware, providing an integrated development experience from AI-native coding to on-device inference.

flatkey-ai/flatkey-cli

A command-line interface for unified multimodal AI generation, allowing media teams and AI agents to generate images, videos, audio, and text through a single API key and credit balance.

hezo-ai/hezo

Hezo is an open‑source server + web UI that lets you build and run organised teams of AI agents (CEO, Captains, engineers, designers, etc.) inside sandboxed containers. You supply your own LLM provider keys, set token and container‑hour budgets, and the platform handles secret protection, Git integration, and a unified “meta‑harness” that normalises different model runtimes. Install with a single binary, self‑host or use Hezo Cloud, and manage projects via an org‑chart‑style UI.

dreamers-laboratory/image-to-3d-pipeline

A pipeline for turning images into 3D meshes that allows users to run and compare multiple open-source reconstruction models to find the best output.

AkshitIreddy/Interactive-LLM-Powered-NPCs

A system that adds dynamic, LLM-powered voice conversations and synchronized facial animations to NPCs in any existing open-world game without modifying the game's source code.

Hchen1218/heytea-style

A creative toolkit that transforms photos into hand-drawn posters and interactive desktop pets by extracting visual elements to create consistent characters.

chflame163/ComfyUI_LayerStyle_Advance

A collection of advanced ComfyUI nodes for high-quality image captioning, VLM inference, and complex image composition using local models and external APIs.

dexhunter/seedance2-skill

A Claude‑compatible “skill” that provides a markdown guide for writing correct Seedance 2.0 video‑generation prompts (ByteDance’s multimodal model). Install by copying the markdown into `~/.claude/skills` or via `npx skills add`. The guide covers constraints, reference syntax, camera language, prompt structures, and templates for ads, dramas, MVs, education, etc., and is sourced from ByteDance’s official docs. MIT‑licensed.

11273/mooc-work-answer

A cross‑platform Python tool that automates learning‑progress tracking, discussion posting and AI‑assisted homework hints on Chinese MOOC services (智慧职教, AI 优课, 资源库). It talks to the official APIs and optionally calls DeepSeek to generate reference answers (60‑100 % relevance). The app can be run as a pre‑built executable or from source, and includes safety delays to mimic human usage.

taruma/SceneFlow

A script-to-screen synchronization tool for AI filmmakers to analyze prompt adherence and evaluate how AI video models visualize screenplay instructions.

geometric-kernels/GeometricKernels

GeometricKernels is a Python library that provides heat and Matérn kernels for non‑Euclidean domains (manifolds, graphs, meshes). It lets Gaussian‑process models run on those spaces and offers adapters for TensorFlow (GPflow), PyTorch (GPyTorch), and JAX (GPJax). Install via `pip install geometric_kernels` and add the backend you need. The repo includes extensive notebooks, a JMLR paper citation, and a clear contribution workflow.

JuliaDiff/ReverseDiff.jl

ReverseDiff.jl is a Julia library for fast, tape‑based reverse‑mode automatic differentiation, capable of computing gradients, Jacobians, Hessians and higher‑order derivatives of native Julia code. It offers tape reuse, non‑allocating linear‑algebra optimisations, mixed‑mode compatibility with ForwardDiff, and is well‑suited as a backend for machine‑learning or scientific‑computing projects that require efficient gradient calculations.

mir-group/flare

FLARE (Fast Learning of Atomistic Rare Events) is a Python library that builds Bayesian interatomic force fields using sparse Gaussian‑process regression and ACE‑type descriptors. It provides uncertainty‑aware molecular dynamics, on‑the‑fly active learning, and a LAMMPS pair‑style implementation, enabling fast, accurate simulations of materials and rare events.

BINE022/EEGPT

EEGPT is a pretrained transformer model designed for universal EEG feature extraction using a dual self-supervised learning method to overcome low signal-to-noise ratios.

IBM/materials

IBM/materials is a genuine open‑source suite of large‑scale foundation models for chemistry and materials science. It offers pre‑trained uni‑modal models (SMILES, SELFIES, graphs, 3‑D structures) and multimodal fusion tools, all accessible via a unified Python wrapper (FM4M‑Kit) or a Hugging Face web UI. The repo includes installation instructions, example notebooks, and links to the hosted models on Hugging Face.

hankcs/HanLP

HanLP is an open‑source multilingual NLP toolkit (Python core, Java/Go clients) that provides production‑ready pretrained models for tokenization, POS, NER, parsing, SRL, AMR, summarization, sentiment, language detection and more. It offers two APIs – a lightweight RESTful service (no GPU needed) and a native Python library that runs on PyTorch/TensorFlow. Models cover 130 languages, with joint multi‑task models for speed and single‑task models for higher accuracy. Installation is a single `pip install hanlp` (native) or `pip install hanlp_restful` (REST). The library returns consistent JSON output and includes visualisation tools, custom dictionary support, and extensive documentation.

Dooy/chatgpt-web-midjourney-proxy

A web UI that merges ChatGPT with Midjourney and dozens of generative‑AI back‑ends (image, video, music, dance, face‑swap). Deploy via Docker, Vercel, or a pre‑built binary; configure through environment variables. Supports voice input, image upload for GPT‑4‑vision, real‑time chat, and a “GPT‑Store” for swapping model endpoints.

Relaxed-System-Lab/Flash-Sparse-Attention

Flash‑Sparse‑Attention (FSA) is a Triton‑based, drop‑in PyTorch module that implements an optimized version of Native Sparse Attention. By reordering loops and splitting the work into three specialized kernels, it cuts memory traffic and removes atomic adds, giving 2‑3× speed‑ups on long‑sequence LLM training and inference on modern NVIDIA GPUs. It works with PyTorch ≥ 2.4, transformers ≥ 4.45, and supports fp16/bf16 on Ampere/Hopper hardware.

fcakyon/phd-skills

A zero-dependency plugin for Claude Code that adds research guardrails, auto-triggering skills, and background agents to help ML researchers avoid costly AI-assistant mistakes (wrong configs, unverified claims, destructive commands) during paper reproduction, experiment debugging, and training launches.

jmiao24/Paper2Agent

Paper2Agent is a multi‑agent Python system that automatically converts a research paper (and its code) into a runnable MCP server, exposing the paper’s methods as AI‑callable tools. Users install the “paper2agent” skill into a coding assistant (Claude Code, Codex, Gemini CLI), then prompt the assistant to *agentify* a paper URL. The system spawns specialist sub‑agents to parse the paper, set up an isolated environment, wrap the original code, run verification tests, and bundle everything into a ZIP with usage instructions. The resulting server can be run locally or hosted on Hugging Face, and then connected back to the coding assistant for interactive scientific queries.

huggingface/meshgen

A Blender addon that uses AI agents to enable natural language control of 3D modeling, supporting both local and remote LLM backends.

marcellodebernardi/loss-landscapes

`loss‑landscapes` is a PyTorch library that samples scalar metrics (loss, gradient norm, curvature, expected return, etc.) on low‑dimensional slices of a model’s parameter space, enabling easy creation of loss‑landscape visualisations. It provides a flexible `Metric` abstraction, sub‑space generators (lines, planes, etc.), and a `ModelWrapper` to work with both standard models and RL agents. The original 2019 release is broken; a rewrite is underway, so the API may change.

20000419/fauxnix

fauxnix is an npm‑based tool that translates a curated subset of bash commands into PowerShell, letting LLM agents run familiar Linux‑style commands on Windows without a VM. It offers deterministic translation, MCP server integration, batch execution, and extensive testing, aiming to improve reliability of AI‑driven code assistants on Windows environments.