AvaLovelace1/BrickGPT

BrickGPT is an AI system that generates physically stable and buildable toy brick models from text prompts by fine-tuning a Llama-3.2 model.

johnson7788/MultiUserClaw

MultiUserClaw is a Docker‑based SaaS framework that lets you run isolated AI assistants (Hermes agents) for many users. It includes a React front‑end, a FastAPI gateway for auth, container management and LLM proxying, per‑user Hermes containers, PostgreSQL storage, and features like knowledge bases, skill store, cron jobs, multi‑channel bots, and multi‑model support.

video-db/call.md

A desktop app that records and transcribes meetings in real-time, providing live AI coaching, conversation metrics, and automated post-meeting summaries.

new-sankaku/manga-editor-desu

A web-based manga creation tool that combines professional layout and lettering tools with optional AI image generation and LLM prompt assistance.

EvolvingLMMs-Lab/NEO

NEO is a series of native vision-language models using an encoder-free dense architecture to unify pixel and word processing for efficient multimodal understanding.

TIGER-AI-Lab/VLM2Vec

A unified framework for training and evaluating omni-modality embedding models that produce fixed-dimensional vectors for text, images, videos, audio, and visual documents.

roboflow/maestro

A streamlined tool for accelerating the fine-tuning of multimodal vision-language models like Florence-2, PaliGemma 2, and Qwen2.5-VL.

R3gm/SoniTranslate

SoniTranslate is a web application for translating videos into different languages with synchronized audio, combining transcription, translation, and voice dubbing in one interface.

ailia-ai/ailia-models

A collection of 419 pre-trained AI models across various domains like vision, audio, and text, all runnable via a unified CLI with automatic weight downloads.

ohdearquant/lionagi

lionagi is a Python library and CLI for building, running, and governing multi‑agent LLM workflows. It offers typed, persisted conversation state, supports parallel fan‑outs and DAG‑based flows, integrates both model‑specific CLIs and API providers, and includes a web UI (Lion Studio) for visual run management.

nikvdp/cco

cco is a sandboxing wrapper for AI coding agents (Claude Code, Codex, etc.) that runs them in an isolated environment to protect your system from prompt injections and accidental damage, while preserving a seamless terminal experience.

Sportinger/MasterSelects

A browser-based media editor for video, audio, and 3D that allows AI coding agents to build new tools and effects directly into the workspace during the creative process.

OfficeDev/microsoft-365-agents-toolkit

A Microsoft‑provided toolkit (VS Code/Visual Studio extensions and CLI) that scaffolds, debugs, and deploys AI‑enabled agents, chatbots, and other extensions for Microsoft 365 Copilot, Teams, and Office. It bundles authentication helpers, Azure Functions support, hot‑reload debugging, and CI/CD pipelines, targeting both JavaScript/TypeScript and .NET developers.

shrimbly/node-banana

A visual node-based workflow editor for building complex AI media generation pipelines across multiple providers for image, video, audio, and 3D content.

modelstudioai/cli

Aliyun Model Studio CLI (bailian‑cli) is a terminal tool for Alibaba Cloud’s AI platform, providing commands for multimodal generation (text, image, video, speech), asset understanding, managed‑agent orchestration, dataset validation, fine‑tuning, deployment, and account management. Install via npm or a one‑line script, authenticate with an API key or OAuth, then issue `bl …` commands or let an AI Agent translate natural‑language prompts into CLI calls.

NetManAIOps/ChatTS

ChatTS is a multimodal LLM designed for native understanding and reasoning over multivariate time series, allowing users to perform conversational question-answering on numerical data.

jimmysu0309/shinkansen

Shinkansen is a real AI-powered browser extension that translates web pages, YouTube subtitles, and documents (PDF, Word, EPUB, etc.) using Google Gemini, Google Translate, or custom OpenAI-compatible models. It preserves original layouts and timestamps, offers bilingual outputs, and processes files locally for privacy.

Graylab/IgFold

IgFold is a deep learning tool that predicts 3D antibody structures from amino acid sequences, enabling fast and accurate structural modeling for drug discovery.

apple-aiml-research/ml-fastvlm

FastVLM is an efficient vision encoding framework that uses a hybrid encoder called FastViTHD to drastically reduce latency and token count for high-resolution images in Vision Language Models.

apple-aiml-research/ml-flextok

FlexTok is a system for resampling images into 1D token sequences of flexible length, allowing images to be represented as sequences that can be truncated while still enabling image reconstruction.

merveenoyan/smol-vision

A collection of recipes for shrinking, optimizing, and customizing cutting-edge vision and multimodal AI models to improve efficiency and performance.

sign/translate

A real-time bidirectional translation system that converts sign language videos into spoken text and audio, and converts spoken language into sign language via avatars or GANs.

Pangu-Immortal/MagicWX

An Android app that enables fully offline text chat and Stable Diffusion image generation using local runtimes like ONNX, MediaPipe, and MNN.

AIFrontierLab/TorchUMM

A unified framework for the inference, evaluation, and post-training of multimodal models, supporting image understanding, generation, and editing through a single interface.

aiqm/torchani

TorchANI 2.0 is a PyTorch library for training and using ANI‑style neural‑network interatomic potentials, providing GPU‑accelerated energy/force predictions and tools for molecular dynamics and ML/MM simulations.

escalante-bio/mosaic

mosaic is a JAX‑based Python library that provides a unified, gradient‑optimizable interface for dozens of protein‑property and structure‑prediction models. By letting users compose differentiable loss terms (binding affinity, stability, solubility, inverse‑folding likelihood, etc.) and run continuous‑relaxation design with custom optimizers, it enables multi‑objective protein engineering in a single, JIT‑accelerated framework. The repo includes many state‑of‑the‑art models (Boltz, AlphaFold, OpenFold‑3, Protenix, ProteinMPNN, ESM, etc.), example notebooks, and detailed documentation, but it requires users to tune hyper‑parameters and have a GPU/TPU‑compatible JAX installation.

PozzettiAndrea/ComfyUI-Sharp

A ComfyUI wrapper for Apple's SHARP model that enables monocular 3D Gaussian Splatting from a single image in under one second.

2U1/Qwen-VL-Series-Finetune

A training toolkit for fine-tuning Qwen-VL series models (Qwen2-VL, Qwen2.5-VL, Qwen3-VL, Qwen3.5) supporting SFT, DPO, GRPO, and multimodal data.

tong-io/tongflow

TongFlow is an open‑source visual studio that lets you build multi‑modal generative AI pipelines by connecting simple “add”, “transform”, and “combine” nodes. It supports text, image, video, audio, 3‑D and documents, offers a rich plugin ecosystem (official API plugins, routers, and GPU/CPU compute plugins), and can be used via a tiny desktop client or self‑hosted backend. Licensed under AGPL‑3.0.

Simbastack-hq/framedex

A queryable knowledge base for video and photo archives that generates plain-text AI descriptions, transcripts, and metadata sidecars for every file.

geekyutao/Inpaint-Anything

A unified framework that combines SAM and generative inpainting models to remove, fill, or replace objects in images, videos, and 3D scenes.

dai-hongtao/InkTime

An AI-powered e-ink photo frame that uses vision models to score and caption memories, displaying the best photo from "today in history" daily.

AlexGladkov/claude-in-mobile

`mcp-devices` (aka `claude‑in‑mobile`) is a modular Node/Rust server that lets Claude‑style AI agents control Android, iOS, web, desktop, Aurora, and HarmonyOS devices. Install the base (`npm i -g mcp-devices`), add platform plugins (e.g., `@mcp-devices/plugin-android`), enable them, and point an MCP‑compatible client at the server. The tool provides actions like screenshots, taps, UI inspection, live debugging, and testing, enabling AI‑driven device automation and embodied‑agent research.

OTeam-AI4S/ODesign

An all-atom generative world model for biomolecular interaction design that generates proteins, ligands, and nucleic acids tailored to bind to specific targets.

qntx/ovo

Ovo is a Rust library that provides an embeddable multi‑agent runtime kernel. It lets a host application spawn lightweight agents (turns) that run Rhai scripts, with optional OS‑level sandboxing (macOS Seatbelt or Linux Landlock). The crate offers configurable approval policies, a sandbox‑backend trait, and optional toolkit features for jailed filesystem access. It is pre‑stability (0.9.x) and dual‑licensed MIT/Apache‑2.0.

facebookresearch/MetaCLIP

A worldwide scaling recipe for CLIP that enables high-performance multilingual image-text alignment without degrading English performance.

skrub-data/skrub

skrub is a Python library that supplies dataframe‑focused cleaning, encoding, and similarity‑grouping tools, all wrapped as scikit‑learn‑compatible transformers, to streamline the preprocessing stage of tabular machine‑learning pipelines.

liangnjupt/VisTouch

A large-scale synchronized dataset of video, audio, and tactile force data from robotic sliding contact used for material recognition and cross-modal learning.

livekit/rust-sdks

A Rust client SDK for LiveKit, enabling real‑time video/audio/data streaming, room management, and hardware‑accelerated encoding across desktop and mobile platforms.

pinellolab/DNA-Diffusion

A diffusion-based generative model for creating cell-type specific synthetic regulatory DNA sequences of 200bp.

OpenMOSS/AnyGPT

AnyGPT is an open‑source multimodal LLM (text, speech, images, music) that uses discrete tokenisation to unify all modalities. It provides base and chat checkpoints, inference scripts, and the AnyInstruct dataset for training and zero‑shot multimodal tasks.

MoonshotAI/Kimi-K2.5

Kimi K2.5 is an open-source, native multimodal agentic model with a 1T parameter MoE architecture that integrates vision and language understanding for complex task execution.

muthuishere/mcp-server-bash-sdk

A pure‑Bash implementation of the Model Context Protocol (MCP) 2026‑07‑28, providing zero‑overhead stdio and local‑HTTP bindings for AI‑agent tool calling. It auto‑discovers `tool_*` functions, validates JSON against the official schema, includes full test suites, and offers step‑by‑step docs for building custom Bash‑based MCP servers.

receptron/mulmocast-cli

An AI-native presentation platform that uses a JSON-based intermediate language called MulmoScript to generate multi-modal content including videos, podcasts, and slide decks.

lupantech/MathVista

MathVista is a benchmark for evaluating the mathematical reasoning of foundation models in visual contexts, combining 6,141 examples from 31 multimodal datasets.

jiaxiaogang/HE

he4o is an AGI system based on Helix Entropy Reduction Theory that combines pre-designed rules with dynamic learning to achieve embodied intelligence and lifelong learning.

Yu-Yang-Li/StarWhisper

An AI framework for astronomy that automates telescope observations, classifies light curves and pulsars, and provides a specialized skill pack for astronomical research.

gokayfem/ComfyUI_VLM_nodes

A production-oriented suite of ComfyUI nodes for vision-language models, providing structured detection, segmentation, and adaptive video reasoning with VRAM optimizations.