openbezal/rhema
A real-time AI-powered desktop app that detects Bible verses in live audio feeds and renders them as broadcast-ready overlays via NDI.
lucidrains/transfusion-pytorch
A PyTorch implementation of the Transfusion architecture that unifies next-token prediction for text and flow matching for continuous modalities like images in a single model.
landing-ai/ade-python
A Python library for the LandingAI ADE API that parses PDFs and images into structured Markdown and extracts typed data using Pydantic models.
NEKOparapa/ReaDreamAI
An AI-powered platform that automates the creation of novels, including text generation, consistent character illustrations, and video adaptations.
OpenEnvision/WorldFoundry
An open-source infrastructure for world models that provides a unified stack for inference, asset staging, and benchmark evaluation across video generation, 3D/4D representations, and embodied AI.
alexiglad/EBT
A framework for Energy-Based Transformers (EBTs) that enables scalable reasoning and System 2 Thinking across text, image, and video modalities.
potamides/DeTikZify
DeTi*k*Zify is an open‑source multimodal LLM that converts sketches or raster scientific figures into editable TikZ code. It supports image‑to‑TikZ generation, text‑to‑TikZ via Ti*k*Zero adapters, and iterative refinement with Monte‑Carlo Tree Search. Installable via pip, it requires a full TeX Live, Ghostscript, and Poppler installation, and runs on GPUs (8‑b models) or CPU (1‑b). The repository provides a CLI/web UI, Python API examples, model weights on Hugging Face, and scripts to recreate the training datasets.
apple-aiml-research/ml-hierarchical-confusion-matrix
Neo is an open‑source JavaScript library (npm package `@apple/hierarchical-confusion-matrix`) that provides an interactive visualisation for confusion matrices with hierarchical and multi‑output class labels. Install via npm, feed it a spec and an array of `{actual, observed, count}` records, and embed the matrix in any web page. It supports flat, hierarchical (`:`), multi‑output (`,`), and combined label structures, and comes with a live demo, unit tests, and a full TypeScript source tree.
lxtGH/OMG-Seg
A unified framework for visual perception and reasoning that combines image, video, and pixel-level segmentation into a single model, reducing the need for specialized specialist models.
baxtree/subaligner
Subaligner is an open‑source Python/CLI tool that synchronises subtitles with video/audio, can transcribe speech via Whisper, translate subtitles with HuggingFace models, and even train custom alignment models. It supports many subtitle and media formats, offers fast global‑shift (`single`) and high‑accuracy two‑stage (`dual`) alignment, and provides Docker, pip, and optional extras for LLM‑based features.
Softlandia-Ltd/vision-is-all-you-need
A Vision RAG (V-RAG) demo that embeds PDF pages as images using a VLM to eliminate the need for text chunking during document retrieval.
openaiotlab/CUHK-X
A large-scale multimodal dataset and benchmark for human activity recognition and reasoning, integrating seven synchronized sensor modalities to enable complex action understanding.
EleutherAI/gpt-neox
GPT‑NeoX is an open‑source, Megatron‑based library for training billion‑parameter language models on multi‑GPU clusters. It adds DeepSpeed‑style optimizations, supports many launchers (Slurm, MPI, etc.), works on NVIDIA and AMD GPUs, and includes modern features like flash attention, Mixture‑of‑Experts, and preference‑learning fine‑tuning. Ideal for research labs with large compute budgets; for inference‑only use, Hugging Face `transformers` is recommended.
mbzuai-oryx/Video-ChatGPT
A video conversation model that combines LLMs with a spatiotemporal visual encoder to enable detailed, natural language discussions about video content.
mbzuai-oryx/groundingLMM
GLaMM is a pixel-grounding large multimodal model that integrates natural language responses with object segmentation masks for precise visual grounding.
mbzuai-oryx/LLaVA-pp
LLaVA++ enhances LLaVA 1.5 by integrating LLaMA-3 and Phi-3 LLMs to improve visual instruction-following and academic task performance.
NVlabs/OmniVinci
OmniVinci is an open-source omni-modal LLM that jointly understands vision, audio, and text using specialized temporal and alignment architectures to improve cross-modal reasoning.
blendi-remade/falcraft
A Minecraft Fabric mod that uses fal.ai to generate 3D structures and live AI video broadcasts directly within the game world from text prompts.
TetreesEX/TetreesAgent_EX
Tetrees Agent EX is a Node.js SDK/CLI that lets developers search, quote, run, extend, and publish AI Packs on the Tetrees EX marketplace via the public MCP contract. It provides example scripts for buyer, builder, and seller flows, and keeps all credentials local.
JavisVerse/JavisDiT
A Diffusion Transformer framework for joint audio-video generation that ensures semantic and temporal alignment between sound and visuals from text prompts.
caiyuanhao1998/Open-DiffusionGS
A single-stage image-to-3D generation and reconstruction framework that integrates Gaussian Splatting into a diffusion denoiser for fast, scalable 3D creation.
open-gigaai/giga-models
GigaModels is an open‑source Python toolbox that wraps dozens of pre‑trained vision, diffusion and multimodal models (e.g., Grounding DINO, Depth Anything, GigaBrain‑0, Pi0, Cosmos‑Predict2.5) behind a unified `load_pipeline` API. It supports both inference and training, offers ready‑to‑run scripts, and is installable via conda + pip.
huggingface/huggingface-gemma-recipes
A collection of minimal recipes and notebooks for implementing multimodal inference, fine-tuning, and RAG with the Gemma family of models.
microsoft/XPretrain
XPretrain is a Microsoft Research repository that provides code, pre‑trained checkpoints, and the HD‑VILA‑100M video‑language dataset for large‑scale multimodal pre‑training on video‑text and image‑text pairs. It includes models such as HD‑VILA, LF‑VILA, CLIP‑ViP, Pixel‑BERT, SOHO, and VisualParsing, each linked to recent conference papers.
OliverDOU776/Few-step-probabilistic-glucose-forecasting-from-continuous-glucose-monitoring-and-meal-images
A fast probabilistic forecasting tool that uses conditional rectified flow to predict glucose levels, transferring knowledge from large source datasets to small target datasets.
wit-ai/pywit
A Python SDK for Wit.ai that enables developers to easily integrate text and speech natural language understanding into their applications.
6551Team/opentwitter-mcp
opentwitter‑mcp is a Python MCP server that lets AI assistants fetch Twitter/X profiles, tweets, searches and real‑time follower events via a set of ready‑made tools and a WebSocket push API.
wladradchenko/wunjo.wladradchenko.ru
An all-in-one AI media tool for face swapping, voice cloning, and video generation that runs locally for privacy.
erdogant/pca
A Python library that simplifies Principal Component Analysis (PCA) by wrapping scikit‑learn’s PCA/SparsePCA/TruncatedSVD and adding ready‑made biplots, explained‑variance charts, outlier detection (Hotelling T², SPE/D‑ModX), feature‑importance extraction, model saving/loading, and utilities for normalising variance and projecting new data.
MemeMeow-Studio/MemeMeow
A natural language meme retrieval tool that uses embedding models and vision AI to find memes based on moods or scene descriptions.
tensorflow/model-analysis
TensorFlow Model Analysis (TFMA) is a library for large‑scale, slice‑aware evaluation of TensorFlow models. It runs distributed Beam pipelines, visualizes results in Jupyter, and integrates with TFX/Kubeflow pipelines.
tensorflow/data-validation
TensorFlow Data Validation (TFDV) is a scalable Python library for computing data statistics, inferring schemas, and detecting anomalies in machine‑learning datasets. It integrates with TensorFlow/TFX pipelines, uses Apache Beam for distributed processing, and provides visual UI tools for inspecting data quality.
polyaxon/haupt
Haupt is a Polyaxon service that provides APIs for lineage metadata, streaming artifacts/metrics, interactive sandbox sessions (SSH/tmux), a local viewer API, and hosted notebook spaces—tools that help manage, monitor, and reproduce machine‑learning experiments.
tonywu71/colpali-cookbooks
A collection of notebooks for implementing, fine-tuning, and interpreting ColPali and ColQwen2 models for vision-based document retrieval.
vincentarelbundock/marginaleffects
`marginaleffects` is an R/Python library that simplifies extracting predictions, contrasts, marginal effects, and hypothesis‑test results from over 100 statistical and machine‑learning model types, backed by a free companion book.
eeeXun/gtt
gtt is a Go‑based terminal UI that lets you translate text using many online services (Google, DeepL, Bing, Apertium, LibreTranslate, Reverso, ChatGPT). It supports API‑key configuration, custom key bindings, colour themes, clipboard integration, and text‑to‑speech, and can be installed via binaries, package managers, Docker, or building from source.
deeplearning4j/deeplearning4j-examples
A set of Maven‑based Java examples that demonstrate how to use the Eclipse Deeplearning4J ecosystem (DL4J, ND4J, SameDiff, DataVec, RL4J, etc.) for tasks ranging from basic neural‑network training to Spark‑distributed and GPU‑accelerated learning, plus model import and Android deployment.
wandb/examples
A curated set of ready‑to‑run scripts and Colab notebooks that show how to plug the Weights & Biases tracking library into popular ML frameworks (PyTorch, TensorFlow/Keras, Hugging Face, XGBoost, scikit‑learn, etc.). It demonstrates logging runs, configs, metrics, gradients, and artifacts, helping users add experiment tracking and reproducibility to their models with just a few lines of code.
ai-ng/2txt
A fast image-to-text conversion tool built with Next.js and the Vercel AI SDK using the GPT-5-nano model.
leamsigc/ShortsGenerator
An automated pipeline for creating short-form videos that handles AI script generation, stock footage sourcing, TTS voiceovers, and social media scheduling.
win4r/openclaw-a2a-gateway
OpenClaw A2A Gateway – a Node.js plugin that implements Google’s Agent‑to‑Agent v0.3.0 protocol for OpenClaw agents. Provides zero‑config installation, automatic peer discovery (DNS‑SD/mDNS), multi‑transport fallback (JSON‑RPC, REST, gRPC), bio‑inspired routing, circuit‑breaker resilience, bearer‑token auth, and file‑transfer tooling.
duolahypercho/fusion-fable
Fusion‑Fable is a Claude Code skill that runs a prompt in parallel on multiple LLMs (Opus 4.8, GPT‑5.5, Gemini 3.1 Pro), lets Opus 4.8 judge the independent answers, and synthesises a final, structured response. It installs as a Claude skill with slash commands, saves provenance files, and includes an iterative planning mode (`/fusion‑plan`) that integrates with the OMC planning system. Zero‑setup panel (two Opus 4.8 runs) works out‑of‑the‑box; richer panels need the `codex` and `agy` CLIs. The trade‑off is higher token cost and latency, but the result is a higher‑quality, auditable answer.
nikkigallery/Whimbox
An AI agent that uses LLMs and image recognition to automate daily tasks, navigation, and activities in the game Infinity Nikki via screen capture and input simulation.
commaai/commavq
A world model framework and dataset for autonomous driving that uses VQ-VAE for video compression and a GPT-based model to predict future driving scenes.
kyegomez/MultiModalMamba
A multi-modal AI model that integrates Vision Transformer (ViT) and Mamba to efficiently process and interpret text, images, audio, and video data simultaneously.
ashnkumar/sketch-code
A deep learning model that converts hand-drawn web mockups into HTML code using an image captioning architecture.
mlcommons/training
A repository of reference (unoptimized) implementations for the MLPerf Training benchmark suite, covering vision, language, recommendation, and graph models. Each benchmark includes model code, Dockerfiles, dataset scripts, and a training‑run script that measures time to reach target quality. Intended for benchmark submissions, hardware evaluation, and learning, but not for real‑world performance.
luciddreamer-cvlab/LucidDreamer
LucidDreamer is a system for domain-free generation of 3D Gaussian Splatting scenes from a single image and a text prompt.