HITsz-TMG/Uni-MoE

Uni-MoE is a Mixture-of-Experts based omnimodal large model capable of understanding and generating across multiple modalities, including text, images, and speech.

J3n5en/EnsoAI

EnsoAI is an Electron‑based desktop workbench that couples Git worktree management with persistent AI‑agent sessions (Claude, Gemini, Codex, etc.). It provides a built‑in Monaco editor, visual Git UI, 3‑way merge tool, and one‑click IDE bridging, enabling developers to work on multiple branches in parallel while each branch enjoys its own dedicated LLM context.

tensorflow/java

TensorFlow‑Java provides Maven‑publishable Java/JVM bindings for TensorFlow, including low‑level JNI wrappers (`tensorflow‑core`), a high‑level neural‑network API (`tensorflow‑framework`), and an independent ndarray library. It supports Linux (x86_64, arm64) and macOS (Apple Silicon) binaries, with optional GPU builds, and integrates with standard Java build tools.

mistralai/client-python

Official Python SDK for Mistral AI’s cloud APIs (chat, embeddings, audio, files, agents, etc.). Install via pip/uv/poetry, set `MISTRAL_API_KEY`, then call methods like `client.chat.complete(...)` synchronously or with `asyncio`. Includes Azure and GCP wrappers and support for streaming, pagination, retries, and more.

nextcloud/recognize

A smart media tagging app for Nextcloud that uses local AI models to automatically categorize photos, videos, and music.

inclusionAI/Ming

Ming-flash-omni 2.0 is an open-source omni-MLLM that unifies multimodal perception and generation across text, image, audio, and video in a single MoE-based architecture.

cgnomads/GSOPs

A SideFX Houdini plug-in that provides a comprehensive toolset for importing, editing, and animating Gaussian splatting scenes for VFX production.

sambanova/bloomchat

BLOOMChat is a 176‑billion‑parameter multilingual chat LLM fine‑tuned from the open‑source BLOOM model. The repo provides data‑prep, tokenisation, and training scripts (the latter for SambaNova’s RDU hardware) plus detailed GPU inference instructions using Hugging Face’s Bloom inference code. It’s a genuine, large‑scale LLM project aimed at researchers and developers who want to reproduce or run the model.

remyxai/VQASynth

A pipeline for transforming image datasets into spatial VQA datasets to improve the 3D spatial reasoning and distance estimation capabilities of Vision-Language Models.

smartscanapp/smartscan-android

An Android app that turns images and videos into a searchable personal knowledge base using on-device processing and automatic media grouping.

AIGeeksGroup/3D-R1

3D-R1 is a generalist 3D Vision-Language Model that enhances spatial reasoning and scene understanding through synthetic CoT data and reinforcement learning.

kozistr/pytorch_optimizer

A PyTorch‑compatible library that aggregates 100+ research optimizers, several LR schedulers and loss functions behind a single, consistent API, with optional integrations for low‑precision training.

OmniCustom-project/OmniCustom

OmniCustom is a joint audio-video generation framework that creates synchronized videos preserving a specific visual identity and voice timbre based on reference inputs and text prompts.

kyegomez/ScreenAI

An implementation of the ScreenAI vision-language model designed for understanding user interfaces and infographics.

Cerlancism/chatgpt-subtitle-translator

A Node.js CLI/Web UI that uses the OpenAI ChatGPT API (or compatible services) to translate SRT subtitle files line‑by‑line, preserving timing. It batches lines, supports structured JSON output, prompt‑caching, optional moderation, and an advanced multi‑pass “agent” mode for richer translations.

HUANGLIZI/LViT

LViT is a multimodal framework that combines language and vision transformers to improve the accuracy of medical image segmentation in CT scans and other medical imagery.

MME-Benchmarks/Video-MME-v2

A robust benchmark for evaluating video understanding in MLLMs, using a three-level progressive difficulty scale and grouped non-linear scoring to measure true temporal reasoning.

software-mansion-labs/private-mind

A fully offline mobile AI application that enables private, on-device chat, document analysis, and multimodal interactions without cloud dependency.

VrchStudio/comfyui-web-viewer

A custom node collection for ComfyUI that enables real-time interactive AI art by integrating streaming and external controls like MIDI, OSC, and gamepads.

ixaxaar/pytorch-dnc

A PyTorch library implementing Differentiable Neural Computers (DNC), Sparse DNC (SDNC) and Sparse Access Memory (SAM) with install instructions, API usage, debugging support, and example training tasks (copy, addition, argmax).

tianclll/Ace-Translate

A privacy-focused, offline document translation tool that translates various file formats while preserving original layouts using local LLMs and OCR.

PipeNetwork/kimi-k3-mlx

An MLX port of Moonshot's Kimi-K3 multimodal MoE model, featuring a streaming converter and REAP pruning to enable execution on Apple Silicon.

pipecat-ai/pipecat-client-web

A web client SDK and React library for connecting to and interacting with voice and multimodal AI applications built with the Pipecat framework.

facebookresearch/mmf

MMF is a modular PyTorch-based framework for vision and language multimodal research, providing reference implementations of state-of-the-art models and tools for distributed training.

om-ai-lab/VLM-FO1

VLM-FO1 is a plug-and-play module for Vision-Language Models that enhances fine-grained perception and spatial awareness without compromising general reasoning capabilities.

OpenEnvision/Awesome-Multimodal-Modeling

Awesome‑Multimodal‑Modeling is a community‑curated “awesome‑list” that surveys multimodal AI models. It defines four clear families—Traditional, MLLMs, Unified, and Native—explains their architectural differences, and links to open‑source model collections on Hugging Face. The repo serves researchers and engineers who need a structured view of the rapidly evolving multimodal landscape.

Sharrnah/whispering-ui

A native Windows UI for the Whispering Tiger application that enables real-time transcription and translation of audio streams and in-game images.

ardha27/AI-Waifu-Vtuber

An AI-powered VTuber assistant that integrates speech recognition, LLMs, and text-to-speech to create an interactive virtual character for live streaming and personal use.

tongjingqi/Thinking-with-Video

A new multimodal reasoning paradigm and benchmark (VideoThinkBench) that evaluates the ability of video generation models to solve complex visual and textual reasoning tasks.

AceDataCloud/Nexior

Nexior is an MIT‑licensed, Vue‑based app that aggregates 40+ chat, image, audio and video AI models into a single self‑hostable UI. It offers one‑click Vercel or Docker deployment, BYOK or a single AceData key, and includes built‑in user accounts, payments and referral mechanics, making it a ready‑made SaaS starter for AI products.

awekrx/ChatGPT-MidJourney-prompt

A Python library that uses LLMs to automatically generate detailed and creative prompts for MidJourney images based on simple text hints.

HorizonWind2004/reconstruction-alignment

A self-supervised post-training method for unified multimodal models that improves image generation and image editing by training models to reconstruct images from their own visual features.

Licoy/GoAmzAI

A private, multimodal AIGC platform for individuals and enterprises that integrates text, image, video, music, and PPT generation into a unified management system.

a616567126/GPT-WEB-JAVA

A Java-based web platform that integrates GPT, Midjourney, and Stable Diffusion into a commercial AI service with user management and payment systems.

lone-cloud/gerbil

A cross-platform desktop application for running local text and image generation on your own hardware, featuring integrated model searching and support for various GPU accelerators.

SkyworkAI/UniPic

A unified multimodal model series for image editing, generation, and understanding, featuring high-speed multi-image editing and autoregressive modeling.

noahgsolomon/brainrot.js

A tool for automating the creation of AI-generated videos featuring celebrity voice clones and generated content using Groq, OpenAI, and Speechify.

Fabric-Project/Fabric

A visual node-based creative coding environment for rapid prototyping of interactive 3D graphics, image/video processing, and AI-enhanced visuals on macOS.

askui/python-sdk

A Python SDK that enables AI agents to control desktop and mobile devices using vision-based automation instead of brittle UI selectors.

IliasHad/edit-mind

A local video knowledge base that indexes videos via transcription and visual analysis to enable natural language semantic search of video content.

nv-tlabs/XCube

XCube is a generative model for high-resolution sparse 3D voxel grids that uses hierarchical latent diffusion and VDB data structures to create detailed 3D objects and large-scale scenes.

FoundationVision/Liquid

Liquid is a unified autoregressive multimodal generator that integrates visual comprehension and image generation into a single LLM without needing external visual embeddings.

WHU-USI3DV/VistaDream

VistaDream is a training-free framework that reconstructs high-quality 3D scenes from a single-view image by ensuring consistency across generated novel views.

OpenDCAI/OpenWorldLib

A unified codebase and framework for advanced world models, integrating various open-source research for video generation, 3D scene generation, and visual-action reasoning.

livekit/python-sdks

LiveKit Python SDK provides async client‑side tools to join LiveKit rooms, publish/receive video‑audio tracks, run RPC calls, and control hardware codecs. The companion `livekit‑api` package wraps LiveKit’s server REST endpoints for room creation, token generation, SIP, egress, etc. It also integrates with LiveKit Agents for voice‑AI bots that combine STT, LLM, and TTS models.

ChilloutCharles/BrainFlowsIntoVRChat

A tool that translates EEG and biometric data from biosensors into OSC parameters for VRChat, enabling brain-controlled avatar animations and effects.

Nutlope/make-comics

An AI-powered comic book creator that generates stories, characters, and panels using Google Flash Image 2.5 and Qwen3 80B.

vual/lobe-chat-pro

Lobe Chat Pro is an open‑source, self‑hosted AI chat platform (fork of lobe‑chat) that adds a full admin console, user/group management, multi‑vendor model pricing, payment integration (WeChat, Stripe, etc.), and multimodal “infinite canvas” panels for image, music, and video generation. Deployable via Docker in three flavors: full back‑office, gateway‑only with login, or gateway‑only without login.