openaiotlab/CUHK-X
[MobiSys 2026] A large-scale, multimodal dataset and benchmark for Human Action Recognition, Understanding and Reasoning
解决的问题
CUHK-X 解决了人类行为识别(HAR)和推理领域缺乏全面且同步的多模态数据集的问题。它提供了一个大规模资源,推动研究从简单的动作分类迈向复杂的人类行为理解(HAU)和下一步动作推理(HARn),这对于医疗监测和智能环境至关重要。
如何工作
该项目提供了一个包含 64,267 个样本的多模态数据集,涵盖七种同步模态:RGB 视频、红外(IR)、热成像、深度图、毫米波雷达、骨骼数据和 IMU 传感器。数据被组织为两类:"小模型数据"(用于单一动作分类)和"大模型数据"(用于序列化活动)。仓库包含用于小模型基线的 PyTorch 流水线,以及用于评估大视觉-语言模型(VLM)如 QwenVL 和 Video-LLaVA 的基准代码,支持动作描述、情绪分析和下一步动作预测等任务。
适用对象
从事多模态学习、传感器融合、人类行为识别,以及在真实物理场景中评估大语言模型(LLM)的研究人员和开发者。
亮点
- 七种同步模态:整合视觉(RGB、IR、热成像、深度)、空间(骨骼)和非视觉(mmWave 雷达、IMU)数据。
- 推理基准:包含序列重排、时间推理和人类意图因果推断等特定任务。
- LLM 驱动的标注:使用基于提示的框架生成逻辑性和时空性场景描述。
- 多样化评估:支持跨试验、跨被试(LOSO)和跨领域性能分析。
相关
- 项目
vual/lobe-chat-proLobe Chat Pro is an open‑source, self‑hosted AI chat platform (fork of lobe‑chat) that adds a full admin console, user/group management, multi‑vendor model pricing, payment integration (WeChat, Stripe, etc.), and multimodal “infinite canvas” panels for image, music, and video generation. Deployable via Docker in three flavors: full back‑office, gateway‑only with login, or gateway‑only without login.
- 项目
- 项目
pwilkin/trellis.cppA standalone C++ implementation of the TRELLIS image-to-3D pipeline using GGML, enabling the generation of textured 3D models from images without Python at runtime.
- 项目
hassancs91/claude-youtube-editorAn open-source pipeline that automates YouTube video editing by turning raw talking-head recordings into polished videos with AI-generated visuals, audio cleaning, and automated uploads.