openaiotlab/CUHK-X
[MobiSys 2026] A large-scale, multimodal dataset and benchmark for Human Action Recognition, Understanding and Reasoning
解決的問題
CUHK-X 解決了人類行為識別(HAR)與推理領域缺乏全面且同步的多模態資料集的問題。它提供了一個大規模資源,推動研究從單純的動作分類,進展至複雜的人類行為理解(HAU)與下一步動作推理(HARn),這對於醫療監測與智慧環境至關重要。
如何運作
本項目提供一個包含 64,267 個樣本的多模態資料集,涵蓋七種同步模態:RGB 影像、紅外線(IR)、熱成像、深度圖、毫米波雷達、骨骼資料與 IMU 傳感器。資料被分為兩類:「小模型資料」(用於單一動作分類)與「大模型資料」(用於序列化活動)。倉儲包含用於小模型基線的 PyTorch 流水線,以及用於評估大視覺-語言模型(VLM)如 QwenVL 與 Video-LLaVA 的基準程式碼,支援動作描述、情緒分析與下一步動作預測等任務。
適用對象
從事多模態學習、感測器融合、人類行為識別,以及在真實物理場景中評估大語言模型(LLM)的研究人員與開發者。
亮點
- 七種同步模態:整合視覺(RGB、IR、熱成像、深度)、空間(骨骼)與非視覺(mmWave 雷達、IMU)資料。
- 推理基準:包含序列重排、時間推理與人類意圖因果推論等特定任務。
- LLM 驅動的標註:使用提示式框架生成邏輯性與時空性場景描述。
- 多樣化評估:支援跨試驗、跨受試者(LOSO)與跨領域性能分析。
相關
- 專案
vual/lobe-chat-proLobe Chat Pro is an open‑source, self‑hosted AI chat platform (fork of lobe‑chat) that adds a full admin console, user/group management, multi‑vendor model pricing, payment integration (WeChat, Stripe, etc.), and multimodal “infinite canvas” panels for image, music, and video generation. Deployable via Docker in three flavors: full back‑office, gateway‑only with login, or gateway‑only without login.
- 專案
- 專案
pwilkin/trellis.cppA standalone C++ implementation of the TRELLIS image-to-3D pipeline using GGML, enabling the generation of textured 3D models from images without Python at runtime.
- 專案
hassancs91/claude-youtube-editorAn open-source pipeline that automates YouTube video editing by turning raw talking-head recordings into polished videos with AI-generated visuals, audio cleaning, and automated uploads.