Hierarchical and Multimodal Data for Daily Activity Understanding

  • 2025-04-25 17:07:50
  • Ghazal Kaviani, Yavuz Yarici, Seulgi Kim, Mohit Prabhushankar, Ghassan AlRegib, Mashhour Solh, Ameya Patil
  • 0

Abstract

Daily Activity Recordings for Artificial Intelligence (DARai, pronounced"Dahr-ree") is a multimodal, hierarchically annotated dataset constructed tounderstand human activities in real-world settings. DARai consists ofcontinuous scripted and unscripted recordings of 50 participants in 10different environments, totaling over 200 hours of data from 20 sensorsincluding multiple camera views, depth and radar sensors, wearable inertialmeasurement units (IMUs), electromyography (EMG), insole pressure sensors,biomonitor sensors, and gaze tracker. To capture the complexity in human activities, DARai is annotated at threelevels of hierarchy: (i) high-level activities (L1) that are independent tasks,(ii) lower-level actions (L2) that are patterns shared between activities, and(iii) fine-grained procedures (L3) that detail the exact execution steps foractions. The dataset annotations and recordings are designed so that 22.7% ofL2 actions are shared between L1 activities and 14.2% of L3 procedures areshared between L2 actions. The overlap and unscripted nature of DARai allowscounterfactual activities in the dataset. Experiments with various machine learning models showcase the value of DARaiin uncovering important challenges in human-centered applications.Specifically, we conduct unimodal and multimodal sensor fusion experiments forrecognition, temporal localization, and future action anticipation across allhierarchical annotation levels. To highlight the limitations of individualsensors, we also conduct domain-variant experiments that are enabled by DARai'smulti-sensor and counterfactual activity design setup. The code, documentation, and dataset are available at the dedicated DARaiwebsite:https://alregib.ece.gatech.edu/software-and-datasets/darai-daily-activity-recordings-for-artificial-intelligence-and-machine-learning/

 

Quick Read (beta)

loading the full paper ...