Six papers accepted at IMWUT 2025

Published on

We have six new papers accepted at IMWUT, to be presented at UbiComp 2026 in Shanghai!

Choube et al., IMWUT 2026

Akshat’s paper introduces DAIMON, an AI-augmented research dashboard designed to enable new human-AI collaborative workflows for researchers conducting longitudinal sensing studies.

Researchers conduct longitudinal passive sensing studies in in-the-wild settings, often spanning months or years, to uncover naturalistic behavioral patterns. These studies are not "set-and-forget" deployments; they require continuous monitoring as technical failures and declining participant compliance can lead to substantial missing data, undermining study validity and downstream models. Thus, conducting these studies involves multiple detail-oriented, cognitively demanding, and time-consuming tasks, making it a burdensome and stressful process. Existing research dashboards, the primary tools for data monitoring, offer limited support in easing this burden. Leveraging recent advances in AI for passive sensing data, we explore the design of human-AI collaborative workflows enabled through research dashboards to improve the effectiveness and efficiency of monitoring and associated tasks. We begin with a co-design study with 13 researchers involved in longitudinal sensing studies to identify desired AI capabilities and interactions through semi-structured interviews, brainstorming, and sketching activities. We operationalize novel human-AI workflows our participants envisioned by implementing an AI-augmented dashboard prototype DAIMON, and use it as a research probe in two studies: a task-based study and a deployment within an ongoing real-world sensing study. Our findings demonstrate the promise of AI-augmented dashboards in supporting researchers' day-to-day data monitoring and decision-making tasks. It also surfaces concerns around transparency and expectations with AI systems. Consolidating insights across all three studies, we present design guidelines for AI-augmented dashboards for longitudinal passive sensing research and discuss directions for future work.

Das Swain et al., IMWUT 2026

Vedant’s paper examines seamful design for human-in-the-loop digital phenotyping of mental health, exploring how users can better understand, evaluate, and influence sensing-based mental health models.

Digital Phenotyping of Mental Health (DPMH) through passive sensing is a promising approach for personal health informatics and digital wellbeing. Its appeal lies in unobtrusiveness, making it appear seamless. However, this very quality leads users to find it impersonal, untrustworthy, and disengaging. To counteract challenges of seamlessness, researchers propose seamful design to deliberately engage users. Yet, it remains unclear how this principle can be incorporated into digital phenotyping. To address this, we conducted a formative study by developing DYMOND. It is a technology probe that estimates depression, explains estimates, reveals discrepancies, and provides user control over the underlying model. In a 6-week deployment, 22 individuals with moderate-severe depression monitored their state with DYMOND. They interviewed every two weeks with researchers to collaboratively reconfigure the model and co-design new interfaces. Our analysis of 57 sessions revealed (i) seams---friction points---across data, modeling, and output, and (ii) design requirements helping users evaluate and mitigate seams. These findings inform the design requirements for human-in-the-loop DPMH to support agency, transparency, and collaborative reflection. This study provides insight into theoretical re-conceptualization for passive sensing, opportunities to integrate large language models and human-AI interaction for better interfaces for digital mental health, and pathways to involve expert stakeholders in DPMH.

Le et al., IMWUT 2026

Ha’s paper presents GLOSS4HAR, a multi-agent LLM-based system that combines passive sensing data and participant self-reports to correct activity annotations and support lower-effort activity labeling.

Accurate measurement of behaviors is critical for research in human-computer interaction, ubiquitous computing, and personal health informatics, because it underpins many health tracking and intervention systems. Most data collection studies, however, still rely on participants' self-reports and manual annotations, which often under- or over-estimate activity duration and type, and require substantial effort from researchers to clean and validate. An automated system that can combine passive sensing data with participants' self-reports might detect inconsistencies and suggest corrections. We introduce GLOSS4HAR, a multi-agent LLM-based system designed to mimic human sensemaking and assist researchers in cleaning and refining activity annotations. We demonstrate the potential of GLOSS4HAR in two key tasks: (1) correcting and reconciling participant self-annotations, and (2) triangulating passive sensing data with different forms of lightweight self-reports to generate accurate activity timelines. Our evaluation shows that GLOSS4HAR improves annotation quality by up to 9.9% in F1 score and can reconstruct activity timelines that align with human annotations at 75-92% F1. Based on our findings, we discuss the implications of our work for the next generation of activity annotation systems that might use human-AI collaboration.

Naito et al., IMWUT 2026

Yuna’s paper introduces a person- and context-aware framework for adaptively selecting filter parameters when estimating pulse rate variability from wearable wrist PPG signals.

Wearable biosensing allows for continuous monitoring and intervention in daily settings. Wrist photoplethysmography (PPG) is a commonly used method for measuring cardiac activity ambulatorily. A typical data preprocessing step is to apply a fixed, one-size-fits-all band-pass filter before peak detection and pulse rate variability (PRV) calculation. However, our analyses reveal that fixed filtering can cause significant errors in PRV estimation even with minimal motion artifacts, and the best filter settings differ across individuals and contextual states. Based on these findings, we introduce a person- and context-aware framework to adaptively select filter parameters at a window level. We instantiate this framework using models that gate segments by signal quality and select band-pass cutoff frequencies for each window. In tests across three datasets, our adaptive method improved PRV (RMSSD) accuracy by up to 279ms compared to a fixed filter and enhanced stress detection. It also increased the amount of usable inter-beat intervals without losing accuracy. Our collective results suggest shifting focus from solely removing motion artifacts toward adaptive selection of filter configurations that considers who is being measured and their state.

Negi et al., IMWUT 2026

Samarth’s paper, Beyond the Binary, introduces PRISM, a framework that moves beyond binary definitions of receptivity by modeling receptivity to digital health interventions as a probabilistic time-to-event distribution.

Just-In-Time Adaptive Interventions (JITAIs) aim to support health behavior by providing the right support at the right time. A critical determinant of JITAI efficacy is timing delivery such that the user is receptive, defined as the cognitive and behavioral capacity to receive, process, and use support. While prior work has explored context sensing to predict receptivity, standard approaches typically operationalize this construct as a binary outcome within a fixed window, despite theoretical definitions characterizing availability as a continuous, time-varying state. Modeling receptivity at this granularity increases learning complexity, and deep sequence models are further constrained by the scarcity of labeled interaction data in mHealth settings. To address these challenges, we propose PRISM, a deep learning framework for modeling receptivity as a probabilistic time-to-event distribution from longitudinal mobile sensing data. PRISM combines a Channel-Independent Transformer (PatchTST) encoder with a discrete-time survival objective and employs self-supervised pre-training on unlabeled sensor traces via Masked Patch Reconstruction to mitigate label scarcity. We evaluate PRISM on the LvL UP intervention dataset, leveraging data from preliminary studies for pre-training and a large-scale efficacy trial for evaluation. Our results demonstrate competitive performance with established receptivity benchmarks, with self-supervised initialization yielding up to 15.5\% improvement in median AUC across heterogeneous user groups. Beyond single-window evaluation, PRISM remains stable across multiple decision horizons from a single trained model, in contrast to binary baselines that degrade or require retraining at each cutoff. A follow-up evaluation on an independent dataset with variable prompt timing provides preliminary evidence that the learned representations transfer across cohorts and schedules. These findings suggest that PRISM provides a data-efficient pathway for deploying resilient, time-aware receptivity models in real-world mHealth systems.

Negi et al., IMWUT 2026

Samarth’s second paper presents the first systematic evaluation of cross-study generalization for receptivity models, examining whether models can transfer across interventions, populations, and sensing configurations.

ust-in-Time Adaptive Interventions (JITAIs) offer a promising paradigm for delivering personalized mobile health support. Receptivity, i.e., the user's ability to receive the intervention plays a critical role in the effectiveness of JITAIs. While prior work has shown that receptivity can be inferred from mobile sensing data, the generalizability of these models across interventions, populations, and sensing configurations remains largely unexplored. Developing a new model for each intervention is burdensome and costly, whereas transferring models across studies is complicated by heterogeneity in populations, sensors, and data collection protocols. This paper presents the first systematic evaluation of cross-study generalization in receptivity detection. Using data from four diverse studies (yielding seven distinct datasets), we establish within-study performance benchmarks (median AUCs = 0.687-0.844) and quantify the performance drop when models transfer using only shared features (mean generalization AUCs of up to 0.676). We evaluate strategies that leverage the full feature union, employing imputation-based methods, correlation alignment (Deep CORAL), and a missingness-aware neural network. These approaches close the generalization gap, improving mean generalization AUCs of up to 0.722. In two target datasets, cross-study models matched or exceeded within-study baselines. These findings demonstrate the feasibility of robust, generalizable "warm-start" receptivity models, offering a practical pathway for deploying effective JITAIs without a costly "cold-start" learning phase.

Congratulations to Akshat, Ha, Vedant, Yuna, Samarth, and all collaborators!