Feasibility of Using a Multi-Agent LLM System to Correct Annotations and Support Low-Effort Activity Labeling

Picture of Stephen S. Intille
Stephen S. Intille
Published at Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. (IMWUT) 2026

Abstract

Accurate measurement of behaviors is critical for research in human-computer interaction, ubiquitous computing, and personal health informatics, because it underpins many health tracking and intervention systems. Most data collection studies, however, still rely on participants' self-reports and manual annotations, which often under- or over-estimate activity duration and type, and require substantial effort from researchers to clean and validate. An automated system that can combine passive sensing data with participants' self-reports might detect inconsistencies and suggest corrections. We introduce GLOSS4HAR, a multi-agent LLM-based system designed to mimic human sensemaking and assist researchers in cleaning and refining activity annotations. We demonstrate the potential of GLOSS4HAR in two key tasks: (1) correcting and reconciling participant self-annotations, and (2) triangulating passive sensing data with different forms of lightweight self-reports to generate accurate activity timelines. Our evaluation shows that GLOSS4HAR improves annotation quality by up to 9.9\% in F1 score and can reconstruct activity timelines that align with human annotations at 75-92\% F1. Based on our findings, we discuss the implications of our work for the next generation of activity annotation systems that might use human-AI collaboration.

Materials