Accepted papers
Daily Routines, Sport Access, and Wellbeing: DiversityOne with a MIDUS Benchmark
- Aleksei Osipov, Higher School of Economics, Russian Federation
- Anton Nikolaev, Higher School of Economics, Russian Federation
- Egor Kozlov, Higher School of Economics, Russian Federation
Abstract
Daily wellbeing varies within people, but evidence rarely shows whether deviations from personal routines recur across populations and measurement systems. This paper studies sleep, physical activity, screen use, social activity, and same-day wellbeing in two panels. DiversityOne combines an international student smartphone panel with diaries. Midlife in the United States (MIDUS) provides a U.S. adult diary benchmark through its National Study of Daily Experiences (NSDE). Depending on the behavior, the fixed-effects analyses include 1,122–2,400 person-days from 123–183 DiversityOne participants and 2,321–3,008 person-days from 296–386 working MIDUS adults. We estimate person fixed-effects models separately in each panel and compare them with alternative specifications. These observational models remain sensitive to unobserved daily shocks. Fixed effects associate above-usual sleep, physical activity, and social interaction with higher wellbeing in both panels. We test whether future behavior predicts current wellbeing and whether prior wellbeing predicts subsequent behavior. Neither pattern remains statistically significant after adjustment for multiple testing. The DiversityOne sport-access-by-weather instrument strongly predicts physical activity, and the two-stage least-squares estimate provides evidence of a positive causal effect of physical activity on wellbeing under the instrument assumptions.Training-Free Few-Shot Personalization for Mobile Mood Inference with TabPFN
- Panyu Zhang, KAIST, Korea
- Shohruh Shokulov, HumbleBeeAI, Korea
- Humoyunbek Abdukarimov, HumbleBeeAI, Korea
- Jumabek Alikhanov, Bee Intelligence Global, Korea
- Surjya Ghosh, BITS Pilani Goa, India
- Uichin Lee, KAIST, Korea
Abstract
Mobile sensing models often perform poorly for users and geographic populations not represented during model development. Prior work on DiversityOne showed that partial personalization with a Random Forest Hybrid Model improves mood inference by incorporating labeled examples from target users. However, this Hybrid Model must fit a new task-specific classifier whenever the target user, source population, or personalization budget changes. We investigate whether TabPFN can instead use the same labeled target-user examples through in-context learning, without task-specific parameter updates. We reconstruct the Hybrid Model on the DiversityOne release and compare it with TabPFN under matched experimental conditions: identical preprocessing, source-population data, labeled target-user personalization data, and fixed held-out target-user test data. We evaluate two-class and three-class valence prediction across six geographic configurations: country-specific, continent-specific, Country-Agnostic I, Country-Agnostic II, natural multi-country pooling, and balanced multi-country pooling. To examine the effect of labeled target-user support data, we report results separately at five personalization budgets: 10%, 20%, 30%, 40%, and 50%. Across these settings, TabPFN generally matches or improves over the Random Forest Hybrid Model, with the most consistent gains in Country-Agnostic I and country-specific three-class prediction. However, its advantage is not universal: performance depends on task formulation and pooling strategy, and the Hybrid Model remains stronger in some pooled three-class settings. These findings position tabular in-context learning as a strong training-free personalization baseline for mobile mood inference.
From Shared Corridors to Everyday Context: Linking DiversityOne GPS Mobility with Time-Diary Socio-Behavioral Signals Across Cities
- Senih Kirmac, Sabancı University, Turkey
- Mustafa Bozyel, Sabanci University, Turkey
- Selim Balcisoy, Sabanci University, Turkey
- Ömer Mert Özel, Sabancı University, Turkey
Abstract
Population-level “shared path” maps and individual everyday life need not tell the same story. We analyze DiversityOne smartphone GPS across eight cities with an OpenStreetMap (OSM) corridor pipeline, then join per-user mobility metrics to time-diary reports of activity, place, and social context. The same data-centric pipeline yields institutional campus corridors in Jilin that co-occur with high diary work/study place share (50.3%, 𝑛 = 35), versus home- dominated diaries in London (82.3%, 𝑛 = 36) and Trento (86.4%, 𝑛 = 98) alongside central mixed-use or arterial corridor ecologies. Social framing also diverges: Jilin reports lower alone share (33.9%) and higher work/study social context (17.3%) than London/Trento (≈ 65% alone; ≈ 4% work/study social). Case studies of top shared- corridor contributors (Jilin User 363; London User 10) show that heavy population-path use can coexist with weak personal GPS routines and sharply different diary profiles. Median splits further link higher diary travelling share to lower GPS routine lock-in. We frame results as observational computational social science on everyday-life diversity shaped jointly by institutional structure and pandemic-era policy—not isolated culture effects—and discuss privacy-preserving 50 m aggregation (25–75 m sensitivity) under jurisdiction-aware, GDPR-aligned minimization.Can Smartphone Sensing Safely Replace Time-Diary Prompts?
- Sizhe Xu, New York University, New York, United States
- Zhaonan Wang, New York University, Shanghai, China
Abstract
Time diaries require repeated participant input, but automatically filling an entry replaces a self-report with a potentially wrong sensor-based inference. We ask whether smartphone sensing can safely suppress selected prompts. On 31,668 participant-held-out DiversityOne slots, we evaluate four activity classifiers as selective pre- dictors: each model fills only its highest-confidence cases and leaves the remainder as prompts. The audit measures site-balanced selective accuracy together with fixed-class breadth, activity-specific coverage and safety, cross-site stability, and prompt-versus-error cost. Supervised scores identify an accurate subset at every site, whereas zero-shot language-model scores do not. Yet high aggregate accuracy is dominated by Sleep and coexists with unsafe accepted predictions for several activities; broadening allocation across predicted classes sacrifices accuracy. Selective filling therefore cannot be authorized by a global confidence threshold alone. It requires class-specific validated error bounds, an explicit cost model, and provenance for every inferred row.Smartphone Interaction Patterns and Momentary Mood: Within-Person Evidence Across Eight Countries
- Ambika Grover, Harvard Medical School, United States
- Chirag Patel, Harvard Medical School, United States
Abstract
Phone-use fragmentation can describe several distinct interaction patterns, yet these patterns are often combined into a single behavioral measure. We examined five patterns using DiversityOne, a four-week smartphone-sensing dataset collected across eight countries. The analytic sample comprised 493 participants and 190,568 active 60-minute windows preceding mood reports. Linear mixed-effects models estimated within-person associations while adjusting for phone-use volume, time of day, weekend status, and country. Mood was scored from 1 (happy) to 5 (sad), so positive coefficients indicate worse mood. The five measures showed different, and sometimes opposing, associations. A greater proportion of short sessions was associated with better mood (β = −0.021), whereas app-switch rate, touch rate, and burstiness were associated with worse mood (β = 0.007–0.012). Notification rate had a smaller positive association (β = 0.003). All estimated differences were 0.021 points or less on the five-point scale per one-standard-deviation within-person increase. An exploratory composite showed heterogeneous associations across country cohorts (Wald χ²(7) = 57.2, p = 5.4 × 10⁻¹⁰), spanning positive, negative, and null estimates. Findings were similar using 30-minute windows, non-overlapping windows, a logistic model of negative mood, and short-session cutoffs from 15 to 60 seconds. These results suggest reporting interaction patterns separately rather than treating fragmentation as one transferable correlate of mood.Understanding Behavioral Dark Patterns of High BMI Individuals
- Manjeet Yadav, Indian Institute of Technology, India
- Prasenjit Karmakar, Indian Institute of Technology, India
- Suchetana Chakraborty, Indian Institute of Technology, India
Abstract
Understanding how everyday behaviors influence body weight is essential for designing effective and personalized health interventions. Existing studies largely rely on self-reported questionnaires or limited sensing modalities, making it difficult to capture the temporal dynamics of daily behavior. In this work, we analyze the DiversityOne dataset, comprising four weeks of passive smartphone sensing and ecological momentary assessments collected from 453 university students across eight countries. We extract behavioral features spanning dietary habits, physical activity, screen time, and smartphone usage, and investigate their associations with self- reported Body Mass Index (BMI). Beyond feature-level analysis, we employ Hidden Markov Models (HMMs) to uncover latent behavioral patterns. Our analysis reveals that higher BMI is associated with more frequent consumption of soda, alcohol, and processed meat. We further reveal that overweight and obese individuals spend longer periods in food delivery apps and are more likely to transition back to unhealthy eating and drinking routines after starting to exercise. In contrast, normal-weight individuals lead a more balanced lifestyle. These findings highlight key behavioral patterns that make weight loss particularly challenging.