Source-linked AI summary

"Am I Just That Dumb?": Applicability, Action and Verification in Consumer IoT Security Advice

Veerle van Harten, Carlos Hernández Gañán, Michel van Eeten, Simon Parkin

arXiv:2608.25225v1cs.HC

TL;DR

Public campaigns ask IoT users to change default passwords and keep devices updated, assuming they can determine whether those instructions apply. This study observed 28 participants applying two government-issued advice items to six popular consumer devices across 168 sessions. Password and update outcomes varied across devices and targets, showing that users’ determinations, actions, and verification could diverge from the device configuration reached.

  • Problem

    Research had not directly examined whether people can determine a generic IoT advice item’s applicability, target, actionable pathway, and resulting device state.

  • Method

    The study observed 28 participants interpreting and demonstrating two government-issued advice items on three of six popular consumer devices, using observations, think-alouds, interviews, and workload assessments.

  • Results

    Across 84 password sessions, 33 reached no password setting, 50 an account-level setting, and one a device-level setting; across 84 update sessions, 27 reached no update, 19 a companion-app update, and 38 verified firmware updates.

  • Takeaways & Limitations

    Generic advice and device environments jointly shape what users can determine, do, and verify, while visible completion can fail to match the configuration reached.

  • Takeaways & Limitations

    The exploratory findings are based on observed session endpoints and decomposed mechanisms rather than a market census, and absence of a located device-level credential is not proof of absence.

Abstract

from arXiv · show

Public campaigns urge people to change default passwords on Internet of Things (IoT) devices and keep them updated, assuming users can independently determine whether the advice applies. We gave 28 participants in the Netherlands two pieces of government-issued advice reflecting guidance in several countries and asked them to try to apply each to three of six consumer devices selected from bestseller lists, not confirmed feature availability (168 sessions). The protocol asked for each action to be demonstrated rather than completed. Of 84 password sessions, 33 reached no password setting, 50 an account-level setting, and one a device-level setting. Of 84 update sessions, 27 reached no update, 19 a companion-app update, and 38 a verified firmware update. No product had a manufacturer-set credential shared across units as described by the advice; the single device-level credential was unique to its unit. We contribute an account of what generic advice and the devices it addresses let users determine, act on, and verify.

1 Introduction

This study examines whether generic IoT security advice enables users to determine its applicability, act on it, and verify the resulting device state. Across 168 sessions, outcomes varied sharply by device and advice item, and participants’ interpretations often diverged from the configuration reached.

  • Study design: 28 participants completed 168 sessions applying two government-issued advice items to three of six popular consumer devices.Participants interpreted the advice without assistance, while researchers recorded achieved states and participants’ beliefs about those states.
  • Results: 33 of 84 password sessions reached no password setting, 50 reached an account-level setting, and one reached a device-level setting.The sole device-level credential was unique to its individual unit; no sampled product had a shared manufacturer-set credential of the advised kind.
  • Results: 27 of 84 update sessions reached no update, 19 reached a companion-app update, and 38 reached a verified firmware update.Every companion-app update session ended with participants concluding they had applied the advice, while five sessions left the resulting state unclear.
  • Results: Outcomes differed sharply across devices and advice items: one device had 13 of 14 password sessions reach no password setting, while another had 11 of 13 update sessions reach verified firmware updates.Participants consulted support in 66 sessions, but only one password consultation found a device-level password path and nine update consultations found firmware-status paths.
  • Contribution: The study contributes an account of what generic advice and its addressed devices let users determine, act on, and verify.It also argues that some device properties remove the need for the advised action, including unique credentials and automatic updating.

2 Related Work

Prior research documents security behaviors, advice usability, default credentials, and update difficulties, but has rarely observed whether users can apply generic advice to a specific IoT device. This study focuses on applicability, target identification, actionability, and verification.

  • Security usability in the home: Home security outcomes depend strongly on everyday configuration and maintenance practices, while smart-home interfaces are often minimal, non-standardized, and weak at displaying security state.Configuration abandonment has also been reported alongside these interface inconsistencies.
  • Security usability in the home: Existing work often attributes non-secure behavior to motivation, skills, risk perception, or self-efficacy, but does not establish whether the recommended action is available.It also leaves open whether users can identify the advice target, locate a pathway, and verify the resulting state.
  • Advice and default credentials: Generic campaigns promote changing default passwords and keeping devices updated, even though the advice does not distinguish hardware credentials from app or account credentials.Campaigns therefore require recipients to determine applicability and target themselves.
  • Advice and default credentials: Shared default credentials are only one credential class: standards increasingly require passwords unique per device or set by the user, while installed products still contain both types.A unique per-unit credential is not a shared secret exposed through standard lists.
  • Updating: IoT updating varies across devices, may be undocumented or inconsistently described, and can leave users uncertain about update status.Automatic updates remove the user task, whereas manual updates require verifying that the displayed version matches firmware running on the device; the examined advice does not distinguish these cases.

3 Methodology

The study observed 28 participants independently interpret and demonstrate two widely promoted IoT security actions on popular consumer devices, preserving the ambiguity of unsupported home use. Researchers recorded interactions, verbalizations, subjective measures, and post-task reasoning without committing credentials or installing firmware.

  • Study design: 28 participants each tried two government-issued advice items on three of six popular consumer devices, producing 168 sessions.Devices were selected from bestseller lists rather than confirmed feature availability.
  • Study design: Participants interpreted printed advice independently, used available device materials, and received no navigational or interpretive assistance.Researchers used neutral think-aloud prompts and deferred clarification until debriefing.
  • Task protocol: Participants demonstrated how they would change credentials or apply updates without committing changes or installing firmware, keeping device states stable across sessions.Two sessions involved app-initiated updates without participant selection.
  • Measures: Researchers combined video annotations, think-aloud data, NASA-TLX ratings, and debrief interviews to document actions, subjective experience, and reasoning.Interviews were conducted before participants learned what researchers had verified.
  • Advice materials: The advice targeted changing preconfigured default passwords and checking or applying available firmware updates, retaining the public campaigns’ differing levels of procedural detail.Password advice generally lacked locational cues, whereas update advice specified opening an interface, finding Settings, and checking status.

4 Results

Across 168 sessions, participants most often reached account-level password settings or verified firmware status, while device-specific password targets were rare and no sampled product had the shared manufacturer credential described by the advice.

  • Password advice: 33 of 84 password sessions reached no password setting, 50 reached an account-level setting, and one reached a device-level setting.The sole device-level setting was an HP Deskjet PIN unique to that individual unit.
  • Update advice: 27 of 84 update sessions reached no update, 19 reached a companion-app update, and 38 reached a verified firmware update.Verified firmware status included devices already current or updates located without installation.
  • Password advice: No sampled product carried a manufacturer-set credential shared across units and publicly documented or retrievable online.The device-level HP Deskjet credential was manufacturer-set but unique to its unit.
  • Participant conclusions: 51 of 84 password sessions and 54 of 84 update sessions ended with participants concluding that they had applied the advice.Among applied conclusions, 49 password sessions reached account-level settings and 35 update sessions reached verified firmware status.
  • Support use: Support consultations occurred in 66 of 168 sessions; 33 returned no path, while all nine consultations returning a firmware path reached verified firmware status.Password consultations returned 21 no paths and 15 account-level paths.
  • Task experience: Password-session frustration was highest for the Echo Dot at 82.1, where 13 of 14 sessions reached no password setting.Account-level password sessions had a mean NASA-TLX Performance rating of 26.3.

I’ve completed the task.”

Participants often struggled to identify which interface layer generic advice addressed and to verify whether their actions matched the intended security state. Ambiguous labels, absent feedback, and inaccessible support materials turned navigation into exploratory trial and error.

  • Distinguishing the companion app from the firmware: The update task created uncertainty between companion-app updates and firmware updates, even after participants inspected update status.Participants questioned whether a firmware screen was the correct target and sometimes later recognized that they had updated only the app.
  • Distinguishing the companion app from the firmware: The word “firmware” discouraged some participants because they lacked a clear understanding of the term or feared damaging the device.One participant said they had no idea what firmware was, while another viewed it as deep system-level functionality.
  • Determining candidate targets across interface layers: No interface offered a way to establish that a setting was absent, so unresolved searches continued through circular reasoning and settings loops.Participants described navigation as exploratory when advice vocabulary did not match device labels.
  • Support and navigation: Support materials frequently failed to provide a usable path, while virtual assistants looped and external searches required filtering mixed manufacturer and third-party advice.Of 66 consultations, 33 returned no path and 20 returned a path to a setting or update.
  • Participant experience: Participants often attributed navigation difficulty to their own abilities, and some weighed the effort against limited time.Comments included asking whether they were “just dumb,” describing themselves as “tech-illiterate,” and rejecting a ten-minute search.

5 Discussion

Across the six products, generic password and update advice often led participants toward ambiguous or alternative targets, while device interfaces and mental models shaped whether actions could be located and verified.

  • Applicability and target identification: 19 of 84 update sessions ended at a companion-app update rather than the device’s firmware status.All sessions reaching a companion-app update ended with participants concluding that they had applied the advice.
  • Pathways and verification: None of the advice texts named a starting point inside an interface, and support consultations usually returned alternative layers rather than device-level controls.Of 40 password consultations, one returned a device-level path; of 26 update consultations, nine returned a firmware-status path.
  • Interpretation and implications: On the HP DeskJet, 12 of 13 sessions failed to reach the setting where the preconfigured device-level credential could be replaced.Thus, the hardware control remained at factory settings when participants did not reach that setting.
  • Pathways and verification: 85% of Tapo update sessions reached verified firmware updates, compared with 23% on the HP DeskJet and 36% on the Echo Dot.These outcomes reflected how choices were presented rather than navigation depth; Tapo’s terminology and explicit status string facilitated verification.
  • Interpretation and implications: Users often followed familiar application-level pathways and treated visible account-level completion as confirmation that the advice had been applied.This interaction between device environments, generic advice, and pre-existing mental models produced persistent uncertainty about the resulting security state.
  • Applicability and target identification: None of the six products had a shared publicly known credential; the HP DeskJet’s device-level PIN was unique to its hardware unit.The Tapo’s device-level facility was disabled by default, and no device-level facility was found for the remaining four products.

6 Conclusion

The study tested whether generic consumer IoT security advice enables users to determine applicability, identify targets, reach settings, and verify outcomes. Across six bestselling devices, the assumed credential and update architectures frequently did not align with what interfaces exposed, leading users to mistake account-level or app-level actions for completion.

  • 6 Conclusion: 33 of 84 password sessions reached no setting, 50 reached an account-level setting, and one reached a device-level per-unit PIN.No sampled product carried the shared manufacturer-set credential described by the advice.
  • 6 Conclusion: 27 of 84 update sessions reached no update, 19 reached a companion-app update, and 38 reached a verified firmware update.Participants concluded they had applied the advice in all 19 sessions ending at companion-app updates.
  • 6 Conclusion: The six products lacked the common credential and update architecture implicitly assumed by the advice, while interfaces failed to distinguish device-, application-, and cloud-level controls.Users consequently relied on familiar application-level pathways and treated visible account-level completion as confirmation.
  • 6 Conclusion: Future guidance should help users verify when no manual configuration is required and diagnose whether and where a setting exists on specific hardware.The paper recommends standardized terminology distinguishing device-, application-, and cloud-level controls.

Ethical Considerations

The study followed established ethical principles and received institutional ethics approval.

  • Ethical Considerations: The research procedures were guided by Respect for Persons, Beneficence, and Justice.These principles were identified as outlined in the Menlo Report.
  • Ethical Considerations: The host university’s Human Research Ethics Committee reviewed and approved the study.

Stakeholders and Principles

The paper frames the study around participants, advice-makers, and manufacturers, while emphasizing that observed difficulties reflected systemic advice–device misalignment rather than participant skill alone.

  • Stakeholders: The study identified participants, policymakers and advice-makers, and IoT manufacturers as its three primary stakeholder groups.
  • Interpretation: The paper characterizes unsuccessful intended configurations as systemic misalignment between advice and device design rather than lack of participant skill.
  • Research design: The device-selection process did not confirm feature availability before testing, preserving whether users could independently determine advice applicability as an empirical question.Pre-screening only devices with confirmed features would have assumed the answer to the central research question.
  • Implications: The findings identify false-positive outcomes in which app-level proxy tasks can leave devices vulnerable while producing misplaced accomplishment.The paper connects this mismatch to proposed usability mandates for feature verifiability and diagnostic feedback.

Data Management and Demographics

Demographic data were collected by a recruitment agency and linked to randomized participant IDs, while researchers avoided collecting directly identifying information during laboratory sessions.

  • Demographic data covered age, gender, and education and were supplied as a de-identified dataset linked to randomized Participant IDs.
  • Researchers collected no direct demographic information during the laboratory sessions.
  • Qualitative quotations were stripped of contextual details that could enable participant re-identification.

Open Science

The study deposited de-identified session and analysis materials, scripts, and supporting artifacts, while withholding video recordings and transcripts.

  • Open Science: Video recordings and transcripts were not deposited, although the qualitative evidence matrix and codebook were to be added.
  • Open Science: The deposited session dataset contains all 168 sessions, endpoints, participant conclusions, task durations, and six NASA-TLX subscale ratings.Interpreting the outcome values requires the study’s endpoint vocabulary and demonstrate-not-complete protocol.
  • Open Science: The repository includes R analysis materials for descriptive counts and exploratory models, plus a Python session-allocation script.Device and advice-item order was randomized per participant, and the allocation script was rerun for participants P21–P28.
  • Open Science: The study materials are available at the stated OSF repository.
  • User-study materials: The update instructions directed participants to access the device, find Settings, and check the update status.
  • User-study materials: The user-study briefing presented government-derived advice on updating device software and changing default passwords.

E Analysis of Information Sources

The pre-study source review distinguished the advice card’s default-password criterion from broader device-authentication facilities and documented what interfaces and sources could locate.

  • Information-source analysis: Table 5 examined each device interface and 14 documentary sources for password and update information before data collection.Sources included manuals, quick-start guides, manufacturer websites, a Google AI-generated overview, Google results, and YouTube results.
  • Information-source analysis: No sampled product carried a credential satisfying the advice card’s definition of a preconfigured, simple, publicly known default password.“None located” records what the reviewed sources returned and does not establish that no such credential or pathway exists.
  • Information-source analysis: The audit identified an embedded web-server PIN that was manufacturer-set and unit-unique, while a camera account was disabled until the owner created credentials.
  • Information-source analysis: The device interface was examined alongside the documentary sources and was excluded from the documentary-source count.

F Exploratory Models of the Verified-Firmware Endpoint

Exploratory models examined whether session-level frustration, duration, and device characteristics were associated with reaching a verified firmware update, with substantial uncertainty limiting interpretation.

  • Endpoint and model specification: 38 of 84 update sessions reached a verified firmware update, the binary endpoint modeled in both exploratory specifications.
  • Exploratory results: Firmware sessions had median Frustration 10 and median duration 131 seconds, compared with higher values for sessions without an update.Companion-app sessions resembled firmware sessions, so the association was not distinguishable from zero when no-update sessions were excluded.
  • Limitations: Medians were calculated over sessions rather than participants, and NASA-TLX coefficients were associational because ratings followed each session.Password sessions were not modeled.
  • Endpoint and model specification: Panel A modeled session-level measures separately because they were collinear, while Panel B entered device indicators in full.The models used population-averaged estimates with standard errors clustered on participant.
  • Exploratory results: The widest Panel B interval spanned a factor of forty and included 1, so no Panel B coefficient was interpreted.Adding device increased between-participant variance from 3.85 to 7.59, consistent with partial device–participant confounding.
  • Robustness: Cluster-bootstrap intervals were [0.12, 0.44] for Frustration and [0.11, 0.51] for duration.

J Codes clustered into Themes and Subthemes

Table 12 clusters coding-framework codes into analytic themes and subthemes describing participants’ session activities. These themes were interpretively synthesized rather than counted as outcome categories.

  • J Codes clustered into Themes and Subthemes: Table 12 organizes codes into analytic themes and subthemes describing what participants did during the sessions.The displayed codes cover feature identification, navigation, support seeking, device contrasts, mental models, effort, emotions, and reactions to applying advice.
  • J Codes clustered into Themes and Subthemes: Themes and subthemes are not frequency-based outcome categories, and their frequencies or co-occurrence with session outcomes are not reported.The code labels are reproduced from the archived codebook, while the themes were constructed through interpretive synthesis.
  • J Codes clustered into Themes and Subthemes: Participants’ activities included identifying password or update targets, navigating device and app interfaces, seeking support, and comparing devices.The code clusters include direct routes, difficult searches, trial-and-error navigation, external information seeking, manuals, and device contrasts.
  • J Codes clustered into Themes and Subthemes: The result-oriented code clusters capture mental models, simulated real-world behavior, interaction quality, workload, frustration, perceived value, and lingering uncertainty.These subthemes describe participants’ experiences and reactions during the sessions, including fatigue, abandonment, skepticism, and mixed feelings about applying the advice.
Loading 2608.25225v1…