Source-linked AI summary

Distribution of Cognitive Load in Web Search

Jacek Gwizdka

arXiv:1005.1340v2cs.HCcs.IR

TL;DR

Web-search cognitive load varies dynamically with task stages and interface demands, but task-level averages can obscure those changes. Using a controlled study of 48 participants, semi-automatic log-based stage segmentation, and a dual-task measure, the paper found higher load during query formulation and relevant-document tagging than during results-list examination and document viewing, while semantic result categories reduced demands in some stages.

  • Problem

    Task-level cognitive-load averages can obscure dynamic changes, limiting understanding of when search tasks and interfaces impose mental demands.

  • Method

    A controlled web-search experiment with 48 participants used interaction logs to segment tasks into stages and a dual-task method to assess cognitive load.

  • Results

    Cognitive load was significantly higher during query formulation and relevant-document tagging than during results-list examination and document viewing.

  • Takeaways & Limitations

    Dynamic task-stage assessment can detect changes in cognitive load and clarify demands imposed by search tasks and interactive systems.

  • Takeaways & Limitations

    The dual-task method does not separately measure germane cognitive load.

Abstract

from arXiv · show

The search task and the system both affect the demand on cognitive resources during information search. In some situations, the demands may become too high for a person. This article has a three-fold goal. First, it presents and critiques methods to measure cognitive load. Second, it explores the distribution of load across search task stages. Finally, it seeks to improve our understanding of factors affecting cognitive load levels in information search. To this end, a controlled Web search experiment with forty-eight participants was conducted. Interaction logs were used to segment search tasks semi-automatically into task stages. Cognitive load was assessed using a new variant of the dual-task method. Average cognitive load was found to vary by search task stages. It was significantly higher during query formulation and user description of a relevant document as compared to examining search results and viewing individual documents. Semantic information shown next to the search results lists in one of the studied interfaces was found to decrease mental demands during query formulation and examination of the search results list. These findings demonstrate that changes in dynamic cognitive load can be detected within search tasks. Dynamic assessment of cognitive load is of core interest to information science because it enriches our understanding of cognitive demands imposed on people engaged in the search process by a task and the interactive information retrieval system employed.

Introduction

Web search engages perception and cognition, with demands shaped by the task, system, and searcher. The article examines how cognitive load varies during search and how interfaces affect those demands.

  • Search behavior is affected by task, system, and individual searcher characteristics.
  • Understanding cognitive load can identify which search tasks and interface features place greater demands on users.
  • The article presents and critiques cognitive-load measurement methods, examines load across task stages, and studies factors affecting load levels.
  • The article emphasizes cognitive load’s dynamic properties and their implications for measurement in information science.

Background and Related Work

Cognitive load reflects limited mental resources and can be assessed through subjective, performance, and physiological methods. Prior work indicates that the unit used to average dynamic measures can determine whether task-stage differences are detected.

  • The Concept of Cognitive Load: Cognitive load concerns demands on limited mental resources relative to a task and a person’s capabilities.
  • The Concept of Cognitive Load: Intrinsic load comes from the problem, extraneous load from the task environment, and germane load from intentional learning effort.
  • Measurement Techniques: Cognitive-load assessment includes subjective, performance, and physiological measures, each capturing different aspects of load.
  • Measurement Techniques: Subjective measures are usually collected after a task, making them unsuitable for detecting dynamic changes during performance.
  • Measurement Techniques: Dual-task and physiological methods support dynamic, real-time measurement, while dual-task assessment offers an inexpensive objective measure of effort.
  • Related Work: Prior findings suggest that task-level averaging often misses differences that emerge when cognitive load is averaged at the task-stage level.

Research Motivation and Objectives

The paper motivates finer-grained cognitive-load analysis because averaging can obscure changes during search. It focuses on peak and average load at task-stage level to clarify mental demands across search activities and interfaces.

  • Different measurement methods and averaging units capture different aspects of cognitive load and should generally complement rather than directly replace one another.
  • The study examines peak and average cognitive load at the task-stage level to characterize dynamic load patterns.
  • Peak-load analysis is intended to identify situations where cognitive capacity may be exceeded and task performance may be degraded.
  • Understanding cognitive-load dynamics can illuminate search processes at the sub-task level and the effects of interfaces on mental effort.

Method

The controlled study involved 48 participants completing varied web-search tasks in a university laboratory. Tasks differed in type, structure, and objective difficulty, and interaction logs supported task-stage analysis.

  • Participants: 48 participants completed a question-driven web-search study in a controlled experimental setting.
  • Procedure: Each session lasted 1.5–2 hours and included practice, questionnaires, six search tasks, and post-session assessment.
  • User Tasks: The study used Fact Finding and Information Gathering tasks, with structures including Simple, Hierarchical, and Parallel information needs.
  • User Tasks: Tasks were assigned low, medium, or high objective difficulty based on desired outcomes and a priori determinability.
  • User Tasks: Each participant performed six tasks of differing type and structure, yielding 288 tasks across the study.
  • User Tasks: Task orders were balanced using sequences that increased or decreased in objective difficulty.

User Interfaces

The study compares two Wikipedia search interfaces: familiar Google search and unfamiliar ALVIS search, which adds category information to result listings. A Stroop-based secondary task measures cognitive load during the primary search task.

  • Search interfaces: UI1 was Google’s Wikipedia search, while UI2 was the unfamiliar ALVIS Wikipedia search.Both interfaces displayed search results in a list; ALVIS additionally displayed categories beside and beneath each result.
  • Search interfaces: ALVIS added categories to search results, shown on the left side of the list and beneath each result.
  • Secondary task: The secondary task used a Stroop-based color-word pop-up to obtain indirect objective measures of cognitive load on the primary task.The pop-up appeared at random intervals and remained visible for a random period.
  • Secondary task: The pop-up presented a color name and colored font that could either match or not match.

Independent Factors (IF)

The experiment treated task characteristics, search interface, and user cognitive abilities as independent factors. User abilities were assessed through operation-span and mental-rotation measures, while participant characteristics were also recorded.

  • Task and interface factors: Independent factors included objective task difficulty and search interface.Objective task difficulty was classified as low, medium, or high; the interfaces were Google and ALVIS Wikipedia search.
  • User characteristics: Participants were tested for operation span and mental rotation because these cognitive abilities were expected to affect search-task performance.
  • User characteristics: Operation span was measured as a ratio from 0–100%, with higher scores indicating higher ability.
  • User characteristics: Mental rotation was measured using mean reaction time and the ratio of correct responses.
  • Participant characteristics: Five of forty-eight participants, or 10%, were left-handed, and the small number of left-handed participants did not significantly affect cognitive-load results.

Dependent Variables

The study measured cognitive load objectively through secondary-task performance and also collected selected self-reported task assessments. Reaction time was more sensitive than missed-event rate for the objective dual-task measure.

  • Objective measures: Objective cognitive load used secondary-task miss rate and reaction time, with reaction time found more sensitive than miss rate.These measures were used to assess average cognitive load over a task-stage duration.
  • Subjective measures: Subjective variables included pre-task familiarity with and interest in the topic and post-task assessment of task difficulty.
  • Measurement scope: The paper analyzes cognitive load at the task-stage level rather than focusing on the task-level subjective variables examined in prior work.

Task Stage Segmentation

Interaction logs were segmented into search-task stages using observable actions, visited-page types, and a state-machine procedure. The resulting stages represented query formulation, result-list examination, document examination, and saving or describing relevant results.

  • Segmentation framework: Task segments were based on selected sub-processes from Marchionini’s information seeking process.
  • Page classification: Visited pages were classified as search-engine home, search-results list, individual result page, or bookmarking and tagging page.Unknown URLs were generally categorized as other content pages.
  • Segmentation limitation: Query-formulation boundaries were less directly observable because this stage was segmented mainly from keyboard activity.The approach assumed that captured interaction logs were sufficient to identify task stages.
  • Task stages: Search-result-list examination involved assessing results in response to an entered query.
  • Task stages: Individual-result examination involved visually scanning or reading a visited content page.
  • Task stages: Relevant content documents were saved through bookmarking and tagging after information extraction and relevance decisions.
  • Segmentation procedure: The segmentation state machine used URLs, active program windows, keyboard activity, and mouse activity to process logged events.A ten-event history window was used to detect text finding through Ctrl-F sequences.

Results

Cognitive load varied across search task stages, with the highest average demands during query formulation and bookmarking. User-interface effects depended on task stage and secondary-task condition, while residual reaction times isolated stage-related differences from individual variability.

  • Average cognitive load per task stage: Task stage significantly affected miss rate, with B higher than C and borderline higher than L.Miss rate was less sensitive than reaction time overall, but the stage effect remained significant (F(3,900)=3.53, p<.05).
  • Individual variability and residuals: Stage-related reaction-time differences remained after separating participant effects: residual reaction time was related to task stage but not cognitive abilities.Participant characteristics explained RTperson, whereas RTtask_stage+C captured stage-related differences.
  • Average cognitive load per task stage: Query formulation (Q) and bookmarking (B) produced longer secondary-task reaction times than viewing content (C), with Q also exceeding results-list examination (L).For residual reaction time, Q exceeded L (p=.007) and C (p<.001), while B exceeded C (p=.001).
  • User interface: User-interface effects reversed across secondary-task conditions: Alvis increased load for color-match cases, whereas Google produced longer times for no-color-match cases.Color-match differences arose mainly at L, while no-color-match differences arose at Q and L.

Discussion

Average cognitive load was higher during query formulation and bookmarking than during results-list or document viewing, while peak load was higher during document viewing than results-list examination. Interface complexity also altered load differently across secondary-task conditions.

  • Average cognitive load: Query formulation and bookmarking imposed higher average cognitive load than examining results lists or viewing individual documents.The paper links this pattern to producing query terms or tags, which requires recall and additional working-memory support; results and documents rely more on recognition.
  • Peak cognitive load: Peak cognitive load was higher during viewing individual documents than during examining search-result lists.The paper plausibly relates document-viewing peaks to longer focused reading episodes, despite low average load in both stages.
  • Secondary-task conditions: Task stage and objective task difficulty affected no-color-match cases, whereas color-match cases showed no such task-stage and difficulty effect.The authors associate the conditions with different sensitivities to interface-related and task-related load, with working memory and mental rotation implicated in no-color-match performance.
  • User-interface effects: Search interfaces affected cognitive load in opposite directions across color-match and no-color-match conditions.Alvis increased effort during results-list examination for color-match cases, while Google produced longer times during query and results-list stages for no-color-match cases.

Conclusions

The study shows that cognitive load varies dynamically across search-task stages and that the dual-task method can detect these changes. It also identifies methodological and study-design limitations that constrain interpretation.

  • Method: The Stroop-based dual-task method assessed cognitive load during web search tasks.The method was used to characterize users’ cognitive load throughout the search process.
  • Findings: Cognitive load was significantly higher during query formulation and relevant-document tagging than during result examination and individual-document viewing.Semantic result categories in Alvis reduced mental demands during query formulation and result-list examination.
  • Measurement: Average cognitive-load measures detected changes between task stages even when they did not detect differences between tasks.This supports measuring load at a finer temporal granularity.
  • Limitations: The secondary-task method may separate extraneous interface-related load from intrinsic task-related load, but it does not separately measure germane load.The method therefore captures only some components of cognitive load.
  • Method: Semi-automatic task-stage classification achieved about 95% accuracy, but its rules were derived from the studied search engines and websites.The rules could be modified for other cases.
  • Limitations: The study included a relatively large proportion of participant-rated easy tasks and unequal familiarity with the two interfaces.These limitations constrain task-difficulty and interface-effect interpretations, although stage-level analysis partly mitigates the former.
  • Implications: Measuring mental effort by task stage can inform search-system design and indicate when users may have spare capacity for additional interaction.The paper gives relevance feedback and notification delivery as potential applications.
  • Implications: Dynamic, on-task assessment is needed because task-level cognitive-load measures cannot show when increased mental demands occur during search.Sequences of load changes across task phases provide a fuller picture of cognitive demands.

Appendix A

The appendix provides an example information-search task about choosing between conventional home heating and solar panels. It also identifies a table containing one example task for each task-type and structure combination.

  • Task design: The study’s example search-task table contains one task for each combination of task type and structure.The table is presented as Table 19.
  • Example task: The example task asks participants to research issues relevant to choosing conventional home heating or solar panels.The scenario involves helping friends who know little about home heating decide between the alternatives.
  • Appendix contents: The appendix combines a concrete home-heating research scenario with a structured set of task examples.Together, the passages show both the content context and the task-design organization.

Appendix B

Figure 7 presents pseudocode for the simplified task-segmentation algorithm.

  • Figure: Figure 7 presents pseudocode for the simplified task-segmentation algorithm.The passage identifies the figure’s subject but does not describe its individual steps.
  • Figure: The figure documents a simplified algorithm for segmenting search tasks.Its format is pseudocode rather than a narrative description.
  • Figure: Figure 7 serves as a procedural representation of task segmentation.The passage does not specify the algorithm’s inputs, outputs, or rules.
Loading 1005.1340v2…