Source-linked AI summary

Remote Collaboration Fuses Fewer Breakthrough Ideas

Yiling Lin, Carl Benedikt Frey, Lingfei Wu

arXiv:2206.01878v4cs.CYecon.GN

TL;DR

The paper compares how scientists contribute across onsite and remote teams and measures inventive contributions using article and patent data. Across repeated collaborations, remote teams show lower participation in conceptual research tasks, and this difference remains after accounting for several team and individual characteristics.

  • Problem

    The paper investigates whether differences between onsite and remote collaboration contexts are associated with how scientists contribute to research and inventive output.

  • Method

    The study combines disambiguated patent and geographic data, classifies fields using patent taxonomies, and measures disruption with a D-score based on citation patterns.

  • Results

    Remote teams show lower participation in conceiving research, including a decline from 33.5% to 28.3% for jointly contributing scientists, robust across fields, periods, and team sizes.

  • Takeaways & Limitations

    Remote teams’ reduced conceptual engagement is not fully explained by team size, expertise composition, weak ties, age, or individual selection differences.

Abstract

from arXiv · show

Theories of innovation emphasize the role of social networks and teams as facilitators of breakthrough discoveries. Around the world, scientists and inventors today are more plentiful and interconnected than ever before. But while there are more people making discoveries, and more ideas that can be reconfigured in novel ways, research suggests that new ideas are getting harder to find-contradicting recombinant growth theory. In this paper, we shed new light on this apparent puzzle. Analyzing 20 million research articles and 4 million patent applications across the globe over the past half-century, we begin by documenting the rise of remote collaboration across cities, underlining the growing interconnectedness of scientists and inventors globally. We further show that across all fields, periods, and team sizes, researchers in these remote teams are consistently less likely to make breakthrough discoveries relative to their onsite counterparts. Creating a dataset that allows us to explore the division of labor in knowledge production within teams and across space, we find that among distributed team members, collaboration centers on late-stage, technical tasks involving more codified knowledge. Yet they are less likely to join forces in conceptual tasks-such as conceiving new ideas and designing research-when knowledge is tacit. We conclude that despite striking improvements in digital technology in recent years, remote teams are less likely to integrate the knowledge of their members to produce new, disruptive ideas.

A. Research design summary

The study compares onsite and remote teams using large, geographically mapped datasets of scientific papers and patents. Teams are classified by city-level collaboration distance, with multiple checks supporting dataset quality and robustness.

  • Team classification: Onsite teams have members in the same city, whereas remote teams have members distributed across two or more cities.The study calculates collaboration distance from city locations rather than institution locations.
  • Team classification: Among papers, 68% of authors are in the same city and 32% are distributed across cities.Among remote teams, 60% of papers involve collaboration distances above 600km.
  • Data quality: The patent dataset uses name-disambiguated inventors and residences from PatentsView, with extensive manual verification of inventor and location assignments.The verified patent data cover 87,937 unique cities.
  • Robustness: The study tests onsite-versus-remote differences across alternative city-distance thresholds of 0km, 1km, 5km, and 10km.These thresholds are applied after scientists are mapped to cities.

C. Calculating D-scores

The paper measures disruption by comparing how later work cites a focal work and its references. Positive scores indicate disruption, negative scores indicate development, and zero indicates balance.

  • D-score calculation: The D-score is calculated as D = p_i - p_j = (n_i - n_j) / (n_i + n_j + n_k).It quantifies the divergence between later papers that cite only the focal work and those that cite both the focal work and its references.
  • D-score interpretation: p_i is the proportion of later papers citing only the focal work, while p_j is the proportion citing both the focal work and its references.A third category cites only the focal work’s references.
  • D-score interpretation: Scores from 0 < D < 1 indicate disruptive discoveries, whereas scores from -1 < D < 0 indicate developing contributions.D = 0 represents a balance between disruptive and developmental characteristics.

D. Quantifying timezone differences underlying inventive teamwork

The study measures temporal separation within teams by mapping members to city time zones and averaging pairwise hour differences.

  • Timezone measurement: For 20,273,444 papers and 3,709,940 patents, each team member’s time zone is mapped from the latitude and longitude of their city.The study uses PYTZ for the time-zone mapping.
  • Timezone measurement: A team’s temporal separation is the average time-zone difference across all pairs of team members.This average serves as a proxy for the underlying temporal separation of the team.

E. Identifying the fields of study for research articles and patent applications

The study assigns broad field labels to papers and patents using established hierarchical taxonomies. Papers use the highest-confidence MAG level-zero label, while patents use the most popular CPC section label.

  • Research articles: Scientific papers are classified with MAG’s six-level taxonomy, whose level-zero contains 19 research fields.When a paper has multiple labels, the analysis selects the label with the highest confidence at level zero.
  • Patent applications: Patent applications are classified with the four-level Cooperative Patent Classification system, whose level zero contains nine sections.The analysis assigns labels to level-zero sections and selects the most popular section.

F. Quantifying knowledge diversity of inventive teamwork

The study measures knowledge diversity through team-member interdisciplinarity, using it as a proxy for the diversity of knowledge available to teams and as a control for team heterogeneity.

  • Team-member interdisciplinarity serves as a proxy for the diversity of knowledge accessible to the team.
  • The measure allows the regressions to account for team heterogeneity.
  • Scientists’ home disciplines are identified from 19 top-level MAG field-of-study labels when they have published at least three papers.

G. Quantifying tie strength within research and innovation teamwork

The study constructs a large co-authoring network and defines tie strength by the extent to which two scientists share collaborators.

  • The social network contains 22,566,650 scientists and 67,226,924 co-authoring relationships.
  • Tie strength is calculated as the ratio of common collaborators to total collaborators.
  • The measure defines stronger ties as greater overlap in scientists’ collaborators.

H. Evaluating the robust, negative relationship between remote teams and disruption

The study evaluates the relationship between remote collaboration and disruption using regression models built from a large longitudinal dataset of scientists and papers.

  • The analysis uses regressions to evaluate the negative relationship between remote teams and disruption.
  • 7,681,669 scientists who published at least two papers were selected for the scientific-team dataset.
  • These scientists produced 13,711,470 papers and 45,078,179 paper-author records between 1960 and 2020.

I. Identifying author contributions to scientific papers

The study analyzes author-contribution disclosures from four journals to identify functional research activities using natural language processing.

  • 89,575 contribution disclosures were collected from four journals between 2003 and 2020.
  • The journals are Nature, Science, PNAS, and PLOS ONE.
  • Natural language processing identifies four activities: conceiving research, writing the paper, performing experiments, and analyzing data.

J. Evaluating the robust, negative relationship between remote teams and conceiving research

The same scientists engage less in conceiving research when they collaborate remotely, and this role shift remains robust across team sizes, periods, and fields. Pair- and group-level analyses, supported by machine-learning inference on millions of papers, reinforce the reduced conceptual engagement of remote teams.

  • 63% to 51%: the same scientists’ probability of contributing to “conceiving research” falls when switching from onsite to remote teams.The decline is statistically significant (p-value < 0.001).
  • The reduced engagement in conceptual tasks cannot be explained by larger team sizes, different research fields, or different time periods.
  • 33.5% to 28.3%: collaborating scientist pairs become less likely to jointly contribute to “conceiving research” in remote teams.The reduction remains significant after accounting for fields, periods, and team sizes.
  • 21.6% to 17.8%: groups of three or more scientists show a lower probability of contributing to “conceiving research” remotely.This difference is statistically significant (p-value < 0.01).
  • A neural network infers conceptual and technical author roles across 16,397,750 papers to test whether the role difference generalizes beyond explicit contribution disclosures.

K. Examining alternative explanations for the reduced disruption of remote teams

The lower disruption associated with remote teams is not fully explained by team size, composition, career age, selection, or weak ties. Although remote teams contain more weak ties and have grown faster in size, the negative relationship with disruption remains after these factors are considered.

  • Accounting for team size and periods does not alter the negative remote-team coefficient, despite faster team-size growth among remote teams.In papers, remote-team size increased 100% versus 65% for onsite teams; in patents, it increased 40% versus 32%.
  • Differences in team heterogeneity are unlikely to explain the observed disruption gap between remote and onsite teams.
  • Controlling for career age leaves the negative impact of remote teams on disruption unchanged, despite onsite members having lower average career ages.Average career age is 9.6 for onsite members versus 11.8 for remote members.
  • The same scientists act differently across team contexts, so individual selection differences cannot fully explain the observed disruption gap.
  • Remote teams include more weak ties, but the negative relationship between distance and disruption remains after weak-tie collaboration is included.The authors conclude that remote teams do not exchange, fuse, and integrate the diverse knowledge available through these ties sufficiently to generate disruptive ideas.

Data availability

The supplied passages document the paper’s datasets, robustness checks, contribution-role analyses, statistical conventions, and data-access information. Together they cover large-scale science and patent samples, alternative distance and innovation measures, author-role inference, and regression specifications.

  • The paper’s datasets are available through the listed project website and Figshare repository.
  • The main science and technology analyses use 20,134,803 papers and 4,060,564 patent applications across the stated publication and filing periods.
  • Alternative distance measures include maximum team-member distance, average distance between unique cities, and a continuous colocation index.
  • Alternative innovation measures include D-score percentiles and the probabilities of proposing new scientific concepts or introducing new technology codes.
  • Author-contribution analyses compare conceptual and technical activities using disclosures from PNAS, Nature, Science, and PLOS ONE.
  • A neural-network model infers conceptual and technical author roles for papers without explicit contribution information.
  • The robustness tables assess disruption in science and technology and reduced conceptual engagement using large author and author-pair samples.
  • Reported tests are two-sided t-tests, with clustered standard errors specified at author, inventor, or author-pair levels depending on the model.
Loading 2206.01878v4…