Source-linked AI summary

Art and the science of generative AI: A deeper dive

Ziv Epstein, Aaron Hertzmann, Laura Herman, Robert Mahari, Morgan R. Frank, Matthew Groh, Hope Schroeder, Amy Smith, Memo Akten, Jessica Fjeld, Hany Farid, Neil Leach, Alex Pentland, Olga Russakovsky

arXiv:2306.04141v1cs.AI

TL;DR

Generative AI can produce high-quality artistic media and is likely to reshape creative production, raising interdisciplinary questions about culture, law, labor, and media. This paper analyzes those questions across four themes and argues that generative AI is a new medium with distinct affordances, while identifying research directions for policy and beneficial uses.

  • Problem

    Generative AI’s expanding creative capabilities and reliance on human-made training data create poorly understood implications for creativity, authorship, ownership, labor, culture, and media policy.

  • Method

    The paper presents a large-scale interdisciplinary collaboration organized around four themes: culture and aesthetics, legal authorship, creative-work economics, and the media ecosystem.

  • Results

    The paper concludes that generative AI is not necessarily art’s demise but a new medium with distinct affordances that can reshape creative roles, labor models, and the media ecosystem.

  • Takeaways & Limitations

    The paper identifies research questions and directions intended to inform policy and beneficial uses of generative AI.

  • Takeaways & Limitations

    The paper does not cover several societal impacts, including model externalities such as climate effects and crowd-worker exploitation.

Abstract

from arXiv · show

A new class of tools, colloquially called generative AI, can produce high-quality artistic media for visual arts, concept art, music, fiction, literature, video, and animation. The generative capabilities of these tools are likely to fundamentally alter the creative processes by which creators formulate ideas and put them into production. As creativity is reimagined, so too may be many sectors of society. Understanding the impact of generative AI - and making policy decisions around it - requires new interdisciplinary scientific inquiry into culture, economics, law, algorithms, and the interaction of technology and creativity. We argue that generative AI is not the harbinger of art's demise, but rather is a new medium with its own distinct affordances. In this vein, we consider the impacts of this new medium on creators across four themes: aesthetics and culture, legal questions of ownership and credit, the future of creative work, and impacts on the contemporary media ecosystem. Across these themes, we highlight key research questions and directions to inform policy and beneficial uses of the technology.

1 Introduction

Generative AI can produce high-quality media across artistic domains and may fundamentally alter how creators develop and produce work. The paper frames it as a new medium while examining its cultural, legal, labor, and media-ecosystem implications.

  • Generative AI produces high-quality visual art, concept art, music, fiction, literature, video, and animation.Diffusion models synthesize images, while large language models produce prose and verse.
  • Its generative capabilities are likely to fundamentally alter how creators formulate ideas and put them into production.
  • The paper argues that generative AI is a new medium with distinct affordances rather than necessarily a harbinger of art’s demise.This framing parallels historical cases in which technological change recast creative roles and practices.
  • Generative AI may threaten existing jobs and labor models in the short term while enabling new models of creative labor and reconfiguring the media ecosystem.
  • The paper examines aesthetics and culture, ownership and credit, creative work, and the contemporary media ecosystem to identify research directions for policy and beneficial uses.

2 Perceptions and Conceptualization of Generative AI

The paper treats generative AI as a sociotechnical tool rather than a human-like agent, emphasizing how terminology, interfaces, and human control shape perceptions and responsibility. It uses meaningful human control to organize questions about intent, predictability, accountability, and agency.

  • Broad use of the term “artificial intelligence” spans algorithms with different computational complexity and problem domains.
  • Natural-language interfaces, the “I” pronoun, and terms such as “hallucination” can encourage users to attribute human-like intent, agency, or self-awareness to AI systems.
  • Anthropomorphizing AI can undermine credit to creators and deflect responsibility from developers and decision-makers when systems cause harm.
  • The paper discusses generative AI as a tool supporting human creators, not as an agent with its own intent or authorship.
  • Meaningful human control requires systems to incorporate human intent while supporting exploration, predictability, and progressively clearer goals.
  • Because generative AI systems are diffuse sociotechnical systems, research must examine how people perceive the interplay among human actors and computational processes.The paper also identifies collective systems that use feedback from many users as an important regime beyond direct prompting.

3 Shifts in Culture & Aesthetics

Generative AI’s training data, built-in aesthetics, and platform context shape both the diversity of artistic outputs and how viewers interpret intention and credit. The paper calls for stronger forms of human agency and research into output diversity, bias, and cultural effects.

  • Large computational infrastructure is controlled by a few companies, which can control access to and functionality of generative AI technology.
  • Generative tools can impose strong aesthetics, including hyperrealism, making creators’ intentions and individuality difficult for casual viewers to identify.
  • Creators may need to express artistic intention through training-data selection, prompt construction, and downstream uses of generated artifacts.
  • Training-data biases can be reflected or amplified in outputs, while web-derived data may reproduce racial, gender, and geographic inequalities.
  • AI-generated content may enter future training sets, creating a self-referential flywheel that reinforces generative AI’s aesthetics and biases.
  • Knowledge that an artifact was AI-generated can influence viewers’ perceptions, while anthropomorphism may shift perceived credit from human artists toward technology creators.
  • Generative AI operates within an attention-economy media ecosystem where opaque recommendation algorithms shape content production and consumption.

Legal Dimensions of Authorship

Generative AI raises distinct legal questions about training data and ownership of outputs, requiring technical, social-scientific, and legal analysis. The paper surveys competing approaches involving permission, licensing, compensation, similarity, and creative contribution.

  • Generative AI’s reliance on training data creates legal and ethical authorship challenges involving both the data used to train models and their outputs.
  • Proposals for training-data use range from treating it as non-infringing or fair use to requiring creator licenses, opt-outs, or compulsory compensation schemes.
  • Key research questions concern copyright infringement, direct copying versus novel creation, protection of individual styles, and mechanisms for artist compensation or opt-out.
  • Resolving training-data questions requires technical research on systems, social-science research on perceived similarity, and legal research applying precedents to novel technology.
  • The paper’s discussion of these legal questions represents an American legal perspective, and coverage ultimately depends on jurisdiction.
  • Ownership of outputs depends partly on the creative contributions of users, developers, and training-data creators.
  • Deliberate emulation of existing works may produce derivative outputs, while proposed responses include compulsory licenses or joint ownership for prompt artists.

4 The Labor Economics of Creative Work

Generative AI may displace some creative work while increasing productivity, accelerating ideation, and creating new occupations. Existing labor frameworks are insufficiently precise for tracing how specific creative-process steps and workplace requirements will change.

  • Labor framework: A new framework is needed to map creative-process steps to generative AI’s effects on workplace requirements and activities.The paper proposes examining which steps are affected and how those effects vary across cognitive occupations.
  • Human-in-the-loop work: Human-in-the-loop interactive paradigms could advance worker productivity while identifying opportunities for tools that better complement workers.These paradigms connect labor research with the design of future creative systems.
  • Productivity and creativity: Generative AI can produce hundreds of outputs per minute, accelerating ideation and potentially reducing production time and costs.The same passage also notes that rapid ideation may undermine prototyping and envisioning associated with a tabula rasa.
  • Employment effects: Although some occupations may be threatened, generative AI could increase productivity in others and perhaps create new ones.The paper gives historical examples in which automation enabled more musicians to create, even as earnings became more uneven.

5 Impacts on the Media Ecosystem

Generative AI may make synthetic media cheaper and more abundant, increasing risks to trust, attention, and information integrity. The paper discusses provenance, authentication, forensic detection, and misinformation research as complementary responses, while noting important vulnerabilities and scope limits.

  • Downstream harms: Lower costs and shorter production times for synthetic media may increase vulnerability to impersonation, manipulation, disinformation, distraction, fraud, and nonconsensual sexual imagery.Photorealistic synthetic media may also undermine trust in authentic media through the liar’s dividend.
  • Provenance and authentication: Provenance and authentication tools can help preserve information integrity, but digital signatures require widespread adoption across the media ecosystem.The C2PA protocol cryptographically binds media to provenance metadata, while visual watermarks can be cropped or manipulated.
  • Forensic detection: Post-hoc forensic methods can detect statistical and physical artifacts of manipulation, but they are vulnerable to adversarial attacks and context shift.These approaches passively identify artifacts left by manipulated visual and auditory content.
  • Misinformation: Photographs can slightly increase susceptibility to fake news headlines without directly evidencing the headlines’ claims, while synthetic visual components may reveal artificial origins.The paper presents this as preliminary evidence about how visual media affects misinformation susceptibility and detection.
  • Scope boundary: The paper leaves the broader information effects of fluent but unverified or ideological LLM-generated writing beyond its scope and calls for future research.Provenance and watermarking may help mitigate these issues, but their impact requires further diagnosis.
  • Attention and collective action: The proliferation of AI-generated information may decrease collective attention and hamper discussion and action on issues such as climate and democracy.The paper connects this concern to evidence that more available information can reduce collective attention spans.

Discussion and Limitations

The paper identifies important societal impacts that it does not comprehensively cover and proposes future research on creativity, including algorithmic methods for embedding improvisation in AI systems.

  • Limitations: The discussion does not cover several societal impacts, including climate change, crowd-worker exploitation, and concentration of market power.The authors call for better carbon-impact measurement, scrutiny of energy use, and comprehensive mapping of additional harms.
  • Future research: Generative models are trained by reducing reconstruction error on training data, which bounds them by reproducing what they have already seen.The paper identifies novel algorithmic methods for embedding improvisation as a key future direction.

Research Questions Raised by Generative AI

The paper frames generative AI as a subject for interdisciplinary research and collective governance, with research questions spanning technology, society, and beneficial use. It also points toward future systems that interact more closely with human creativity.

  • Human creativity: Future systems could use generative AI earlier in workflows for speculation and idea generation or explicitly interact with distinct modes of human creativity.These directions concern how system design can structure collaboration with human creators.
  • Governance: The uses and impacts of generative AI will be shaped by collective decisions made by developers, users, regulators, and civil society.The authors emphasize artists and creative laborers as particularly important stakeholders.

Contributions and Position

The paper is a 14-author academic collaboration organized around four themes concerning generative AI’s effects on culture, authorship, creative work, and the media ecosystem. It also discloses relevant professional affiliations and prior consulting relationships.

  • The paper brings together 14 authors with expertise spanning four themes: culture and aesthetics, authorship, creative labor economics, and the media ecosystem.
  • The collaboration was initiated and conceptualized at the 2023 International Conference of Computational Creativity by four named researchers.
  • Two authors work for Adobe, a company that makes generative AI tools, while three authors had previously consulted for OpenAI by red teaming its systems.
Loading 2306.04141v1…