Source-linked AI summary

Diffusion TV: Experiencing Diffusion Models through Tangible, Embodied Interaction

Sihwa Park

arXiv:2609.05404v1cs.HCcs.AI

TL;DR

Diffusion TV addresses the limited public exposure of diffusion models’ intermediate states by creating an embodied, tangible alternative to conventional technical explanation. It maps CRT antenna and knob interactions to audiovisual denoising and channel selection, and reports that audiences generally found the interaction intuitive and engaging while using it to interpret the generative process thematically. The authors caution that this experiential engagement does not guarantee precise comprehension and call for structured audience studies.

  • Problem

    Diffusion models’ intermediate states are rarely exposed to broader audiences, limiting opportunities for interactive or embodied exploration.

  • Method

    Diffusion TV uses a modified CRT television to map antenna and knob interactions to audiovisual denoising stages and animal channels.

  • Results

    Audiences generally learned the interaction through exploration and perceived the antenna-based denoising interaction as intuitive and engaging, supporting thematic interpretation of the generative process.

  • Takeaways & Limitations

    Embodied, metaphorical engagement can serve as an alternative mode of explainable AI beyond explicit technical explanation in conventional 2D GUIs.

  • Takeaways & Limitations

    Tangible interaction with denoising supports intuitive and experiential engagement but does not guarantee precise comprehension, and the work relies on informal feedback and qualitative reflection rather than systematic evaluation.

Abstract

from arXiv · show

Diffusion TV is an interactive AI art installation that offers a tangible and embodied experience of diffusion models through a modified CRT TV. By physically manipulating the TV's antenna, audiences control the clarity of AI-generated images and sounds, metaphorically enacting the denoising process that underlies diffusion-based generation. Using the tuning knob, participants switch between three channels featuring AI-generated animals from the Past (extinct species), Present (endangered species), and Future (speculative creatures), situating the interaction within a temporal and ecological narrative. Through continuous audiovisual feedback and physical interaction, Diffusion TV foregrounds the generative process over final outputs, allowing audiences to explore intermediate states as experiential material. Rather than providing explicit technical explanation, the work presents an alternative, embodied mode of explainable AI that invites exploratory engagement with and reflection on generative technologies.

1 Introduction and Background

Diffusion TV addresses the limited public exposure of diffusion models’ intermediate states by turning denoising into an embodied, sensory interaction. The installation offers an alternative to conventional technical explanation through tangible engagement with generative processes.

  • Motivation: Diffusion models generate structured outputs through an iterative transformation from noise, but their intermediate states are rarely exposed to broader audiences.Outside research contexts, these states are typically shown as pre-rendered or static visualizations, limiting interactive or embodied exploration.
  • Related work: The approach builds on artistic, tangible, and experiential forms of explainable AI that mediate between opaque computational systems and human understanding.These practices support engagement with AI systems intellectually, perceptually, and emotionally.
  • Contribution: Diffusion TV is an interactive AI art installation that explores complex AI mechanisms through embodied interaction with a modified CRT television.The installation maps physical controls to image and sound generation stages.
  • Contribution: The antenna and knob map audience actions to intermediate stages of diffusion-based image and sound generation.This mapping reframes denoising as a performative and exploratory act rather than only a technical process.
  • Contribution: Rather than prioritizing precise technical explanation, the work emphasizes intuitive sensory engagement with generative processes.It invites reflection on disappearance, preservation, and relationships among humanity, technology, and the environment.

2 Diffusion TV Design

Diffusion TV uses a nostalgic CRT interface to make denoising physically manipulable and organizes its content across Past, Present, and Future animal channels. Antenna rotation controls audiovisual clarity and continuous progression through generated content.

  • 2 Diffusion TV Design: Figure 1 depicts the interaction sequence from a noisy animal image through intermediate denoising to a clear image with additional information.Without rotation input, the system eventually displays the noisiest image of the next animal.
  • 2 Diffusion TV Design: The CRT television’s tuning knobs and antenna serve as a conceptual and physical interface for the denoising process.These familiar broadcast metaphors let audiences engage with generative transformation through physical interaction.
  • 2 Diffusion TV Design: Rotating the antenna changes the visual and auditory clarity of an AI-generated animal across intermediate stages from noise to structured output.After a sequence reaches clarity, the system resets and introduces a new animal.
  • 2 Diffusion TV Design: The VHF tuner knob selects three channels: Past, Present, and Future.Past presents extinct animals, while Present focuses on endangered animals and Future features speculative creatures.
  • 2 Diffusion TV Design: At maximum clarity, broadcast-style lower-third information provides minimal identifiers, temporal references, or speculative descriptors.The system can also regenerate animal images and sounds at configurable intervals so repeated visits produce different audiovisual experiences.

3 Animal Data

The installation combines sourced historical and conservation data with manually designed prompts and AI-generated speculative species. Its animal dataset supports the Past, Present, and Future channel structure.

  • 3 Animal Data: The Past and Present channels use conservation sources to identify sixteen extinct and fourteen endangered species.The dataset includes contextual information such as last-seen year, habitat, and estimated population size.
  • 3 Animal Data: Image and sound generation prompts were manually designed and iteratively refined for each sourced species.The prompts were developed alongside the animal dataset and contextual information.
  • 3 Animal Data: The Future channel contains twelve speculative species generated with ChatGPT o3.Each includes a name, emergence year, environmental trigger, short biography, and prompts for image and sound generation.

4 Technical Overview

Diffusion TV combines a modified CRT, Raspberry Pi, sensors, pre-generated audiovisual media, and custom software to map physical controls onto denoising stages and channel content.

  • 4 Technical Overview: A modified CRT TV, Raspberry Pi 5, sensor system, and HDMI-to-RF modulator deliver generated audiovisual content through the television.Figure 2 presents the installation’s overall schematic.
  • 4 Technical Overview: Stable Diffusion XL and Stable Audio Open generate pre-rendered animal images and sounds, with custom callbacks extracting intermediate denoising outputs.The outputs are stored for access by the client.
  • 4 Technical Overview: Antenna rotation is captured through a mechanically linked rotary encoder and mapped to an index selecting corresponding image and sound pairs from different denoising stages.This connects continuous physical movement to discrete audiovisual content states.
  • 4 Technical Overview: A magnetic encoder attached to the VHF tuner detects channel selection among the Past, Present, and Future content sets.The channel state determines which generated content collection is available.
  • 4 Technical Overview: The Processing client dynamically updates audiovisual output according to antenna position and channel state.It can asynchronously incorporate newly generated content without interrupting interaction.

5 Exhibition and Discussion

The exhibition showed that audiences could learn Diffusion TV through exploration, while the antenna interaction was generally intuitive and engaging. The tuning metaphor supported experiential engagement with the generative process, but responses varied with prior familiarity and some participants did not explore all content.

  • The installation was exhibited at the International Symposium on Electronic/Emerging Art 2025 in Seoul from May 23 to 29, 2025.
  • The reflections were based on exhibition observations and video recordings rather than a structured empirical evaluation.
  • Audiences typically learned the interaction through exploration, with prior familiarity with CRT TVs shaping engagement.Older participants navigated intuitively, while younger viewers sometimes struggled with the interface.
  • The antenna-based denoising interaction was generally perceived as intuitive and engaging.
  • Some participants did not fully explore sequential content or expected additional system behaviors.
  • The tuning metaphor supported experiential engagement with the generative process and encouraged thematic interpretation without explicit instruction.

6 Conclusion and Future Work

Diffusion TV presents embodied, metaphorical engagement as an alternative mode of explainable AI that moves beyond explicit technical explanation. The authors also note that tangible interaction does not ensure precise comprehension and propose structured audience studies and more responsive generative pipelines as future work.

  • 6 Conclusion: Diffusion TV demonstrates the potential of embodied, metaphorical engagement as an alternative mode of explainable AI.
  • 6 Conclusion: The tangible interaction moves beyond explicit technical explanation in conventional 2D GUI-based approaches.
  • 6 Conclusion: Tangible denoising interaction supports intuitive and experiential engagement but does not guarantee precise comprehension of diffusion models.
  • 6 Conclusion: Audience responses varied depending on prior familiarity with AI.
  • 6 Conclusion: The observations relied on informal feedback and qualitative reflection rather than systematic evaluation.
  • 6 Future Work: Future work will use structured audience studies to examine how embodied interaction shapes public interpretation of generative AI systems.
  • 6 Future Work: The system will be extended toward real-time or hybrid inference pipelines to maintain responsive interaction and enable more flexible generative behavior.
Loading 2609.05404v1…