Source-linked AI summary
Govern the Model, Not Only the Data: Storage, Circulation, and Learning in Creative AI
Phoenix Perry, George Simms, Elizabeth Wilson, Yasmine Boudiaf, Nick Bryan-Kinns, Tega Brain, R. Luke DuBois, Alix Rule, Rachel Meade Smith, Kelani Nichole, Atharva Pravin Pawar, Rebecca Fiebrink
TL;DR
The paper examines whether federated learning can support creative communities when computation is distributed but model governance remains centralised. By distinguishing storage, circulation, and learning and examining creator-governed infrastructures, it finds that governance is established for the first two layers but remains open at learning, motivating model-centered commons design.
Problem
Federated learning distributes computation while model aggregation and ownership typically remain with a central organisation, leaving creative communities’ agency over resulting models unresolved.
Method
The paper distinguishes storage, circulation, and learning as governance layers and examines artist-governed trusts, cooperatives, and consent infrastructures.
Results
Creator governance is established at storage and circulation, but contributors can consent to training while governance of the trained model remains under negotiation.
Takeaways & Limitations
A creative data commons must govern models and their federation, not only datasets, while making refusal, provenance, stewardship, and contribution terms actionable.
Takeaways & Limitations
Federation can reproduce inequalities because the capacity to run nodes, steward repositories, interpret licences, and maintain models is unevenly distributed.
Abstract
from arXiv · showhide
Federated learning is increasingly presented as a privacy-preserving advance: personal data remain on the device, and only model updates are shared. It borrows the vocabulary of the federated social web, yet inverts its logic, distributing computation while the resulting model stays with whoever convened the training. We argue that federation is not in itself a remedy for extractive AI, because outcomes depend on who governs the data and the model and who has agency over the practices that shape them. We describe three layers at which a creative community can hold its work: storage, circulation, and learning. Examining artist-governed trusts, cooperatives, and consent infrastructures, we show that creator governance is established at storage and circulation but stops at learning: contributors can consent to training, yet have little say over the resulting model or its federation. We map the research space this opens, pairing technical open problems with the human questions from which they unfold. We propose four design principles for a creative data commons that governs models and their federation, not only datasets: govern the model, not only the corpus; make the terms legible at the moment of contribution; design for refusal as a first-class state; and decide stewardship in the open and account for it.
1 Introduction: Two Federations
The paper distinguishes federated social-web architectures, which distribute publication and community governance, from federated learning, which distributes computation while central organisations typically retain model ownership. It argues that federation is a protocol whose effects on creative agency depend on governance of data and models.
- Federated social-web systems distribute publication, while server communities decide what they host, with whom they federate, and whom they refuse.
- Federated learning keeps personal data on devices but sends local model updates to a central server that aggregates them into one model.
- These paradigms distribute different things: one distributes data with governance, while the other distributes computation and typically retains model ownership centrally.
- Federation can increase or erode creative agency depending on who governs the data and resulting model.
- The paper distinguishes the paradigms, analyzes storage, circulation, and learning as governance layers, and maps technical and artistic questions for governing models and their federation.
2 What creators already do
Creative practitioners use small, self-curated datasets as artistic materials, shaping models toward particular aesthetics and treating data practices as part of the work. Yet governance practices for sharing resulting models across creative communities remain underdeveloped.
- Artists often use small, self-curated datasets to steer models toward particular aesthetics rather than general performance.
- Examples include photographing and labelling ten thousand tulips for Myriad (Tulips) before training a GAN for Mosaic Virus.
- Artists curate, label, refuse, withdraw, and govern data at artwork, gallery, and exhibition scales, shaping outcomes materially and conceptually.
- Practices for governing and sharing resulting AI models across creative communities remain underdeveloped.
3 When the model is the material
The paper treats models as materials through which communities can jointly reshape technical systems and creative practice. Collective care of models could open governance over how models respond, are shared, and are adapted, but current infrastructure does not support this.
- Models can function as materials: engaging with them participates in practices that form both the system and the people working with it.
- Artists reshape models through labelling, training, sharing, and deliberate manipulation of their architectures, timescales, and values.
- Collective care could let communities jointly sense how a model responds, resists, prunes, re-weights, and refuses.
- Learning-layer governance concerns not only rights over training data but also how a model is shared and under what conditions.
- Collective model governance remains inaccessible largely because infrastructure lacks places for models to be configured and manifested on collective terms.
4 Three layers, and where governance stops
The paper distinguishes storage, circulation, and learning as governance layers. Creative communities have established governance over storage and circulation, while consent to training exists but governance of resulting models remains unsettled.
- Storage concerns where data or artefacts sit and who accesses them; circulation concerns attribution, licensing, payment, and withdrawal terms.
- Learning includes who may train models, under which conditions, and who may hold, adapt, deploy, redistribute, or withdraw the resulting model.
- TRANSFER federates storage across member studios, linking separate archives for redundancy without absorbing contributors’ work into one proprietary store.
- Creator governance is established at storage and circulation, while governance at the learning layer remains absent or underdeveloped.
- Serpentine’s Choral Data Trust created a purpose-built dataset with 15 UK choirs, appointed a steward, and established governance frameworks for external training.
- Choirs consented to training but did not decide who could adapt or adopt the model, how it circulated, or what happened after the exhibition.
5 Why this is a design problem and not only a technical one
The learning layer is a governance problem, not merely an aggregation problem: federated systems can preserve asymmetries unless communities can shape data, models, and stewardship.
- Governance is absent at the learning layer when contributors have no agency over models trained on their work.
- Federating a data commons can accelerate power asymmetries because access to stewardship requires unevenly distributed education, time, compute, and social access.
- Unpaid stewardship may fall to those with capacity, while the people an infrastructure supports often have the least capacity to maintain it.
- A commons designed without attention to existing power relations can reproduce inequalities in who accesses and shapes it.
6 The research space this offers
Extending governance to learning opens technical and social questions about aggregation, provenance, withdrawal, feasibility, and preserving creative difference in small federations.
- The research agenda includes aggregation for small, heterogeneous, identifiable client groups rather than anonymous crowds.
- Model merging raises questions about tracing whose work a model rests on and enabling meaningful creator participation in decisions.
- Withdrawal after training is unresolved because models do not forget on request, requiring contingent options and accessible governance interfaces.
- The agenda asks what federated training is feasible for small networks without institutional computation and how technical, legal, and infrastructural designs can reflect creator desires.
- These directions are areas where technical, social, and practical frictions must be worked through together, not discrete problems solved independently.
- Federation can make per-contributor consent, provenance, and refusal more tractable, but preserving creative heterogeneity during aggregation remains an open question.
7 Design principles
The paper proposes four situated principles for creative data commons: govern models, make contribution terms legible, support refusal, and openly account for stewardship.
- Governance should make contribution, attribution, adaptation, and withdrawal expressible for trained models, not only their source corpus.
- Contribution terms should be acknowledged in the contributor’s vocabulary, allowing artists to shape their own licence relations.
- Refusal should be a first-class governance state, with permissions attached to data and practices rather than recorded elsewhere.
- Stewardship should be an explicit, revisable decision covering maintenance, terms, and compensation.
- Unaccounted stewardship falls to whoever can afford to provide it, reproducing exclusions the commons is intended to escape.
- Artists currently lack both infrastructure and vocabulary or legal instruments for sharing self-trained models on terms they set.
8 Conclusion
The paper concludes that creative communities already govern storage and circulation through artist-organized practices, while governance of learning and resulting models remains to be developed.
- Creative communities have embedded governance approaches for storage and circulation, but learning remains the unresolved third layer.
- Future work combines interviews, participatory co-design, a public database of open-source models, and prototypes for community-governed training and sharing.
- The paper frames model governance as a question of collective agency over what AI systems do and become.