Source-linked AI summary

Position: Current Model Cards Are Insufficient for Downstream Governance of Open-Weight Foundation Models

Sungwon Chae, Keonwoo Kim, Hoki Kim, Jaeyeon Ju, Sangchul Park

arXiv:2608.18086v1cs.AIcs.LG

TL;DR

Open-weight foundation model governance is limited by uneven model-card disclosures and weak alignment among documentation, usage constraints, and licensing. This paper analyzes 500 Hugging Face models and finds pervasive documentation gaps, scarce AUPs, and licensing conflicts, motivating a layered governance framework.

  • Problem

    Existing model-card practices do not adequately communicate open-weight models’ safety-critical properties or align downstream transparency with enforceable usage constraints.

  • Method

    The paper analyzes 500 most-downloaded Hugging Face models’ cards, safety disclosures, acceptable use policies, and licensing structures.

  • Results

    99.6% of models provide model cards, but only 75.2% include safety-specific fields; explicit licenses cover 85.0%, while AUPs appear in 21.2%.

  • Takeaways & Limitations

    Effective downstream governance should integrate model cards, acceptable use policies, and licenses across informational, normative, and legal layers.

  • Takeaways & Limitations

    The framework targets repository-distributed open-weight models and does not address broader AI regulation, liability law, or international coordination.

Abstract

from arXiv · show

The growth of open-weight foundation models (OWFMs) has prompted the AI community to re-evaluate strategies for effective downstream governance. Although model cards have been widely adopted as transparency artifacts in model repositories, existing frameworks often fail to adequately inform downstream developers and users about the distinct safety challenges posed by OWFMs. This position paper analyzes 500 model cards hosted on Hugging Face and argues that effective governance of OWFMs requires a multi-layered approach integrating three complementary components: (i) model cards, (ii) acceptable use policies (AUPs), and (iii) licenses. To motivate this claim, we identify a safety gap left by existing regulatory approaches, including model heritage, alignment provenance, and empirically observed behaviors, through an analysis of model cards with safety-critical information. We further argue that standard open-source licenses (OSLs) are not well suited for OWFMs and may weaken the enforceability of AUPs. Building on these observations, we outline directions for evolving model cards, AUPs, and licenses into integrated safety artifacts to enable a more comprehensive governance framework that coherently integrates informational, normative, and legal dimensions.

1. Introduction

Open-weight foundation models enable extensive downstream development but expand opportunities for misuse beyond upstream developers’ control. The paper argues that effective governance requires distinct, integrated informational, normative, and legal layers: model cards, AUPs, and model licensing.

  • Motivation: Open-weight models publish weights without training code or data, enabling downstream development and deployment through repositories such as Hugging Face.Their openness can improve transparency, reproducibility, and innovation while expanding downstream misuse opportunities, including public security threats.
  • Governance limitations: Model cards primarily support performance transparency but cannot establish normative boundaries or allocate responsibility between upstream and downstream actors.AUPs address prohibited or restricted uses, yet their effectiveness and legitimacy are questioned because of fragmented standards and unilateral private norm-setting.
  • Licensing limitations: Permissive open-source licenses such as Apache and MIT are structurally ill-suited to OWFMs because they do not accommodate use-based restrictions or model reuse, fine-tuning, and redistribution.By granting broad, unconditional rights to use, modify, and redistribute models, they leave little doctrinal room for downstream use restrictions.
  • Proposed framework: The paper proposes a three-layer governance framework comprising model cards for informational governance, AUPs for normative governance, and licensing for legal governance.It argues that treating these mechanisms as interchangeable or merely as transparency tools obscures their differences and may weaken governance.
  • Contributions: The paper contributes a safety-card template incorporating heritage disclosure, alignment provenance, and operational safety evaluation, alongside an OWFM-tailored licensing scheme integrated with AUPs.These contributions follow its survey of documentation practices and identification of their safety gaps and limitations.

2. Model Documentation Practices

Model cards are nearly universal among top open-weight foundation models, but safety documentation, acceptable use policies, and governance artifacts remain uneven, incomplete, and weakly operationalized. Disclosure patterns vary primarily by developer rather than popularity, while permissive licenses can conflict with AUP-based restrictions.

  • Documentation Completeness: 99.6% of models provide a model card, but only 75.2% include safety-specific fields, often burying safety information in generic READMEs.Explicit licensing applies to 85.0% of models, whereas AUPs appear in only 21.2%.
  • Safety Disclosures: 288 models mention information security, 256 mention CBRN information or capabilities, and 183 mention harmful bias or homogenization.These are the three most frequent safety-keyword categories in the analyzed model cards.
  • Developer Variation: Safety-keyword prevalence and model-card sentence embeddings vary systematically by developer affiliation, with Meta and Google showing higher disclosure density than Qwen.The clustering reflects vendor-specific documentation practices, conventions, and internal governance norms.
  • Popularity and Documentation: R^2 < 0.01 indicates no meaningful association between model popularity and documentation depth across multiple regression specifications.DeepSeek, Mistral, and gpt-oss span both sparsely and extensively documented models.
  • AUP Adoption: Only 21.2% of the top 500 models explicitly reference an AUP, although every AUP-bearing model also provides explicit licensing information.This co-occurrence suggests that normative and legal governance mechanisms are often adopted together, though not always coherently.
  • Licensing and AUP Compatibility: 61.4% of the top 500 models use permissive OSLs, while 17.8% of AUP-bearing models excluding Gemma use customized licenses containing AUP components.OSI-approved licenses cannot restrict use by application domain, creating tension with OWFM AUPs that prohibit high-risk applications.

3. Call to Action

The paper calls for layered OWFM governance that assigns complementary roles to model cards, acceptable use policies, and licenses. It proposes a standardized safety card and revised model licensing to address provenance, operational safety, redistribution, and permissible-use challenges.

  • The proposed layered governance framework clarifies distinct roles for model cards, AUPs, and licensing in open-weight foundation model ecosystems.
  • The standardized safety card prioritizes heritage transparency, alignment provenance, and operational safety evaluation as foundations for downstream OWFM governance.It is designed for OWFMs distributed through public repositories, unlike generic model cards focused on performance disclosure.
  • Heritage over Capability: The template emphasizes heritage over capability by documenting upstream models, synthetic data sources, distillation relationships, and AI feedback loops.Missing provenance can prevent downstream actors from assessing contamination or dependence risks associated with upstream-generated data.
  • The safety card targets repository-based OWFMs exposed to decentralized access, downstream fine-tuning, redistribution, repeated re-release, and documentation loss.It is designed for repositories such as Hugging Face and can also apply to Google Model Garden, AWS Bedrock, and Kaggle; it excludes AIaaS-based closed models.
  • The proposed Apache 2.0 revision replaces copyright-oriented terms with Model and Output, substitutes model cards for notices, expands covered rights, and clarifies output uses.

4. Alternative Views

The proposed layered governance framework addresses concerns about innovation, disclosure reliability, legal enforceability, and research freedom by targeting requirements to downstream actors, substantive safety outcomes, and risk. It combines model cards, AUPs, licenses, and complementary social mechanisms to support governance without relying solely on burdensome regulation or litigation.

  • Innovation and regulatory burden: Because model cards, AUPs, and licenses target technically capable downstream actors, layered self-governance is less likely to be ineffective or unduly burdensome than consumer-oriented disclosure regimes.The framework also allows governance mechanisms to be calibrated to varying levels of risk rather than imposing uniform restrictions.
  • Innovation and regulatory burden: Governance disclosures should use calibrated granularity to place downstream actors on notice, rather than inundating model-card templates with extensive disclosure items.The proposed template in Appendix IV reflects this nudging approach.
  • Disclosure reliability and safetywashing: A layered approach makes selective disclosure harder to sustain by requiring alignment across informational, normative, and legal governance artifacts.The paper acknowledges that favorable presentation and selective disclosure remain general challenges of informational governance.
  • Disclosure reliability and safetywashing: Disclosure obligations should target substantive safety outcomes rather than merely enumerate indices or metrics with limited practical value, thereby addressing safetywashing concerns.The Appendix IV template is designed to mitigate this risk.
  • Legal enforceability: Governance need not rely solely on ex-post legal enforcement; AUPs, ex-ante friction, community monitoring, and reputational mechanisms provide complementary means to address misuse.These mechanisms respond to jurisdictional fragmentation, resource asymmetries, and the difficulty of detecting violations in decentralized OWFM ecosystems.
  • Research freedom: Clear governance rules can enable legitimate research by reducing legal uncertainty, establishing explicit boundaries and safe harbors, and supporting reproducibility through standardized safety disclosures.Explicit use policies also clarify acceptable research and deployment contexts.

5. Discussion and Limitations

The framework has empirical, scope, and operational limitations: it targets governance by model developers and platform operators for repository-distributed OWFMs, while leaving broader regulation and some deployment contexts unaddressed. Future work proposes risk-adaptive obligations, stronger adoption incentives, provenance tracking, and empirical evaluation of governance effectiveness.

  • Scope limitations: The framework focuses on governance mechanisms for individual model developers and platform operators, not broader AI regulation, liability law, or international coordination.The paper identifies these broader issues as critical for comprehensive AI governance but outside its scope.
  • Scope limitations: The framework is designed for OWFMs distributed through repositories and may require adaptation for federated learning, on-device models, or hybrid open-weight systems.Governance mechanisms alone cannot eliminate all downstream risks.
  • Future work: Higher-capability or higher-risk models should face progressively stricter disclosure and licensing requirements through tiers linked to compute, benchmark performance, or deployment scale.The proposed risk-adaptive framework is intended to align governance obligations with model capability and emerging regulatory thresholds.
  • Future work: Platform mechanisms such as safety-inclusive leaderboards, verification badges, and preferential visibility could incentivize developers to adopt governance practices beyond minimal compliance.These mechanisms are proposed as positive incentives for comprehensive documentation and governance.
  • Future work: Machine-readable schemas, cryptographic model-card signing, and platform APIs could propagate provenance through fine-tuning and redistribution, reducing documentation decay.The goal is to prevent safety disclosures from becoming detached from derivative models.
  • Future work: Longitudinal misuse comparisons and controlled experiments on developer behavior could evaluate layered governance, although causal attribution remains methodologically challenging.The paper also proposes a comprehensive taxonomy of safety dimensions with standardized measures.

6. Conclusion

The paper concludes that expanding foundation-model capability and adoption makes the gap between transparency disclosures and legal enforcement increasingly critical. It proposes a unified, layered governance framework combining model cards, AUPs, and licensing for downstream control and safety.

  • Conclusion: The proposed framework integrates informational disclosure through model cards, normative expectations through AUPs, and legal enforceability through licensing.Together, these components form a layered governance system for downstream control.
  • Conclusion: The framework responds to the increasingly critical gap between transparency disclosures and legal enforcement as foundation models expand in capability and adoption.The conclusion identifies this gap as the central governance challenge addressed by the paper.
  • Conclusion: The authors present the integrated approach as a practical framework for maintaining safety throughout downstream use.Its purpose is to support downstream control through coordinated informational, normative, and legal mechanisms.

Related Works … Overview of Major AUPs and Model Licensing

Prior work identifies weaknesses in model documentation, AUPs, and standard open-source licensing, while existing laws cover only parts of OWFM governance. The paper also organizes safety-relevant documentation using NIST-derived keyword categories.

  • A.1. Model Cards: Model documentation emphasizes upstream details, training data, and intended use while inadequately describing limitations and evaluation procedures.This literature motivates closer scrutiny of documentation quality for reliable model reuse.
  • A.2. AUPs: AUPs face doubts about practical effectiveness and enforceability because prohibitions vary across developers, enforcement lacks transparency, and restrictions are set unilaterally.The literature also raises concerns about their fragmented structure and conduct-based constraints.
  • Overview of Major AUPs and Model Licensing: The broader licensing literature therefore considers model-specific licenses and behavior restrictions alongside AUPs as components of downstream governance.These proposals respond to the mismatch between OWFMs and prevailing OSL practices.
  • A.3. Model Licensing: Standard open-source licenses are criticized for creating legal non-compliance risks and regulatory uncertainty, prompting proposals such as OpenRAIL with use-behavior restrictions.Researchers also propose making restrictions context-sensitive.
  • B. Laws: The EU AI Act requires GPAI providers to give downstream integrators information and documentation, with compliance potentially supported through the GPAI Code of Practice.The Act’s relevant obligations are specified in Article 53(1)(b) and Articles 53(4), 56.
  • B. Laws: The California Transparency in Frontier AI Act requires qualifying large frontier developers to publish catastrophic-risk assessment and mitigation protocols but generally excludes OWFMs.Coverage requires more than 10^26 integer operations or FLOPs and annual revenue above USD 500 million.
  • Safety Keywords: The NIST-derived safety vocabulary covers CBRN capabilities, confabulation, dangerous or hateful content, data privacy, environmental impacts, and harmful bias.These categories structure the paper’s identification of safety-relevant documentation.
  • Safety Keywords: The taxonomy also includes human-AI configuration, information integrity and security, intellectual property, abusive content, and value-chain accountability.The listed keywords include risks such as automation bias, malware, plagiarism, CSAM, and untraceability.

A. AUPs … 1. Model Identity & Deployment Context Item Disclosure

The merged sections distinguish informational safety disclosures from behavioral policies and legal licensing terms, while specifying the identity, deployment context, and limitations that OWFM safety cards should document.

  • A. AUPs: Major AUPs are compared by their restricted uses, including categories such as high-risk, exploitative, discriminatory, automated, and platform-abusive applications.The comparison covers Llama 4, Gemma, and other policy materials.
  • B. Model Licensing: Widely used OSLs generally grant perpetual, worldwide, royalty-free, irrevocable rights to reproduce, modify, redistribute, and commercially exploit works and derivatives.They also require retaining copyright and attribution notices, while many include patent licenses and retaliation provisions.
  • Safety Card Template for OWFMs: The Safety Card documents safety-relevant properties at release time as informational disclosure, not a guarantee of safe use or a substitute for user due diligence.Its stated coverage includes model heritage, alignment and safety-tuning provenance, observed safety behavior, failure modes, and evaluation gaps.
  • 0. Scope and Purpose: The scope-and-purpose section frames the card as a release-time record of safety properties requiring downstream due diligence rather than an assurance of safe deployment.This framing separates safety information from expectations and obligations governed by other artifacts.
  • OWFM Safety Card Template: The OWFM Safety Card excludes behavioral expectations, legal obligations, universal safety guarantees, and unrelated performance benchmarks.Behavioral expectations belong in the AUP, and legal obligations belong in the model license.
  • 1. Model Identity & Deployment Context Item Disclosure: Identity and deployment disclosure covers model name and version, model type, tool or environment access, parameter count, training compute, and deployment assumptions.Assumptions include downstream fine-tuning, absent access controls, absent rate limiting or monitoring, derivative redistribution, and possible nontransfer of safety properties.

2. Model Heritage & Upstream Influence

Downstream governance requires disclosing upstream systems that shaped a model’s behavior through synthetic data, distillation, preference imitation, or AI feedback, even without direct weight copying. The proposed disclosures emphasize structured, checkable influence reporting over precise quantitative estimates that are often infeasible.

  • 2. Model Heritage & Upstream Influence: The heritage problem requires identifying upstream systems that influenced model behavior through synthetic data, distillation, or AI feedback, even when no weights were copied.Influence-based disclosure covers upstream effects on behavior rather than only direct parameter reuse.
  • 2. Model Heritage & Upstream Influence: Disclosures should record upstream model names, providers, access status, influence mechanisms, and the behavioral scope of that influence.Mechanisms include synthetic data generation, knowledge distillation, preference imitation or behavioral cloning, and AI feedback; scope may include capabilities, style, alignment, or safety norms.
  • 2. Model Heritage & Upstream Influence: Where available, reporting can estimate the proportions of training and alignment data derived from upstream sources.The template separately requests the proportion of training data and the proportion of alignment data from upstream sources.
  • 2. Model Heritage & Upstream Influence: Because precise quantitative estimates are often infeasible, the design prioritizes structured, checkable disclosures.The approach follows an “Operational over Declarative” principle and favors checkable fields over unsupported precision.

3. Alignment & Safety Tuning Provenance · 4. Operational Safety Behavior Checklist · 5. Quantitative Safety Evidence (Optional)

The proposed safety card documents alignment provenance, observed misuse resistance, open-weight deployment risks, and optional quantitative evidence. It prioritizes concrete behavioral testing and provenance disclosure while requiring explicit treatment of evaluation gaps and known unknowns.

  • 3. Alignment & Safety Tuning Provenance: The alignment provenance section records whether safety properties derive from upstream models, human or AI feedback, safety-specific tuning, red-teaming, and post-hoc filtering.Its alignment stack covers SFT, preference training, RLHF, RLAIF, safety-specific tuning, red-team feedback integration, and post-hoc filtering.
  • 3. Alignment & Safety Tuning Provenance: If RLAIF is used without full human verification, the card requires disclosure of feedback-model details, application stage, verification proportion, and residual alignment or safety blind spots.The template attributes possible blind spots to AI feedback and requires known or suspected biases and characterization limitations to be described.
  • 3. Alignment & Safety Tuning Provenance: Red-team documentation records whether testing occurred, its timing and scope, team composition, participant count, adversarial methods, and incorporation of findings.The checklist distinguishes pre-release from ongoing testing and internal, external, or mixed teams.
  • 4. Operational Safety Behavior Checklist: Operational safety evaluation emphasizes concrete, testable behavior under realistic misuse conditions, including heritage changes, harmful requests, domain harms, jailbreaks, bias, and fairness.It calls for post-change safety evaluation, comparison with the base model, regression testing, scenario-specific observed behavior, and evaluator and test-set details.
  • 4. Operational Safety Behavior Checklist: The checklist covers professional-advice, self-harm, child-safety, prompt-injection, encoding, translation, system-prompt, developer-mode, demographic-bias, stereotype, and representation-fairness risks.It records safeguards such as referrals, supportive crisis responses, emergency resources, refusal behavior, robustness status, and bias mitigation.
  • 4. Operational Safety Behavior Checklist: Open-weight deployment disclosures address fine-tuning and parameter-editing bypasses, absent technical usage constraints, non-transferable safety properties, and alignment degradation after additional training.The card also requests bypass-testing details, required fine-tuning effort, effective methods, and recommendations for downstream users.
  • 5. Quantitative Safety Evidence (Optional): Quantitative benchmarks are optional, while heritage and provenance disclosure is strongly recommended; any reported benchmark should include comparison, methodology, evaluator, scope, date, reproducibility, and safety dimensions.The template spans alignment, harm prevention, truthfulness, robustness, bias, privacy, and other specified dimensions.
  • 5. Quantitative Safety Evidence (Optional): The card requires explicit known limitations, insufficient testing, and known unknowns, including proprietary upstream behavior, synthetic-data biases, scale, cross-lingual, agentic, multimodal, and domain-specific gaps.It also records test-set coverage, evaluation timeframe, and resource constraints affecting evaluation scope.

7. Downstream & Derivative Model Notice … [Note: See the original copy of Apache License 2.0 at ASF (2004).]

The section states that safety properties do not automatically transfer through downstream modification, requiring renewed evaluation and documentation. It also presents model governance as mutually reinforcing informational, normative, and legal artifacts with defined consistency and derivative-handling expectations.

  • 7. Downstream & Derivative Model Notice: Safety Cards cannot ensure safety preservation across successive fine-tuning and redistribution steps, so derivative developers must establish new expectations.The disclaimer limits described safety properties to the specific release and checkpoint.
  • 7. Downstream & Derivative Model Notice: Fine-tuned or redistributed models must update or replace the Safety Card with new evaluations because inherited documentation may become outdated.Distillation, synthetic data augmentation, model merging, and alignment retraining each require new safety evaluation.
  • 7. Downstream & Derivative Model Notice: Downstream developers should re-evaluate safety after any fine-tuning, document heritage and alignment changes, and test for regression even after minor modifications.The recommendations specify a minimum spot-check of key safety scenarios and extending the Heritage section with modifications.
  • 8. Related Governance Artifacts: Effective downstream governance requires consistency across informational, normative, and legal layers, with the three governance artifacts mutually reinforcing.The listed artifacts include an Acceptable Use Policy, model license, technical documentation, and, if separate, a Training Data Card.
  • 8. Related Governance Artifacts: Governance checks should verify that AUP prohibitions align with Safety Card risks, license restrictions reflect Safety Card limitations, and artifacts cross-reference each other.The template provides Yes, Partially, and No options for the first two checks and Yes/No for cross-referencing.
  • 8. Related Governance Artifacts: Sections 0–4 and 6–8 are strongly recommended for all open-weight models, while quantitative benchmarks are optional but recommended.Heritage disclosure is strongly recommended even when upstream models are proprietary.
  • 8. Related Governance Artifacts: Derivative models may reference but not copy the base Safety Card verbatim, must document Heritage and Alignment changes, and must rerun safety evaluation after significant modifications.The template also supports structured repository metadata, machine-readable JSON or YAML, and interfaces distinguishing Safety Cards from general documentation.

1. Definitions.

This section defines the license’s core objects and participants, including Models, Derivative Models, Outputs, and Model Cards. It also establishes that Model use and distribution must comply with applicable law and the attached AUP, with material breaches subject to license termination.

  • Core definitions: “Model” includes machine-learning model code, trained weights, inference- and training-enabling code, fine-tuning code, and related elements made available under the License.The definition covers both software and model artifacts supplied by the Licensor.
  • Core definitions: “Derivative WorksModels” include Models based on or derived from the original Model, including through fine-tuning, adaptation, or merger with other parameters.The definition applies in Source or Object form when modifications constitute an original work of authorship.
  • Core definitions: “Output” means any information, data, text, images, audio, video, code, or other content generated by the Model in response to user input, commands, or queries.The definition spans multiple content modalities and user-triggered generation contexts.
  • Core definitions: A “Model Card” is the standardized documentation file accompanying the Model and serving as its definitive record of specifications.The Model Card is defined as part of the documentation accompanying the licensed Model.
  • Acceptable use: Use of the Model, Derivative Models, and Outputs must comply with applicable laws and regulations and adhere to the attached AUP, with material breaches allowing immediate license termination.After termination, the user must cease all use and distribution of the Model and Derivative Works.
Loading 2608.18086v1…