Source-linked AI summary
BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation
Zihao Zhu, Ruotong Wang, Siwei Lyu, Min Zhang, Baoyuan Wu
TL;DR
T2V commercialization lacks a seamless way to integrate advertiser brands without disrupting user intent, recognizability, or naturalness. BrandFusion addresses this with offline brand-knowledge construction and online five-agent prompt refinement, outperforming baselines across core integration dimensions and user satisfaction while remaining constrained by underlying T2V capabilities.
Problem
Seamless brand integration in T2V is introduced to embed advertiser brands while preserving semantic fidelity, recognizability, and contextual naturalness.
Method
BrandFusion combines offline Brand Knowledge Base construction with online collaborative prompt refinement by five specialized agents.
Results
BrandFusion significantly outperforms baselines in semantic fidelity, brand visibility, and integration naturalness, while human evaluations confirm superior user satisfaction.
Takeaways & Limitations
The framework establishes a practical pathway for sustainable T2V monetization that supports organic brand exposure, provider revenue, and uninterrupted user content creation.
Takeaways & Limitations
Integration quality is constrained by the underlying T2V model, especially for complex interactions, fine-grained details, and rapid motion.
Abstract
from arXiv · showhide
The rapid advancement of text-to-video (T2V) models has revolutionized content creation, yet their commercial potential remains largely untapped. We introduce, for the first time, the task of seamless brand integration in T2V: automatically embedding advertiser brands into prompt-generated videos while preserving semantic fidelity to user intent. This task confronts three core challenges: maintaining prompt fidelity, ensuring brand recognizability, and achieving contextually natural integration. To address them, we propose BrandFusion, a novel multi-agent framework comprising two synergistic phases. In the offline phase (advertiser-facing), we construct a Brand Knowledge Base by probing model priors and adapting to novel brands via lightweight fine-tuning. In the online phase (user-facing), five agents jointly refine user prompts through iterative refinement, leveraging the shared knowledge base and real-time contextual tracking to ensure brand visibility and semantic alignment. Experiments on 18 established and 2 custom brands across multiple state-of-the-art T2V models demonstrate that BrandFusion significantly outperforms baselines in semantic preservation, brand recognizability, and integration naturalness. Human evaluations further confirm higher user satisfaction, establishing a practical pathway for sustainable T2V monetization.
1. Introduction
Seamless brand integration is introduced as a T2V task that embeds advertiser brands while preserving user intent. BrandFusion addresses the resulting fidelity, visibility, and naturalness challenges with offline knowledge construction and online multi-agent refinement.
- Seamless brand integration embeds advertiser brands into prompt-generated videos while preserving semantic alignment with user intent.
- The task seeks recognizable yet unobtrusive brands that remain visually prominent, contextually harmonious, and semantically coherent with the original prompt.
- Maintaining semantic alignment, ensuring brand visibility, and achieving contextually natural integration are the central challenges.
- BrandFusion constructs brand knowledge offline and uses five specialized agents online to collaboratively refine prompts through iterative refinement.
- Experiments across established and novel brands and multiple T2V models report superior semantic alignment, brand visibility, integration naturalness, and user satisfaction versus baselines.
2. Related Work
Prior work improves video prompts through training-based optimization and multi-agent refinement, while brand embedding methods have largely pursued covert manipulation rather than user-aware integration.
- Training-based prompt optimization uses supervised fine-tuning or reinforcement learning to improve image-text alignment and compositional fidelity.
- Multi-agent frameworks decompose prompt refinement into specialized agents that collaborate on improving user-provided prompts.
- Lightweight fine-tuning methods such as DreamBooth and Textual Inversion enable models to synthesize novel entities.
- Adversarial brand-embedding methods inject or manipulate brand elements covertly, prioritizing stealth and concealment without user awareness.
3. Seamless Brand Integration in Text-to-Video Generation
Seamless brand integration adds brand elements to user-requested T2V videos while preserving semantic fidelity and user intent. The proposed ecosystem connects brand owners, service providers, and end users through contextually integrated branded content.
- T2V models synthesize video sequences from natural-language descriptions, with prompt quality influencing semantic alignment and specificity.
- The task combines a user prompt and advertiser brand profile to generate an integrated video that preserves the user’s intended content.
- The ecosystem involves brand owners registering profiles with T2V providers, which integrate brands into users’ creative requests.
- Users receive branded videos balancing creative intent with brand visibility, while brands gain organic exposure and providers obtain monetization channels.
4. Methodology
BrandFusion uses two synergistic phases: offline construction of a Brand Knowledge Base and online multi-agent integration. The system combines brand preparation, contextual prompt refinement, critique, and experience-based improvement.
- Offline Brand Knowledge Base Construction: The offline phase adaptively builds brand knowledge through prior-knowledge probing and selective model adaptation.
- Offline Brand Knowledge Base Construction: Brands with sufficient model knowledge are registered directly, while unfamiliar brands receive brand-specific adapters through model-level adaptation.
- Offline Brand Knowledge Base Construction: The Brand Knowledge Base stores brand metadata, adapter weights, visual patterns, and successful integration experiences for later retrieval.
- Online Multi-Agent Brand Integration: The online phase uses five specialized agents to select compatible brands, generate strategies, rewrite prompts, critique outputs, and learn from completed integrations.
- Online Multi-Agent Brand Integration: Critic feedback triggers iterative refinement, while experience learning writes completed integrations into long-term memory for continuous improvement.
5. Experiments
BrandFusion is evaluated on established and novel brands using semantic preservation, brand integration quality, visual quality, and human judgments. Across these evaluations, it maintains comparable visual quality while improving semantic fidelity, naturalness, brand presence, and user satisfaction.
- Experimental Setup: The benchmark covers 18 established brands across seven industry categories and 270 brand-prompt pairs with high, medium, and low compatibility.Each brand has 15 diverse prompts spanning three prompt-brand match levels.
- Known Brands: BrandFusion achieves comparable VBench quality while significantly outperforming baselines on semantic fidelity and naturalness across three T2V models.The comparison is reported for well-known brands in Table 1.
- Evaluation Metrics: The evaluation measures semantic preservation with VQAScore, CLIPScore, and LLMScore, while integration quality uses Brand Presence Rate and a 1-5 Naturalness Score.Naturalness averages contextual fit, visual blend, and nonintrusiveness.
- Novel Brands: For novel brands, Wan2.2-T2V-5B achieves the highest Brand Presence Rate and Naturalness Score, while Wan2.1-T2V-1.3B remains reasonably effective.The results are attributed to model capacity and generation capability differences.
- Human Evaluation: Human evaluation with 10 participants finds BrandFusion highest across semantic fidelity, integration naturalness, and overall acceptability.Participants used a 1-5 Likert scale to assess videos generated by Veo3.
6. Analysis
Analysis shows that BrandFusion remains effective across compatibility levels, scene categories, brand categories, and sequential experience-learning stages. Its online optimization averages 7.4 LLM calls and 16 seconds, compared with at least 120 seconds for video generation on Wan2.2-T2V-5B.
- Prompt-Brand Match Levels: BrandFusion maintains strong semantic fidelity and naturalness in challenging Low Match scenarios, whereas Template-based Rewriting drops from NS 4.01 to 1.38.Direct Append also declines significantly in low-match cases.
- Prompt Scene Categories: BrandFusion outperforms baselines across seven scene categories, including everyday contexts and temporally challenging Sci-Fi and Historical scenes.Urban Scenes and Social & Home Life show particularly strong results, while Temporal Themes remain more difficult.
- Brand Categories: Across seven brand categories, BrandFusion consistently outperforms baselines on semantic fidelity and integration naturalness.Apparel & Footwear achieves especially high integration quality because of its natural association with human subjects.
- Experience Learning: Experience learning produces an upward trend in overall acceptability across sequential BMW prompt groups, unlike the relatively flat baseline.The analysis uses 100 prompts divided into 10 sequential groups.
- Efficiency: 7.4 LLM calls and 16 seconds per prompt account for about 11% of the pipeline, while Wan2.2-T2V-5B video generation requires at least 120 seconds.Video generation remains the dominant latency factor.
- Novel Brand Cases: Novel-brand examples show ARUA sportswear and FreshWave beverage integrated visibly into activity and social contexts without disrupting scene coherence.The examples use Wan2.2-T2V-5B and cover lawn mowing and backyard barbecue scenarios.
7. Conclusion
The paper introduces seamless brand integration for T2V and proposes BrandFusion, which combines offline brand knowledge construction with online collaborative prompt refinement. Experiments and human evaluations report superior performance and user satisfaction across brands, models, scenarios, and categories.
- Task: BrandFusion frames seamless brand integration as embedding brands into T2V videos while preserving the user’s original semantic intent.The task is introduced as a new problem for text-to-video generation.
- Framework: The framework combines offline brand knowledge construction with online collaborative refinement by five specialized agents.This architecture supports brand integration while maintaining semantic fidelity.
- Findings: Experiments on 18 established and 2 novel brands across multiple state-of-the-art T2V models show significant improvements over baselines, with human evaluations confirming higher user satisfaction.The reported gains span diverse scenarios and brand categories.
- Implication: The authors position seamless brand integration as a practical pathway toward sustainable T2V monetization without disruptive interruptions.The proposed model supports organic advertiser exposure, service-provider revenue, and continued high-quality user content.
A.1. Known Brand Benchmark
The known-brand benchmark covers 18 established brands across seven industry categories, with prompts spanning different brand–scene compatibility levels. BrandFusion integrates these brands through a knowledge base and coordinated prompt-refinement agents.
- Benchmark Composition: The benchmark contains 18 established brands spanning seven major industry categories and diverse commercial contexts.Categories include food and beverage, technology and electronics, transportation, apparel and footwear, beauty and personal care, home and furniture, and health and wellness.
- Prompt Design: Each brand is evaluated with 15 user prompts covering high-, medium-, and low-match scenarios across varied scene types.The match levels test natural placement, creative integration, and robustness in difficult cases.
- Brand Knowledge: The Brand Knowledge Base stores brand profiles, reference images, and prior-knowledge information for registered brands.Its schema includes brand name, category, description, reference images, and whether the T2V model already has sufficient brand knowledge.
- Online Integration Workflow: BrandFusion’s online workflow uses specialized agents to select brands, generate strategies, rewrite prompts, and evaluate revised prompts.The workflow includes iterative refinement with critic feedback and decision branches for acceptance, revision, or replanning.
- Strategy Generation: Strategy generation balances semantic preservation with brand visibility by using context-aware placement approaches and prior successful experiences.Supported strategies include main-object, background, character-interaction, environmental, lifestyle, contextual, and ambient integration.
- Prompt Evaluation: The critic evaluates semantic fidelity and brand clarity to determine whether a rewritten prompt should be accepted or further refined.Its evaluation considers preservation of the user’s creative intent and recognizable brand integration.
D.3. LLMScore Implementation
LLMScore evaluates semantic consistency holistically by having a multimodal language model compare the generated video with the original prompt. The broader evaluation protocol complements this measure with human ratings of fidelity, naturalness, and acceptability.
- LLMScore: LLMScore uses a multimodal large language model to assess overall semantic similarity between the original text description and generated video.Unlike VQAScore’s discrete questions, it directly evaluates the video after comprehensively analyzing its content.
- LLMScore: The LLM evaluation considers content match and requires an objective score with reasoning about matching and missing video content.The output format contains an LLM score between 0.0 and 1.0 and an explanatory reasoning field.
- Human Evaluation: Human evaluation uses a custom web interface that presents videos, contextual information, evaluation guidelines, and questionnaire controls.The interface is designed to support systematic assessment while minimizing participant cognitive load.
- Human Evaluation: The questionnaire measures semantic fidelity, integration naturalness, and overall acceptability using three questions.The acceptability item captures satisfaction and willingness to accept the video despite brand elements.
- Human Evaluation: All human-evaluation questions use a consistent 5-point Likert scale from strongly disagree to strongly agree.This provides a common response format across the three evaluation dimensions.
E.1.2. Quantitative Results
BrandFusion achieves the strongest overall ablation performance and remains effective across LLM backbones and prompt–brand difficulty levels. Its gains reflect complementary contributions from strategic planning, iterative criticism, and stronger reasoning capacity.
- Ablation Results: The complete BrandFusion framework achieves the best performance across all ablation metrics, while removing components causes progressively larger degradation.The most severe decline occurs when both strategic planning and iterative refinement are removed.
- Ablation Results: Removing the Strategy Generation Agent lowers Naturalness Score by 0.28 points and Brand Presence Rate by 1.85%.Without strategic guidance, prompt rewriting makes integration decisions reactively and produces less optimal placement choices.
- Ablation Results: Removing the Critic Agent lowers Naturalness Score by 0.55 points and Brand Presence Rate by 4.29%.The results identify iterative evaluation and correction as more impactful than strategic planning alone.
- Ablation Results: Removing both agents lowers Naturalness Score by 0.88 points and Brand Presence Rate by 6.62%, exceeding the sum of individual degradations.This pattern suggests a synergistic relationship between strategy generation and critic verification.
- LLM Backbone Results: GPT-4o-mini retains 96.2% of GPT-5’s Naturalness Score and 97.5% of its Brand Presence Rate.Its scores are 4.52 versus 4.70 for Naturalness Score and 0.9234 versus 0.9474 for Brand Presence Rate.
- Difficulty-Stratified Results: In Low Match scenarios, Gemini-2.5-Pro reaches a Naturalness Score of 4.75, 0.33 points above GPT-5 and 0.57 points above GPT-4o-mini.The largest backbone-related gains occur when scene–brand compatibility is inherently challenging.
- Cost–Performance Tradeoff: GPT-4o-mini offers approximately 8× lower inference cost than GPT-5 and achieves a Naturalness Score per dollar of 36.16.The reported efficiency is 7.7× better than GPT-5, supporting cost-sensitive deployment scenarios.
F.1.2. IKEA Brand Integration Across Match Levels
BrandFusion demonstrates IKEA integration across high-, medium-, and low-match prompts by adapting placement strategies to each scene while preserving natural visual coherence. These examples illustrate functional, environmental, and natural-object integration across varying difficulty levels.
- Integration Strategies: The IKEA cases span functional product use, ambient decoration, and natural-object placement across different prompt-brand compatibility levels.The examples complement quantitative evaluation by illustrating strategic planning, iterative refinement, and context-aware reasoning.
- High Match Scenario: Home Organization: High-match wardrobe organization uses IKEA storage boxes functionally, making the branding visible while keeping the organizational activity central.The boxes hold folded sweaters and sit on shelves, with natural lighting highlighting the branding.
- Medium Match Scenario: Birthday Celebration: Medium-match birthday-party integration places IKEA-patterned gift boxes near children as decorative elements within the celebration.The branding appears through distinctive gift-wrap designs and tags without disrupting the party’s aesthetic.
- Low Match Scenario: Forest Meditation: Low-match forest meditation uses a wooden stool with a subtle IKEA logo, blending brand presence into the natural setting.The logo remains identifiable upon closer inspection while the stool functions as mindfulness equipment.
G.1. Technical Limitations
BrandFusion’s deployment remains bounded by technical, contextual, cultural, and operational constraints, alongside ethical risks involving consent, manipulation, privacy, misuse, and unequal access. The paper therefore calls for safeguards, transparency, and responsible deployment.
- Technical Limitations: BrandFusion cannot overcome underlying T2V weaknesses in complex interactions, fine-grained details, or rapid motion, which degrade brand fidelity and temporal consistency.Integration quality depends fundamentally on the capabilities of the underlying generation model.
- Technical Limitations: Extremely low-match prompts can produce forced integrations, while abstract, purely natural, or historical settings offer few natural brand-placement opportunities.These scenarios make it difficult to satisfy semantic preservation and brand visibility simultaneously.
- Technical Limitations: The current framework handles one brand per video; simultaneous multi-brand integration would require additional planning and multidimensional evaluation.Multiple brands may compete for visual attention and advertiser interests may conflict.
- Technical Limitations: Strategies developed primarily in Western commercial contexts may not generalize across cultures with different brand perceptions, advertising norms, and visual aesthetics.The paper proposes expanding the Brand Knowledge Base with region-specific guidelines and cultural-sensitivity considerations.
- Ethical Implications: Natural-looking placements raise risks of subliminal influence and blurred boundaries between user-generated content and commercial messaging.The paper warns that users may be unable to distinguish integrated advertising from organic scene elements.
- Ethical Implications: Ad-supported T2V may broaden access while disadvantaging users who cannot afford ad-free options through more intrusive brand placements.The paper recommends balancing revenue generation with equitable access and reasonable integration constraints.
- Safeguards: Responsible deployment requires safeguards addressing transparency, consent, moderation, unauthorized usage, harmful associations, privacy, and other misuse scenarios.The recommended framework includes disclosure, opt-in controls, and automated and human content review.