Source-linked AI summary
DIRECT: Decomposing Audience Preference and Creative Effect in Visual Content Analytics
Yizhi Liu, Balaji Padmanabhan, Siva Viswanathan
TL;DR
Pooled visual-content coefficients can conflate audience preference with the creative effect relevant to deciding what a creator’s next post should look like. DIRECT separates these patterns in panel data, finding opposite signs on 4 of 11 attributes and that pooled prescriptions forgo 31% of achievable engagement gain.
Problem
Visual content analytics lacks clear evidence about which estimand supports valid recommendations for what creators’ next posts should look like.
Method
DIRECT separates audience preference from creative effect using a panel-based causal framework centered on within-source effects.
Results
31% of achievable within-source engagement gain is forgone when held-out creators follow the pooled coefficient, which points the wrong way on 2 of 11 attributes.
Takeaways & Limitations
Creative-effect signals should guide direction-sensitive recommendations, while pooled coefficients provide audience-preference context rather than a single leaderboard.
Takeaways & Limitations
Residual within-source confounding cannot be fully ruled out.
Abstract
from arXiv · showhide
Which visual choices make a post perform better? A growing literature answers this question with pooled coefficients estimated across many creators, which platforms translate into creative recommendations. We show that these coefficients blend two distinct patterns that can point in opposite directions for the same attribute. The first, audience preference, arises because creators who favor a style attract differently composed audiences, so their posts perform differently because of who is watching, not what any single post does. The second, creative effect, captures how a creator's audience responds when she departs from her usual look. Pooled estimation averages the two, and audience preference can be large enough to reverse the signal that creative direction requires. We propose DIRECT (Decomposed Identification of Response Effects via Causal Tools), a panel-based causal-inference framework that separates them, combining the Mundlak between-within decomposition with double machine learning over latent vision-language treatments that co-vary within a creator. We apply it to 232,088 sponsored Instagram beauty posts across 1,527 creators and 11 CLIP-derived visual style axes. The two carry opposite signs on 4 of 11 attributes, and on 2 of 11 the pooled coefficient itself recommends the wrong creative direction: on skin tone, it favors lighter representations while the creative effect points the other way, since a creator's audience engages more with tones darker than her baseline. On held-out creators, prescribing from the pooled coefficient forgoes 31% of the achievable engagement gain. We contribute a diagnosis of estimand mismatch in visual content analytics, a framework that recovers the decision-relevant estimand from observational panel data, and three portable diagnostics for auditing whether pooled estimates support the decisions they inform.
Abstract
The paper sits at the intersection of visual content analytics, causal inference, double machine learning, influencer marketing, and user engagement.
- The study combines visual content analytics with causal inference and double machine learning.
- Its application domain is influencer marketing, with user engagement as a focal outcome.
Introduction
Pooled visual-content coefficients conflate audience preference across creators with creative effects within a creator, so they can recommend the wrong creative direction. DIRECT separates these effects from observational panel data and provides diagnostics for auditing decision validity.
- Estimand mismatch: Pooled coefficients capture which visual choices accompany better performance across creators, whereas creative decisions require knowing whether changing a choice improves a particular creator’s content.The distinction is between audience preference, which informs creator selection, and creative effect, which informs creative direction.
- Estimand mismatch: Audience preference reflects between-creator associations, while creative effect reflects within-creator responses to posts that depart from a creator’s usual visual style.The pooled coefficient is a variance-weighted average of the two effects.
- DIRECT framework: DIRECT separates the two effects from a single observational panel using panel-based causal inference and double machine learning over latent, co-varying vision-language treatments.Axis-specific recovery requires cross-axis nuisance modeling because multiple visual axes co-vary within a creator.
- Empirical results: 4 of 11 attributes show opposite signs for audience preference and creative effect, while 2 of 11 produce pooled coefficients pointing creative direction the wrong way.The findings come from 232,088 sponsored Instagram beauty posts across 1,527 creators and 11 CLIP-derived visual style axes.
- Empirical results: 31% of the achievable within-source engagement gain is forgone when held-out creators are prescribed pooled-coefficient directions.This result is reported under the magnitude comparison used in the study.
- Contributions: The paper contributes an estimand-mismatch diagnosis, a panel-based DML framework, and three portable diagnostics for assessing whether pooled estimates support applied decisions.The mismatch is not bias in the usual sense: pooled estimates can be unbiased for their own estimand while misaligned with the decision they support.
Related Work
Prior visual advertising research extracts visual features at scale and relates them to engagement with pooled specifications, but this approach does not separate the economic effects underlying engagement. DIRECT instead combines panel-based double machine learning with the Mundlak decomposition to target the appropriate estimand, with benchmarks showing that unorthogonalized within-creator estimates can be severely distorted.
- Prior visual advertising research: Visual advertising research typically extracts visual features at scale and estimates their relationship with engagement using pooled specifications with covariate controls.This strategy consistently describes cross-creator patterns but does not separate the two economic effects underlying engagement.
- Methodological positioning: The methodological literature includes panel-based double machine learning and observational policy learning, while DIRECT shifts the decision-theoretic focus from whom to treat to which estimand should guide treatment design.The paper positions its contribution as applying this stance to estimand choice rather than treatment assignment.
- Methodological positioning: DIRECT combines DML with the Mundlak decomposition, whose pooled-between-within identity alone does not yield axis-specific marginal effects for latent, co-varying vision-language treatments.The decomposition supplies the algebraic identity, but additional treatment-specific handling is required.
- Benchmark evidence: 40 to 100% inflation occurs on most axes, roughly 1,300% on one axis, and sign reversal occurs on 2 of 11 axes when within-creator fixed effects lack cross-axis orthogonalization.The benchmark motivates the distinction between the relevant estimands and unorthogonalized within-creator estimates.
The DIRECT Framework
DIRECT separates audience-related between-source variation from the within-source creative effect that governs visual direction. It combines Mundlak decomposition with panel-based double machine learning and validates whether within-source prescriptions improve engagement over pooled guidance.
- Estimand definition: The direction-relevant estimand is the within-source partial effect τW of a visual attribute on engagement, because a creative brief changes style within a fixed source.Between-source differences can reflect audience composition and other source-level factors, while within-source differences compare posts by the same source.
- Estimand mismatch: Pooled coefficients blend between-source τB and within-source τW, so when their signs differ, pooled guidance can point opposite to the creative-direction rule.The pooled coefficient correctly answers a descriptive question but implicitly treats between-source variation as creative effect when used for direction.
- Four-stage framework: DIRECT projects images onto interpretable vision-language style axes, decomposes each attribute into source means and post deviations, and jointly identifies τB and τW with panel-based DML.Its nuisance functions absorb post-level controls, source-level covariates, and cross-axis confounding; estimated τB informs creator selection while τW informs creative direction.
- Four-stage framework: DIRECT validates decision relevance by comparing τW-based prescriptions with pooled prescriptions on held-out creators and estimating the resulting engagement difference.The validation uses conflation cost in percentage terms and an off-policy doubly-robust value estimate as an independent check.
- Portable diagnostics: DIRECT supplies three diagnostics: D1 counts sign disagreements between τB and τW, D1′ counts pooled-versus-τW sign disagreements, and D2 measures magnitude divergence.D3 reports the percentage effectiveness loss from prescribing on pooled rather than within-source estimates.
- Identification conditions: The causal interpretation of τW requires residual within-source style variation to be approximately exogenous after conditioning on source effects, controls, and cross-axis nuisances.Within-source variation may still reflect campaign objectives, product types, or sponsorship terms; DIRECT addresses this concern with orthogonalization and held-out validation.
Empirical Implementation
The empirical implementation analyzes 232,088 sponsored Instagram beauty posts from 1,527 creators over 24 months, representing visual style with 11 CLIP-derived axes. It estimates between-creator and within-creator effects using a two-level partial linear model with cross-fitted machine-learning nuisance functions, then evaluates prescriptions on creator-disjoint holdout data.
- Data and sample: 232,088 sponsored Instagram beauty posts from 1,527 creator accounts form a 24-month panel spanning 2024–2025.Creators contribute a mean of 152 posts, with substantial within-creator stylistic variation.
- Visual treatments: 11 face-validated visual style axes are extracted with CLIP ViT-L/14 from image–language anchor-prompt similarity differences.The axes span color, lighting, composition, framing, and demographic representation.
- Estimation model: The two-level partial linear model separates each attribute’s between-creator mean from its within-creator deviation while controlling for post-, creator-, and other-axis covariates.For attribute k, ūi(k) is the creator mean and ũi,t(k) is the within-creator deviation; the nuisance function absorbs the remaining 10 axes and observed controls.
- Estimation model: 5-fold group K-fold cross-fitting keyed on source ID, LightGBM nuisance regressors, Neyman-orthogonal scores, and source-clustered standard errors support estimation of τB and τW.A separate full-sample estimate supplies τ̂pool for comparison.
Results
Results show that pooled visual-style coefficients can blend opposing between-creator audience-preference and within-creator creative-effect signals, sometimes prescribing the wrong direction. DIRECT’s within-creator rules improve held-out engagement and remain robust across specifications, while diagnostics quantify the mismatch and its limits.
- Channel decomposition: On skin tone, the pooled estimate favors lighter representations, but the within-creator estimate favors darker tones.The pooled estimate is τ̂pool = +0.95 (p < 0.05), whereas τ̂W = −0.64 (p < 0.001), with darker tones receiving about 1.8% higher engagement per within-creator standard deviation.
- Held-out validation: 30.7% of achievable engagement gain is forgone by using pooled rather than DIRECT rules under the magnitude rule, versus 18.2% under the sign rule.Pooled rules deliver +6.94 to +7.94% engagement gain per standard deviation, while DIRECT rules deliver +8.48 to +11.47%.
- Benchmark and identification: 9 of 11 axes have the same sign under within-creator fixed-effects OLS and DIRECT, but OLS reverses creative direction on soft lighting and selfie/posed.Where signs agree, OLS magnitudes exceed DIRECT’s by 40 to 100%, and by roughly 1,300% on complex-versus-simple.
Discussion and Implications
DIRECT recommends separating creative effects from audience-preference context in analytics and using diagnostics to determine when pooled prescriptions require decomposition. The framework also clarifies reporting standards, exposes fairness-relevant ambiguity, and acknowledges limits from confounding and narrow data coverage.
- Practice: Creative analytics systems should report τ̂W as the creative-effect signal and τ̂B as audience-preference context instead of one leaderboard.This separates the two effects that pooled estimation averages together.
- Practice: 9 of 11 axes fall outside D1′, where pooled prescription is directionally correct though potentially mis-sized; axes within D1′ require decomposition.The diagnostics provide a fast screen for deciding when decomposition is necessary.
- Research: 4-of-11 reversals and a 30.7% conflation cost show why attribute-engagement coefficients should distinguish descriptive from prescriptive aims and report panel-based decomposition when prescriptive.These figures are specific to this panel but indicate that the gap matters for findings embedded in platform tooling.
- Fairness: DIRECT does not eliminate fairness concerns, but it reveals that positive pooled coefficients for lighter representations may reflect audience-recruitment history.The skin-tone case connects the distinction to algorithmic bias in commercial visual advertising systems.
- Limitations: Residual within-source confounding cannot be fully ruled out, and the data cover one category, platform, and time window.A within-creator natural experiment and cross-category replication would address both limitations.