Source-linked AI summary
Personalizing Text-to-Image Generation to Individual Taste
Anne-Sofie Maerten, Juliane Verwiebe, Shyamgopal Karthik, Ameya Prabhu, Johan Wagemans, Matthias Bethge
TL;DR
T2I systems optimize average appeal rather than individual taste, motivating personalized preference modeling. PAM∃LA provides a multi-user benchmark and user-conditioned predictor, which improves preference prediction and enables steering toward user-specific visual preferences.
Problem
T2I models and reward systems optimize for crowd appeal or aggregated quality, despite subjective preferences causing different users to want different results from identical prompts.
Method
PAM∃LA introduces a curated personalized-rating dataset and a predictor conditioned on image, metadata, and user information, trained with complementary aesthetic datasets.
Results
PAM∃LA outperforms existing reward models and user studies find highest preference for images optimized to the evaluating participant’s own taste, scoring 1065.
Takeaways & Limitations
User-conditioned prediction and prompt optimization can steer image generation toward individualized visual qualities rather than generic or stylized outputs.
Takeaways & Limitations
Preference prediction remains difficult when users strongly disagree because near-ties on continuous rating scales introduce evaluation noise.
Abstract
from arXiv · showhide
Modern text-to-image (T2I) models generate high-fidelity visuals but remain indifferent to individual user preferences. While existing reward models optimize for "average" human appeal, they fail to capture the inherent subjectivity of aesthetic judgment. In this work, we introduce a novel dataset and predictive framework, called PAMELA, designed to model personalized image evaluations. Our dataset comprises 70,000 ratings across 5,000 diverse images generated by state-of-the-art models (Flux 2 and Nano Banana). Each image is evaluated by 15 unique users, providing a rich distribution of subjective preferences across domains such as art, design, fashion, and cinematic photography. Leveraging this data, we propose a personalized reward model trained jointly on our high-quality annotations and existing aesthetic assessment subsets. We demonstrate that our model predicts individual liking with higher accuracy than the majority of current state-of-the-art methods predict population-level preferences. Using our personalized predictor, we demonstrate how simple prompt optimization methods can be used to steer generations towards individual user preferences. Our results highlight the importance of data quality and personalization to handle the subjectivity of user preferences. We release our dataset and model to facilitate standardized research in personalized T2I alignment and subjective visual quality assessment.
1 Introduction
PAM∃LA addresses the gap between global preference alignment and individual aesthetic taste by combining user-conditioned prediction with prompt-based steering. Its benchmark and predictor are designed to model subjective evaluations and tailor generated images to particular users.
- T2I models optimize for crowd appeal, so identical prompts can produce outputs misaligned with different users’ subjective aesthetic preferences.
- Existing reward models learn an aggregated notion of quality from often outdated, uncurated AI images, which may steer generations toward older artifacts.
- PAM∃LA introduces around 70,000 ratings for 5,077 images, with each image scored by 15 users across subjective visual domains.
- Prompt optimization with PAM∃LA steers image generation toward individual taste, and user studies report preferences for images optimized for the evaluating users.
- The PAM∃LA predictor conditions preference prediction on image, metadata, and user information, outperforming existing reward models.
2 Related Work
Prior work includes global reward models, personalized image-aesthetics methods, and preference-tuning approaches, but PAM∃LA targets personalized evaluation of AI-generated images. Its benchmark combines subjective domains, dense multi-rater coverage, and user demographics.
- Reward Models for Text-to-Image Generation: Global T2I reward models such as ImageReward, PickScore, and HPS learn from large-scale preference data to improve semantic alignment and generic visual quality.
- Reward Models for Text-to-Image Generation: PAM∃LA is the only compared dataset combining AI-generated images, subjective visual domains, dense per-image multi-rater coverage, and user demographics.
- Personal Preference Prediction: Personalized image-aesthetics research models user-specific preferences using image attributes, demographic metadata, personality traits, graphs, or interaction matrices.
- Personal Preference Prediction: PAM∃LA uses individual preferences with pretrained visual and language backbones to support personalized prediction and steering toward unseen users.
- Personal Preference Prediction: The benchmark spans Art and Photography across 21 thematic categories, separating stylized artistic judgment from photographic evaluation.
3 Personalized Image Evaluation
PAM∃LA constructs a diverse, multi-user image-rating benchmark and predicts personalized scores by fusing image, prompt, metadata, demographic, and user information. Evaluation uses held-out users and compares the model with population-level reward baselines.
- Dataset Construction: PAM∃LA contains 5,077 images from artistic and photorealistic domains generated by Flux 2 and Nano Banana, varying visual style and semantic content.
- Dataset Construction: Users rate isolated images on a five-anchor aesthetic slider, while the dataset records demographic metadata such as age, gender, and nationality.
- Dataset Construction: Models predict a user rating from a prompt, image, and demographic profile in seen-user and few-shot unseen-user settings.
- PAM∃LA Predictor: The predictor fuses frozen SigLIP2 image and prompt features with projected metadata, demographic, and user embeddings through a shallow transformer and regression head.
- PAM∃LA Predictor: Training jointly uses PAM∃LA, LAPIS, and PARA to span AI-generated images, artworks, and photographs.
- PAM∃LA Predictor: For unseen users, the method retrieves similar training users from few-shot context embeddings and interpolates their learned embeddings.
- Evaluation: The model outperforms all baselines on held-out users across user-level and population-average SROCC, PLCC, and pairwise accuracy.
4 Experiments: Personalized Reward Modeling
PAM∃LA evaluates whether explicitly modeling individual preferences improves aesthetic prediction and compares it with population-level reward models. The experiments report significant gains in both user-level and population-level metrics, alongside more natural prompt-optimized images.
- PAM∃LA is evaluated against population-level reward models on user-level and population-level aesthetic metrics.The comparison includes LAION-Aesthetics, ImageReward, Q-Align, DeQA, and HPSv3 on held-out unseen users.
- PAM∃LA-steered prompt refinement produces more natural and appealing samples without reward-hacking, unlike global reward models.
- Explicitly modeling individual preferences enables PAM∃LA to achieve significant improvements in both population-level and user-level aesthetic prediction.
5 Experiments: Personalized Image Steering
The paper uses reward-driven iterative prompt optimization to steer image generation with PAM∃LA. Compared with generic reward models, the personalized predictor preserves realism and adapts composition, lighting, saturation, and other visual attributes to users and demographic groups.
- Experimental setup: PAM∃LA is compared with HPSv3 and Q-Align using reward-driven iterative prompt optimization.Each iteration generates prompt variations, renders candidate images, scores them, and retains the top-scoring prompt for subsequent refinement.
- Comparing reward models: PAM∃LA maintains high-fidelity photorealism while steering composition through lighting, camera angle, and viewpoint rather than artificially boosting saturation.
- Preserving realism: For surreal prompts, PAM∃LA steers outputs toward physically plausible textures and lighting instead of a heavily stylized digital-art aesthetic.
- Individual user profiles: PAM∃LA produces distinct visual qualities for different users, including different camera angles and lighting conditions for identical base prompts.
- Demographic profiles: Age-conditioned optimization yields higher color saturation for younger users, indicating demographic differences in optimization trajectories.
6 Validating PAM∃LA
The validation studies test whether PAM∃LA optimization improves images for target users and whether its discovered preferences are reproducible. Personalized outputs receive the strongest user preference scores, while generic reward-model optimization ranks below the unoptimized baseline.
- User study setup: The user study compares images optimized for the evaluating user, other users, HPSv3, Q-Align, and an unoptimized baseline.It collects 15,300 ratings across 7,650 pairwise comparisons, 18 prompts, and 6 users.
- Preference ranking: 1065: images personalized to the evaluating participant achieve the highest Bradley–Terry preference score, ahead of other-user personalization at 1038.Both personalized conditions significantly outperform the unoptimized image at 1016 based on non-overlapping 95% confidence intervals.
- Preference ranking: 959: HPSv3-optimized images and 922: Q-Align-optimized images rank below the unoptimized baseline at 1016.The reported rankings indicate lower perceptual quality for both generic reward-model conditions.
- Validation summary: Users prefer PAM∃LA-optimized images over unoptimized baselines, while generic reward-model optimization degrades perceptual quality.
- Reproducibility: Independent optimization runs converge to the same compositional and stylistic elements for each user despite different stochastic prompt proposals.Examples include repeated keywords for User 1, low-angle compositions with vibrant colors for User 2, and warm HDR lighting for User 3.
7 Analysis
The analysis examines difficult preference conflicts and near-ties in pairwise evaluation. PAM∃LA predicts conflicting user preferences above chance and approaches 80% accuracy when marginal score differences are treated as functional ties.
- Diverging user preferences: PAM∃LA assigns different reward scores to the same image according to user profiles, enabling evaluation on genuinely conflicting preferences.
- Diverging user preferences: 61.44%: PAM∃LA correctly predicts diverging preference pairs for seen users, compared with 55.23% for unseen users.The evaluation covers 13,000 unseen and 71,700 seen pairwise preferences, where random guessing gives 50% accuracy.
- Near-ties: Treating small score differences as functional ties raises PAM∃LA pairwise accuracy to nearly 80%.The margin threshold excludes pairs whose absolute rating difference falls below the threshold, focusing evaluation on unambiguous preferences.
- Near-ties: The accuracy gains after removing near-ties suggest that continuous-score evaluations penalize models for predicting marginal differences that may reflect random rating variance.
- Analysis summary: PAM∃LA learns robust representations that distinguish clear preferences when the underlying user signal is strong.
8 Conclusion
PAM∃LA shifts text-to-image evaluation toward individualized preference alignment and consistently steers generation toward users’ tastes. However, predicting preferences remains difficult when users strongly disagree, partly because near-ties add noise to continuous ratings.
- PAM∃LA combines 70,000 personalized ratings with a user-conditioned preference predictor to model individual taste and steer image generation.
- The method alters compositional and visual qualities during optimization, while existing approaches produce oversaturated, generic images.
- Human evaluators systematically prefer PAM∃LA’s photorealistic outputs over the stylized “AI look.”
- Preference prediction remains difficult when users strongly disagree, and near-ties on continuous rating scales introduce evaluation noise.
A.1 Ablations
The ablation study evaluates how input modalities affect prediction for seen and unseen users. The full model is selected because it achieves the smallest generalization gap and the most stable performance across user populations.
- The full model achieves a 0.059 generalization gap, compared with 0.084–0.113 for other ablations.The gap measures the difference between seen- and unseen-user performance.
- Removing metadata costs −0.043 SROCC on unseen users, making metadata the largest benefit for that group.
- Removing user identity costs −0.048 SROCC for seen users, making it the most critical component for that group.
- Demographics contribute −0.008 SROCC on unseen users, while jointly removing metadata and demographics costs −0.025 SROCC.The joint ablation indicates partially overlapping contributions.
- Removing text captions slightly degrades unseen performance by −0.019 SROCC but improves seen-user scores.This pattern suggests prompt text can introduce noise for users whose preferences the model has not learned directly.
- The full model uses all input modalities to balance performance between seen and unseen users.
A.2 Hyper-parameter Tuning for Unseen User Evaluation
Unseen users are represented through preference profiles built from rated images, then matched to training users for embedding interpolation. Hyperparameter tuning identifies the configuration used for unseen-user evaluation.
- Performance peaks at N = 15 and K = 5 across all metrics, with diminishing returns beyond this point.Temperature is fixed at 0.1 in the figure and reported configurations.
- Unseen-user profiles are constructed from rating-weighted SigLIP embeddings of rated images.
- The method retrieves K similar training users by cosine similarity and interpolates their learned participant embeddings with softmax weighting.
- The grid search varies the number of context images N across {5, 10, 15, 20, 25}, neighbors K across {1, 5, 10}, and temperature τ across {0.05, 0.1, 0.2}.
B Methods
The methods use continuous aesthetic ratings and pairwise preference judgments to collect annotations and validate personalized steering. Prompt optimization changes composition, photographic settings, and style while preserving semantic content, with repeated runs showing consistent steering patterns.
- Users rate image aesthetic value on a continuous slider with five anchor points, considering beauty, liking, preference, and attraction.Practice trials familiarize users with the task before rating begins.
- Prompt refinement encourages changes in composition, lighting, camera settings, color, mood, and other stylistic elements while keeping semantic content unchanged.
- The steering validation study presents two images side by side and collects preferences across all pairwise comparisons.Comparisons include user-optimized, other-user-optimized, generic-reward-optimized, and unoptimized images.
- Two runs for User 2 produce consistent steering patterns, indicating that the LLM learns recurring prompt patterns.
- Two runs for User 3 likewise produce consistent steering patterns across image-steering outcomes.