Source-linked AI summary
PACEShop: Evaluating Personalized, Actionable, Compositional, and Evidence-grounded Shopping Assistants
Weimin Lyu, Chen Luo, Guangrui Li, Yaochen Xie, Dhineshkumar Ramasubbu, Arief Koesdwiady, Wanqiu Long, Hansu Gu, Yutong Chen, Zheshen Wang, Dakuo Wang, Yi Liu
TL;DR
The paper addresses the lack of a joint evaluation target for structured shopping-assistant responses that must satisfy shopper context, component consistency, evidence support, and actionable diagnosis. It introduces PACESHOP and PACEJUDGE to measure and report these properties, finding that generic judges often miss diagnostic fields while PACEJUDGE improves closure metrics without retraining. The benchmark and protocol are bounded by synthetic construction, limited evidence depth, narrower holdout coverage, and English-only data.
Problem
Existing personalization, grounding, and generic judging benchmarks do not jointly evaluate persona validity, cross-component consistency, evidence support, and localizable defects in structured shopping responses.
Method
The paper introduces PACESHOP, a controlled benchmark with structured personas, auditable evidence, GOOD/BAD labels, and gold defect annotations, plus PACEJUDGE, a training-free structured judging protocol.
Results
Generic judges often recognize broad GOOD/BAD quality but fail to recover diagnostic fields, while PACEJUDGE improves persona-source, cross-component, grounding, and family/location closure without retraining.
Takeaways & Limitations
Realistic evaluation of structured shopping assistants requires a task-aligned diagnostic output contract rather than only scalar quality scoring or a stronger backbone.
Takeaways & Limitations
The benchmark relies on synthetic GOOD/BAD construction, has shallow evidence coverage, narrower persona holdouts, and English-only data.
Abstract
from arXiv · showhide
Shopping assistants are shifting from ranked product lists toward structured decision support, where systems must synthesize shopper context, product evidence, and next-step guidance into a coherent recommendation experience. This changes the unit of evaluation: a fluent response can still fail by ignoring shopper context, contradicting itself across components, or leaving defects too vague to localize. Existing personalization, grounding, and LLM-as-a-judge benchmarks cover pieces of this problem, but they do not define a joint evaluation target for structured shopping-assistant responses. We formulate this missing evaluation target as PACE: Personalized, Actionable, Compositional, and Evidence-grounded evaluation. We instantiate PACE with two artifacts: PACEShop, a benchmark dataset that makes the target measurable through 22,625 controlled records with structured personas, auditable evidence pools, GOOD/BAD labels, and gold defect family and location annotations; and PACEJudge, a training-free judging protocol that makes the target reportable through a structured output contract. Our experiments show that generic judges can recognize broad quality but fail to recover the diagnostic fields required for PACE; PACEShop makes these failures verifiable, and PACEJudge improves persona-source, cross-component, grounding, and family/location closure without retraining, showing that realistic shopping-assistant evaluation requires a task-matched output contract rather than only a stronger backbone or scalar prompt.
1 Introduction
Structured shopping assistants require evaluation of persona-conditioned, multi-field responses rather than fluency or broad relevance alone. The paper formulates PACE and introduces PACESHOP and PACEJUDGE to make these failures measurable and diagnostically reportable.
- Motivation: Structured shopping responses can recommend avoided brands, contradict fields, or leave defects unlocalized despite appearing coherent.These failures make overall quality scores insufficient for practical diagnosis.
- PACE formulation: PACE defines evaluation as Personalized, Actionable, Compositional, and Evidence-grounded.The dimensions require validity to depend on shopper context, failure localization, structured-response consistency, and support from product evidence and persona history.
- Contributions: PACEShop makes PACE measurable with structured personas, candidate responses, auditable evidence pools, GOOD/BAD labels, and gold defect annotations.It addresses the absence of a controlled setting jointly exposing persona validity, cross-component consistency, evidence support, and localizable defects.
- Contributions: PACEJudge makes PACE reportable through a training-free structured judging protocol that requires violated dimensions, broken fields, and supporting evidence.Its backbone-agnostic design separates model capability from the diagnostic output contract.
- Results: Across multiple backbones, generic judges often recognize broad GOOD/BAD quality but fail on persona-source, cross-component, grounding, and actionable defect diagnosis.PACEJudge improves these contract-dependent metrics without retraining.
2 Related Work
Prior work separately studies personalization, grounding, commerce support, and generic LLM judging. The paper argues that these resources do not jointly evaluate the four PACE targets in structured shopping-assistant responses.
- Personalization and commerce: Personalized LLM benchmarks evaluate user-conditioned generation, preferences, memory, likability, and tool use.These resources focus on user adaptation and related behavioral signals.
- Benchmark gap: Table 1 compares existing datasets with the four PACE targets and identifies PACESHOP as jointly covering them.Its coverage uses structured personas, multi-component responses, auditable evidence pools, and gold defect family/location labels.
- Personalization and commerce: Commerce-oriented resources add e-commerce support, search-augmented chat, or online-shopping behavior signals but do not provide the joint PACE evaluation unit.They cannot directly test personalization, actionability, compositional consistency, and evidence grounding together.
- LLM judging and grounding: Generic LLM-as-a-judge methods established rubric-based, scalar, and pairwise evaluation for open-ended outputs.Subsequent work examines judge reliability, alignment, bias, distributional validity, and domain-specific protocols.
3 PACESHOP Dataset
PACESHOP constructs controlled, auditable records for evaluating whether judges recognize GOOD/BAD responses and diagnose failures by family and location. Its pipeline varies personas against fixed query evidence, validates GOOD responses, and injects labeled defects.
- Task formulation: PACESHOP evaluates the generated response layer rather than the full ranking interface or end-to-end shopping session.This scope isolates judging validity and failure diagnosis given inputs, responses, personas, and evidence.
- Task formulation: Each record supplies a query, structured persona, candidate response, auditable evidence pool, and GOOD/BAD validity label.For BAD responses, judges additionally predict a defect family and response-field location.
- Persona and evidence setup: 1,132 retained queries paired with distinct personas produce 4,525 validated query–persona pairs from 97,227 normalized queries, 1,200 personas, and 4,460 evidence records.The repeated-measures design fixes query and evidence while varying shopper context.
- Construction pipeline: Validated GOOD responses become anchors for controlled BAD construction after schema, cardinality, attribution-ID, and parsing checks.BAD variants contain one through four defects drawn from seven defect families with known field locations.
- Construction pipeline: Figure 2 links persona, behavior, intent, and evidence assignment to GOOD generation, controlled BAD construction, validity checks, measurable PACE scenarios, and structured judging outputs.The design uses auditable records and verification layers to expose PACE failures.
- Dataset statistics: 22,625 records comprise 4,525 GOOD and 18,100 BAD examples, with BAD records evenly distributed across single-, double-, triple-, and quad-defect tiers.Every BAD record carries gold defect family and location annotations.
4 PACEJUDGE Evaluation Protocol
PACEJUDGE evaluates shopping-assistant responses as structured, persona-conditioned, evidence-grounded objects and reports diagnostic fields beyond a GOOD/BAD verdict. Its structured contract covers PACE scores, defect family and location, supporting evidence, confidence, and rationale, while results show stronger PACE closure than generic judging protocols.
- Protocol overview: PACEJUDGE evaluates each candidate response as a structured, persona-conditioned, evidence-grounded object rather than a single fluent text.This addresses scalar judging’s inability to identify shopper-context mismatch, cross-component inconsistency, unsupported evidence, or unlocalizable defects.
- Output schema: The protocol returns a verdict, BAD probability, PACE scores, an auxiliary format/safety score, and a diagnostic object.The diagnostic object includes a predicted defect family, response location, supporting evidence IDs, confidence, and a short rationale.
- Output schema: GOOD verdicts require NONE for defect family and location, whereas BAD verdicts require a primary defect family and response field.The response-field set includes overview, categories, related queries, and attributions; defect families come from a seven-family taxonomy.
- Evaluation results: PACEJUDGE achieves the strongest overall performance among the compared multi-target rubric and single-target methods on PACE-closure evaluation.Table 3 averages results over seven LLM backbones and distinguishes broad GOOD/BAD balanced accuracy from diagnostic closure metrics S1–S4.
- PACE coverage: PACEJUDGE maps persona, actionability, compositionality, and evidence-grounding to explicit scores or diagnostic constraints.Persona uses sP; actionability requires what failed, where, and supporting evidence; compositionality uses sC and location; evidence-grounding uses sE with supporting IDs constrained to the evidence pool.
- Evaluation results: The structured output contract enables scenario-level evaluation of diagnostic fields that scalar judges do not require.These fields support evaluation beyond GOOD/BAD accuracy, including violated targets, broken components, defect families, and evidence support.
5 Experiments
The experiments test whether scalar or generic judges recover the diagnostic fields required for PACE, using PACEShop across multiple backbones and strict closure scenarios. PACEJUDGE performs best on task-aligned diagnosis, while backbone capability still limits absolute closure.
- Evaluation setup: The evaluation compares scalar and generic judges with PACEJUDGE across seven LLM backbones and scenarios targeting PACE diagnostic fields.S0 measures broad GOOD/BAD discrimination, while S1–S4 test persona-source diagnosis, cross-component localization, grounding, and actionable defect prediction.
- Evaluation setup: Strict paired-counterfactual gates require target-specific correctness, localization, consistency, evidence containment, and GOOD-record specificity beyond BAD detection.These gates prevent an always-BAD judge from receiving high credit through over-flagging.
- Main results: Broad GOOD/BAD discrimination is not the main difficulty; judges break down when evaluation requires identifying the violated target, failed response field, and evidence-grounded diagnosis.A judge can recognize that a response is flawed while failing to provide the fields needed for debugging.
- Main results: PACEJUDGE improves diagnostic closure rather than merely flagging more responses as BAD, including under strict paired-counterfactual audits.The audit requires GOOD-record specificity, correct target, family and location, axis consistency, and evidence-ID containment.
- Main results: PACEJUDGE achieves the strongest overall PACE average among multi-target rubric ablations and is strongest on S1–S4 diagnostic closure metrics.Figure 3 summarizes persona-source diagnosis, cross-component localization, grounding control, and actionable defect localization.
- Main results: Protocol design helps, but backbone capability sets the ceiling: Sonnet 4.6 and Opus 4.7 are strongest, while smaller or open-weight backbones remain unstable on several diagnostic controls.The strongest frontier backbones combine structured-contract following, GOOD-record specificity, and reasoning over persona and field-level constraints.
6 Conclusion
The paper frames shopping-assistant evaluation as a joint PACE problem and introduces PACEShop and PACEJudge to measure and diagnose it. Results show that task-aligned structured diagnosis matters more than broad quality scoring alone.
- PACE treats shopping-assistant responses as personalized, actionable, compositional, and evidence-grounded objects.
- PACEShop makes PACE measurable with controlled records, structured personas, auditable evidence pools, and gold defect family/location labels.
- Scalar and generic rubric judges often recognize broad GOOD/BAD quality but fail to recover the diagnostic fields needed for debugging.
- PACEJudge improves closure metrics without retraining.
- The results identify task-aligned diagnosis of what failed, where, and why as the central evaluation challenge for structured shopping assistants.
Limitations
PACEShop and PACEJudge have bounded limitations involving synthetic construction, shallow evidence coverage, limited persona holdouts, and English-only scope.
- GOOD responses come from one generator and BAD responses from rule-based defect injection, so the data may not fully reproduce live traffic.Cross-generator and real-assistant slices provide audits but do not fully reproduce live traffic.
- Most evidence records contain product-listing fields, with only 129/4,460 review-augmented records.This bounds how deeply grounding can be audited per claim.
- Held-out persona bundles cover 345/1,200 personas across eight families, narrowing generalization beyond the benchmark.
- The dataset is English-only.
Potential Risks
The artifacts are intended to evaluate structured shopping-assistant responses, not to certify production systems or replace human review. Careless transfer of synthetic persona evaluation could also encourage exploitation of user traits.
- PACEShop and PACEJudge should not directly certify production systems or replace human review.
- Synthetic or structured personas can encourage systems to infer or exploit user traits if transferred carelessly to real deployments.
- Practical use should avoid sensitive attributes, preserve privacy, and treat persona-conditioned evaluation as misalignment detection rather than persuasion maximization.
A Existing Benchmark Datasets
Prior personalization, grounding, shopping, and judging resources measure important components but do not define PACEShop’s joint persona–compositional–grounded–actionable diagnosis. The appendix therefore preserves native protocols, uses only defensible single-target adaptations, and documents construction for auditability.
- Existing Benchmark Datasets: Prior resources target different native tasks, including user-conditioned generation, dialogue preference following, memory retrieval, task success, pairwise search comparison, and RAG faithfulness.
- Existing Benchmark Datasets: These benchmarks do not provide the joint output fields and closure metrics required for persona-aware, compositional, grounded, actionable diagnosis.
- Existing Benchmark Datasets: Personalization benchmarks usually lack persona-swap defect families and response-field locations, while RAG evaluations usually omit invented shopper history and evidence-pool citation checks.
- Existing Benchmark Datasets: PACEShop does not force every prior resource into every scenario; it marks unsupported native judging protocols as “–”.
- Existing Benchmark Datasets: The implemented adapted baselines map ARES-style evaluation to evidence grounding, PersonaLens-style evaluation to personalization, and EtaPP-style diagnosis to actionable key points.
- Existing Benchmark Datasets: The construction appendix makes the benchmark auditable through validity layers, concrete records, JSON examples, prompts, holdout statistics, and source-coverage tables.
B.1 PACESHOP Dataset Statistics
PACEShop is an auditable benchmark built from structured personas, evidence-backed shopping queries, and controlled GOOD/BAD response records. Its construction and validation layers make personalized, compositional, and evidence-grounded defects measurable and localizable.
- Benchmark design: PACEShop organizes benchmark records around structured personas, auditable evidence pools, candidate responses, GOOD/BAD labels, and defect annotations.BAD records are generated deterministically from a source GOOD record, assigned persona, and defect family, producing gold location labels without a separate validator.
- Persona pool: 1,200 personas encode shopper profiles, constraints, preferences, brand affinities, purchase histories, and recent searches.This supports both persona-conditioned correctness and auditable checks for invented or unsupported shopper-history claims.
- Queries and evidence: 97,227 normalized US shopping queries and 4,460 evidence records are filtered to 1,132 queries with at least three evidence-backed products.The catalogs cover 11 retail verticals, 6 shopping missions, and 11 shopping domains; 129 evidence records include review snippets.
- Construction workflow: The construction pipeline normalizes catalogs, builds personas and evidence pools, assigns covered query-persona combinations, generates constrained responses, and packages validation assets.It includes schema validation, deterministic post-processing, four defect tiers, and released JSONL and calibration slices.
- Dataset statistics: Coverage varies across retail verticals and shopping missions because evidence availability, rather than intentional stratification, determines query retention.For example, Beauty & Health retains 33 queries, representing 1.14% of its backbone queries.
C.2 Native Judging-Protocol Map and Evaluation-Target Leakage Audit
The audit compares multi-target rubric judges with faithfully adapted single-target baselines under a shared structured evaluation contract. It shows why native protocols must be restricted to their supported PACE targets and why diagnostic output fields matter.
- Audit findings: Most judges can distinguish broadly GOOD from BAD responses, but diagnostic closure fails when evaluation requires the violated target, broken field, and evidence-grounded diagnosis.PACEJUDGE is strongest among multi-target rubric ablations, while single-target baselines leave unsupported columns or risk target leakage.
- Protocol comparison: Multi-target rubric ablations share the PACEJUDGE inputs and output contract, while ARES-style, PersonaLens-style, and EtaPP-style baselines are scored only on their native PACE targets.The adapted baselines abstain on unsupported targets, with target coverage and leakage documented through the protocol map and audit.
- Diagnostic contract: PACEJUDGE predicts verdict, PACE scores, defect family, defect location, supporting evidence IDs, confidence, and rationale instead of only a scalar quality score.This structured contract turns evaluation into fault localization across seven defect families and four response-field groups.
- Single-target adaptations: ARES-style maps context relevance, answer faithfulness, and answer relevance to evidence grounding, emitting evidence defects only when its grounding score is below 3.0.Persona, compositional, and format/safety axes are fixed at 3.0, and the baseline reports only S3 and S3E.
- Single-target adaptations: PersonaLens-style checks persona slots such as constraints, brand preferences, attributes, and budget, but pins compositional and evidence axes at 3.0 and abstains elsewhere.Its defect vocabulary is restricted to persona-target labels, including explicit conflicts and invented purchase or search history.
D.4 PACEJUDGE Evaluation Protocol and Prompt
PACEJUDGE evaluates each PACESHOP record as a structured, persona-conditioned object and returns a closed, evidence-grounded diagnostic contract rather than a scalar impression.
- D.4 PACEJUDGE Evaluation Protocol and Prompt: The same training-free prompt and PACESHOP input interface are used across every backbone and reported PACEJUDGE row.This makes the judging protocol reproducible across the seven-backbone experiments.
- D.4 PACEJUDGE Evaluation Protocol and Prompt: PACEJUDGE takes typed record fields including the query, persona, response, evidence pool, used evidence, and deterministic checks.The response is supplied in both XML and structured JSON views.
- D.4 PACEJUDGE Evaluation Protocol and Prompt: Closed defect-family and response-location enumerations prevent unsupported free-form diagnostic labels.The taxonomy contains seven defect families plus NONE, and four response-field groups plus NONE.
- D.4 PACEJUDGE Evaluation Protocol and Prompt: Decision rules require GOOD only for correct, coherent, grounded, and structurally valid responses, while supporting evidence IDs must come from the candidate pool.Confidence, probabilities, and axis scores are also constrained to specified numeric ranges.
- D.4 PACEJUDGE Evaluation Protocol and Prompt: A JSON-Schema contract requires verdict, BAD probability, four axis scores, and an actionable defect family, location, evidence set, and confidence.The actionable object encodes defect type, response location, supporting evidence, confidence, and related fields.
- D.4 PACEJUDGE Evaluation Protocol and Prompt: The protocol evaluates persona alignment, cross-component consistency, evidence grounding, and format safety.Its prompt explicitly instructs judges to reason over these four dimensions.
- D.4.1 PACEJUDGE Output Examples: Qualitative outputs demonstrate actionable diagnosis: axis scores identify failed dimensions, while family and location support debugging.One example assigns Consistency = 1.5 to a mismatch between toy categories and costume-themed related queries.
- E.4 Main Table Metrics Definitions: The main metrics separate broad GOOD/BAD accuracy from diagnostic closure across persona source, component localization, grounding, and family/location prediction.S1–S4 are computed on a 3,400-record held-out test set, with false-fire checks preventing indiscriminate defect prediction.
E.6 Expanded Experiments Findings
The expanded experiments show that generic judges often recognize broad validity but struggle with the diagnostic fields required to close PACE targets. PACEJUDGE improves this structured diagnosis, while specificity and backbone capability remain important.
- Finding 1: S0 shows that scalar and generic prompts can perform reasonably on broad GOOD/BAD discrimination, but PACE closure requires additional diagnostic fields.The harder requirements include persona-source diagnosis, cross-component localization, grounding control, and defect-family/location prediction.
- Finding 2: S1 requires identifying PREF_CONFLICT on persona-swapped records rather than merely marking them BAD.The unchanged query and response isolate the swapped persona as the source of invalidity.
- Finding 3: S2 and S3 expose gaps between detecting a flawed response and localizing its failed component or verifying both product-evidence and persona-history grounding.The scenarios are designed to test compositional localization and grounding beyond broad error recognition.
- Finding 4: S4 requires a debugging-ready diagnosis containing both the defect family and the response-field location.This makes actionable evaluation stricter than deciding only whether a response is BAD.
- Finding 5: False-fire checks ensure that closure scores cannot be obtained simply by predicting defects on most records.Tab. 3 measures defect alarms on GOOD records, while Tab. 4 incorporates paired specificity into per-target gates.
- Finding 6: Strict paired-counterfactual gates preserve the conclusion by requiring joint success across paired GOOD/BAD records and target-specific diagnosis.These gates reduce loopholes such as always predicting BAD or assigning generic defect labels.
E.7.1 Concrete Benchmark Examples
The benchmark examples connect PACESHOP’s controlled records to aggregate evaluation: galleries expose concrete defects, persona swaps isolate conflict sources, and tier analyses reveal judge behavior across defect counts.
- Concrete Failure Cases: Concrete failure cases pair gold defect families and locations with PACEJUDGE verdicts, axis scores, and predicted diagnoses.The ten cases span all four response-field groups and all seven defect families.
- Persona-Swap Conflict Examples: Persona-swap examples change one of four persona axes while leaving the response unchanged, flipping paired records from GOOD to BAD.The axes are budget, hard constraint, brand affinity, and attribute preference.
- Expanded Evidence: The supporting audits and galleries explain why the compact S0–S4 tables should be read as PACE-closure results rather than ordinary scalar-judge scores.They include leakage, false-fire, per-backbone stability, tier difficulty, qualitative outputs, and record-level checks.
- Tier Difficulty: Deterministic-baseline accuracy rises from 0.880 for single defects to 0.975 for quad defects on the development set.The increase is attributed to more defects triggering more rule-based checks.
- Tier Difficulty: Qwen3 32B accuracy rises from 0.740 for single defects to 0.990 for quad defects, illustrating a steeper tier gradient for smaller models.The paper contrasts this with flatter tier profiles for frontier models.