Source-linked AI summary

Neural Rating Regression with Abstractive Tips Generation for Recommendation

Piji Li, Zihao Wang, Zhaochun Ren, Lidong Bing, Wai Lam

arXiv:1708.00154v1cs.CLcs.AIcs.IR

TL;DR

Recommendation systems have used ratings, item specifications, and reviews, but not short tips that express users’ experience and suggestions. NRT jointly models rating prediction and abstractive tips generation from user and item representations, achieving better performance than state-of-the-art methods on both tasks while producing tips that vividly predict user experience and feelings.

  • Problem

    Existing recommendation models use item specifications or user reviews but do not consider tips, although tips and numerical ratings jointly express users’ product experiences and feelings.

  • Method

    NRT uses shared user and item latent factors with neural rating regression and gated recurrent networks for abstractive tips generation, trained jointly end-to-end.

  • Results

    NRT achieves better performance than state-of-the-art models on both rating prediction and abstractive tips generation across benchmark datasets.

  • Takeaways & Limitations

    Generated tips can vividly predict users’ experience and feelings while ratings and tips are modeled as two facets of product assessment.

  • Takeaways & Limitations

    NRT generally does not outperform baselines on recall because training tips are short and decoding favors short sentences.

Abstract

from arXiv · show

Recently, some E-commerce sites launch a new interaction box called Tips on their mobile apps. Users can express their experience and feelings or provide suggestions using short texts typically several words or one sentence. In essence, writing some tips and giving a numerical rating are two facets of a user's product assessment action, expressing the user experience and feelings. Jointly modeling these two facets is helpful for designing a better recommendation system. While some existing models integrate text information such as item specifications or user reviews into user and item latent factors for improving the rating prediction, no existing works consider tips for improving recommendation quality. We propose a deep learning based framework named NRT which can simultaneously predict precise ratings and generate abstractive tips with good linguistic quality simulating user experience and feelings. For abstractive tips generation, gated recurrent neural networks are employed to "translate" user and item latent representations into a concise sentence. Extensive experiments on benchmark datasets from different domains show that NRT achieves significant improvements over the state-of-the-art methods. Moreover, the generated tips can vividly predict the user experience and feelings.

1 INTRODUCTION

Recommendation systems commonly use ratings and increasingly incorporate item specifications or user reviews, but mobile E-commerce tips add short expressions of experience, feelings, and suggestions. NRT jointly models ratings and tips, using latent factors and neural generation to predict ratings and produce abstractive tips.

  • Motivation: Tips are short, single-topic texts that directly express user experience, feelings, or suggestions and provide quick insights beyond long reviews.They average about 10 words and complement item specifications and user reviews.
  • Research gap: Existing recommendation models use item specifications and user reviews to enhance latent factors and rating prediction, but do not consider tips.Item specifications describe item attributes, whereas reviews capture users’ usage experiences and preferences.
  • Proposed framework: NRT jointly predicts ratings and generates abstractive tips rather than extracting existing sentences.The generated sentence is intended to simulate how users express experience and feelings after consuming an item.
  • Model design: Gated recurrent neural networks translate user and item latent factors into concise tips, while a multilayer perceptron projects them into ratings.Latent factors and neural parameters are learned end-to-end through multi-task learning.
  • Evaluation: Benchmark experiments report better performance than state-of-the-art models on both rating prediction and abstractive tips generation.The framework is designed to predict precise ratings and generate tips with good linguistic quality.

2 RELATED WORKS

Prior recommendation research centers on collaborative filtering and latent-factor models, with later work incorporating item specifications, user reviews, and neural networks to address sparse or limited representations.

  • Collaborative filtering: Collaborative filtering methods primarily use historical ratings, while latent-factor models map users and items into a shared representation space for rating prediction.Matrix factorization methods include SVD, SVD++, NMF, and PMF.
  • Text-enhanced recommendation: Sparse rating matrices motivate methods that incorporate item specifications or user reviews into rating prediction.Specification-based approaches model item text, while review-based approaches derive user and item factors from review content.
  • Review-based methods: Review-aware models such as HFT, RMR, TriRank, and sCVR integrate topic models to construct latent factors from user reviews.TriRank and sCVR are also described as providing explanations through review-related modeling.
  • Neural recommendation: Neural-network approaches combine deep architectures with collaborative filtering to improve recommendation performance.The related work discusses neural models including restricted Boltzmann machines for modeling user interactions.

3 FRAMEWORK DESCRIPTION

NRT jointly models rating regression and abstractive tips generation from user and item latent factors. It uses neural transformations, GRU decoding, contextual signals, and shared latent representations trained across subtasks.

  • 3.1 Overview: At operation time, NRT receives only a user and an item, predicts a rating, and generates a concise abstractive tips sentence without review or tips text inputs.Training uses users, items, ratings, review content, and tips sentences.
  • 3.2 Neural Rating Regression: NRT maps user and item latent factors through multilayer nonlinear transformations to produce a real-valued rating.The model uses separate user and item latent spaces and projects them into a shared hidden space.
  • 3.3 Neural Abstractive Tips Generation: A GRU sequence decoder translates the user and item latent factors into tips words while incorporating rating sentiment and review-generation context.The context information initializes the decoder hidden state because the first decoding step has no word input.
  • 3.3 Neural Abstractive Tips Generation: The decoder predicts each next word from sequence hidden states using a softmax output over the review-and-tips vocabulary.The generative process is evaluated with Negative Log-Likelihood.
  • 3.3 Neural Abstractive Tips Generation: User and item latent factors are shared by rating prediction and text generation, so both subtasks provide training feedback to their representations.This shared multi-task design is trained end-to-end with the neural parameters and latent factors.

4 EXPERIMENTAL SETUP

The experiments evaluate NRT across multiple recommendation and tips-generation questions using benchmark datasets, standard metrics, and comparative baselines. The setup combines rating prediction, abstractive tips evaluation, and controlled training and implementation choices.

  • Research Questions: Experiments investigate rating prediction, abstractive tips generation, and the relationship between predicted ratings and generated-tip sentiment.The paper explicitly formulates these as RQ1, RQ2, and RQ3.
  • Datasets: Four benchmark datasets from different domains are used, including Amazon Books, Electronics, Movies & TV, and Yelp Challenge 2016.Ratings are integers from 0 to 5; Books is described as the largest Amazon domain dataset, while Yelp has the most users and is the sparsest.
  • Evaluation Metrics: Rating prediction is evaluated with Mean Absolute Error and Root Mean Square Error, while tips generation uses ROUGE-1, ROUGE-2, ROUGE-L, and ROUGE-SU4 precision, recall, and F-measure.ROUGE compares overlapping n-grams between generated tips and user-written ground-truth tips.
  • Datasets: The datasets are split into 80% training, 10% validation, and 10% testing, with dataset-specific vocabularies built after filtering low-frequency words.Model parameters are tuned on the validation set.

5 RESULTS AND DISCUSSIONS

NRT consistently improves rating prediction and abstractive tips generation across benchmark datasets, while its performance depends on decoding choices and reveals limitations in recall and sentiment alignment.

  • Rating Prediction: NRT consistently outperforms comparative methods on both MAE and RMSE across all datasets.A two-tailed paired t-test also finds NRT significantly better than RMR.
  • Rating Prediction: Text-aware topic-modeling methods outperform traditional rating-only collaborative-filtering models because they improve latent-factor representations.CTR and RMR outperform LRMF, NMF, PMF, and SVD++.
  • Abstractive Tips Generation: NRT achieves the best Precision and F1-measure among evaluated tips-generation methods across all four datasets.On Movies&TV, it also achieves the best Recall for every ROUGE metric.
  • Limitations and Case Analysis: NRT usually trails extraction-based baselines on Recall because short training tips and beam search favor short outputs.Extraction-based methods favor longer sentences even under a 20-word restriction.
  • Abstractive Tips Generation: Beam sizes β = 3–5 produce the best tips-generation performance on the Electronics and Movies&TV validation sets.The study evaluates β values from 1–5, 10, and 20.
  • Abstractive Tips Generation: Length-Normalization makes NRT much better on ROUGE F1-measures than the same model without it.It adjusts beam-search log-probabilities to consider longer sentences; the reported settings are n = 2 and α = 0.6.
  • Limitations and Case Analysis: Generated tips generally show good linguistic quality and sentiment consistency with predicted ratings, although some cases exhibit mismatches.Examples include negative tips paired with predicted ratings of 2.25 and 1.46, alongside a neutral tip rated 4.34.

6 CONCLUSIONS

NRT jointly predicts ratings and generates abstractive tips, using shared latent representations and GRU-based decoding. Experiments show better performance than state-of-the-art models on both tasks, while generated tips can predict user experience and feelings.

  • NRT simultaneously predicts ratings and generates abstractive tips with good linguistic quality.GRU with context information translates user and item latent factors into concise sentences.
  • The framework uses multi-task end-to-end learning for neural parameters and user and item latent factors.
  • NRT achieves better performance than state-of-the-art models on rating prediction and abstractive tips generation.
  • Generated tips can vividly predict user experience and feelings.
  • Table 10 presents predicted ratings and generated tips alongside their ground-truth counterparts.
Loading 1708.00154v1…