Source-linked AI summary

A Survey on Conversational Recommender Systems

Dietmar Jannach, Ahtsham Manzoor, Wanling Cai, Li Chen

arXiv:2004.00646v2cs.HCcs.AIcs.IR

TL;DR

One-shot recommender systems have limited access to current, contextual, or evolving preferences, motivating conversational recommender systems. The paper surveys CRS definitions, architectures, technologies, tasks, knowledge, and evaluation, finding broad technical diversity but limited CRS-specific research on explanations and uneven evaluation practice. It identifies helpfulness, user expectations, failure modes, supported intents, end-to-end learning, and evaluation as open research areas.

  • Problem

    One-shot recommendation may not reliably estimate preferences when past observations are unavailable, context-dependent, or still being constructed during decision-making.

  • Method

    The paper provides a detailed survey organized around CRS architectures, interaction modalities, knowledge and data, computational tasks, evaluation, and future directions.

  • Results

    The review finds broad technical diversity across CRS, but little CRS-specific research on explanations, limited explanation support, and challenging cross-system comparisons.

  • Takeaways & Limitations

    Future CRS research should clarify what makes systems helpful, what users expect, which intents to support, and how conversational recommendation should be evaluated.

  • Takeaways & Limitations

    The literature lacks a clear understanding of helpfulness, user expectations, failure, supported intents, and the usefulness of learning-based CRS, while current evaluations can be insufficiently insightful.

Abstract

from arXiv · show

Recommender systems are software applications that help users to find items of interest in situations of information overload. Current research often assumes a one-shot interaction paradigm, where the users' preferences are estimated based on past observed behavior and where the presentation of a ranked list of suggestions is the main, one-directional form of user interaction. Conversational recommender systems (CRS) take a different approach and support a richer set of interactions. These interactions can, for example, help to improve the preference elicitation process or allow the user to ask questions about the recommendations and to give feedback. The interest in CRS has significantly increased in the past few years. This development is mainly due to the significant progress in the area of natural language processing, the emergence of new voice-controlled home assistants, and the increased use of chatbot technology. With this paper, we provide a detailed survey of existing approaches to conversational recommendation. We categorize these approaches in various dimensions, e.g., in terms of the supported user intents or the knowledge they use in the background. Moreover, we discuss technological approaches, review how CRS are evaluated, and finally identify a number of gaps that deserve more research in the future.

1 INTRODUCTION

Conversational recommender systems address limitations of one-shot recommendation by supporting multi-turn interaction for eliciting preferences, explaining suggestions, and processing feedback. This survey reviews CRS approaches, technologies, evaluation, and future research gaps.

  • Challenges of one-shot recommendation: One-shot recommendation can fail when past behavior poorly captures preferences, context, or preferences formed during decision-making.These issues arise especially for high-involvement products, where past observations may be absent.
  • Conversational recommender systems: CRS support task-oriented, multi-turn dialogues that elicit current preferences, provide explanations, and process user feedback.This richer interaction is the central alternative to one-directional ranked recommendations.
  • Technological development: Advances in language technology, voice-controlled assistants, and chatbot use have increased interest in CRS.Machine-learning-based proposals are now more common than systems following entirely predefined dialogue paths.
  • Open challenges: A gap remains between current voice assistants and chatbots and the capabilities needed for truly conversational recommendation, particularly in voice-controlled settings.The survey identifies this gap as an ongoing challenge for CRS development.
  • Survey scope: The survey organizes CRS research by interaction modalities, knowledge and data, computational tasks, evaluation approaches, and future directions.Its review is structured around common building blocks of a typical CRS architecture.

2 DEFINITIONS AND RESEARCH METHODOLOGY

The paper defines CRS as systems that achieve recommendation-related goals through multi-turn dialogue and surveys their architectures, knowledge, and research literature. Its methodology identifies relevant works through database searches, manual screening, and snowballing.

  • Definition: A CRS is a software system that supports recommendation-related goals through a multi-turn dialogue.The definition is task-oriented and distinguishes CRS from one-shot question answering and general chat systems.
  • Definition: CRS support recommendation, preference acquisition, explanations, and dialogue-state management rather than merely producing one-shot answers.Their outputs and inputs may use voice, text, forms, buttons, gestures, speech, or multimedia.
  • Conceptual architecture: Typical CRS architectures combine dialogue management, user modeling, recommendation and reasoning, input processing, and output processing components.Dialogue management updates the dialogue state and user model, then selects recommendations, explanations, questions, or other content.
  • Knowledge elements: CRS commonly use an item database together with domain, background, and dialogue knowledge.Dialogue knowledge can encode states, supported intents, and possible transitions.
  • Research methodology: The review used predefined searches across several digital libraries, manual relevance checks, detailed reading, and snowballing.The process surfaced 121 CRS papers for consideration.
  • Research methodology: The review included papers compliant with the CRS definition and excluded one-shot or non-recommendation dialogue systems and non-dialogue forms of interactive recommendation.It also excluded systems focused only on initial rating acquisition for cold-start users.

3 INTERACTION MODALITIES OF CRS

CRS use varied input, output, device, and initiative designs, including forms, natural language, hybrid interfaces, and multimodal environments. The survey highlights unresolved usability and design trade-offs, especially for voice interaction and modality selection.

  • 3.1 Input and Output Modalities: CRS commonly use forms, natural language, or hybrid combinations for input and output.Natural language may be written or spoken, while hybrid systems can combine language with visual lists or critiquing controls.
  • 3.1 Input and Output Modalities: Voice-only systems offer flexible dialogue but struggle with utterance understanding, intent identification, and presenting multiple recommendations.These challenges are especially relevant on smart speakers such as Alexa or Google Home.
  • 3.2 Application Environment: CRS appear as stand-alone recommenders, embedded chatbots, multimodal interfaces, and applications on voice-based home assistants.In embedded settings, recommendation may be one function within a larger software solution.
  • 3.2 Application Environment: Device capabilities strongly influence CRS design choices, with smart speakers often restricting interaction to voice.Other investigated environments include interactive walls and systems interpreting facial expressions and gestures.
  • 3.3 Interaction Initiative: Most CRS are mixed-initiative systems, balancing user freedom with system guidance to obtain reliable preference information.Fully NLP-based interfaces provide more user control, but the system still typically maintains an agenda for advancing the dialogue.
  • 3.4 Discussion: Mixed natural-language and button interfaces can ease disambiguation, while optimal modality combinations remain difficult because user preferences vary.The survey also notes unresolved questions about dialogue flexibility, system guidance, and voice presentation of recommendation sets.

4 UNDERLYING KNOWLEDGE AND DATA

CRS require item information and recommendation-generation resources, along with knowledge about user intents, dialogue states, and related conversational context.

  • 4 UNDERLYING KNOWLEDGE AND DATA: CRS depend on item information, recommendation rules or constraints, machine-learning models, and knowledge about dialogue behavior.Additional background knowledge can include supported user intents and possible dialogue states.

4.1 User Intents

User intents define what CRS can recognize and support during recommendation dialogues, but the literature provides limited systematic coverage. The survey organizes common domain-independent intents and identifies gaps in existing systems.

  • 4.1 User Intents: Supported user intents vary by application domain, while some intents recur across many CRS.The survey’s overview is intended to help designers identify unsupported user needs.
  • 4.1 User Intents: Only 11 papers explicitly discuss relevant user intents, and few cover most domain-independent intents.Other studies address smaller intent subsets or application-specific group recommendation intents.
  • 4.1 User Intents: Chit-chat represented nearly 80% of recorded user utterances in one study and was associated with reduced dissatisfaction.Although irrelevant to the interaction goal, chit-chat may contribute to an engaging user experience.
  • 4.1 User Intents: Preference elicitation includes stating criteria, answering system questions, revising preferences, and responding to recommendations.Users may add or relax constraints, inspect profiles, reject recommendations, or restate preferences.
  • 4.1 User Intents: 41.1% of users in one voice-controlled movie-recommender study attempted to refine initially stated preferences.In another study, rejecting or restating preferences occurred in 1.5% of interactions.
  • 4.1 User Intents: CRS also support recommendation requests, requests for more options, similarity comparisons, and explanations about recommended items.These intents can occur at the dialogue start or after users revise their preferences.

4.2 User Modeling

CRS model user preferences through session-specific profiles, long-term preference profiles, item-level feedback, item facets, and collective preference information. These representations support recommendation and adaptation across interaction contexts.

  • 4.2 User Modeling: CRS interactively elicit current preferences, which are recorded in user profiles for subsequent relevance inference.Preference information may represent item-level expressions or estimates and preferences for item facets.
  • 4.2 User Modeling: Preference elicitation can use ratings, likes, dislikes, implicit feedback, item attributes, or desired functionalities.These alternatives represent different granularities of user preference information.
  • 4.2 User Modeling: Some systems elicit preferences without engineered item features by using keywords, tags, key phrases, or latent representations.Users repeatedly specify preferences on items, and the system seeks similar items using unstructured features.
  • 4.2 User Modeling: Some CRS maintain long-term profiles that derive supposedly stable preferences across multiple sessions.Examples include preferences for non-smoking restaurant rooms and probabilistic content-based models.
  • 4.2 User Modeling: Collective user preferences and popular items can support cold-start recommendation when little is known about an individual user.Feedback on popular items can then refine the user model.

4.3 Dialogue States

CRS dialogue states organize multi-turn recommendation conversations by defining possible steps and transitions. Approaches range from manually specified state machines to lightweight phases, implicit intents, and neural models.

  • Dialogue-state representations: Finite-state dialogue management defines possible states and allowed transitions, with the next action selected from the current dialogue state.This approach is especially common in knowledge-based CRS.
  • Pre-defined dialogue models: Pre-defined dialogue models combine preference-elicitation steps with states for presenting results, explaining recommendations, and comparing alternatives.Design-time transitions constrain the model, while run-time decision rules determine the path taken.
  • Learned transitions: Some systems learn run-time transitions with reinforcement learning rather than manually engineered decision rules.One stated goal is minimizing the number of interaction steps.
  • Dialogue-state representations: State machines can be represented visually or through textual and declarative formalisms such as dialogue grammars and case-frames.Commercial tools can model linear and non-linear conversation flows.
  • Alternative state designs: Critiquing systems often use only a few generic states, while NLP systems may represent states through adaptive phases or implemented user intents.This makes dialogue-state management relatively lightweight in some systems but implicit in others.
  • End-to-end learning: An end-to-end neural CRS can encode a relatively simple dialogue model centered on movie-related questions, sentiment analysis, and subsequent recommendations.The described system seemingly supports fewer intents and information requests beyond movie names.

4.4 Background Knowledge

CRS rely on diverse background knowledge beyond user-item ratings, including item information, dialogue corpora, interaction logs, user studies, and external lexical resources. These resources support recommendation, preference understanding, dialogue modeling, and entity recognition.

  • Item-related information: CRS use item databases containing ratings, metadata, tags, and extracted keyphrases to support recommendation and explanation generation.Item attributes can also help determine which information to acquire from users.
  • Item-related information: Surveyed systems use standard rating datasets, domain-specific datasets, and researcher-created collections for item-related information.Critiquing applications often rely on datasets created for individual studies.
  • Item-related information: Limited dataset reuse is common because most analyzed papers did not publicly share their datasets.This observation concerns the surveyed research literature.
  • Dialogue corpora: NLP-based CRS commonly train on recorded and sometimes annotated human conversations obtained from crowdworkers, interviews, chatbot logs, or other existing datasets.These dialogue corpora provide interaction histories for building conversational systems.
  • Dialogue corpora: Dialogue corpora may be combined with other knowledge bases when conversations lack sufficient relevant information.One example combines dialogue data with MovieLens for sentiment analysis and rating prediction.
  • Logged interactions and user studies: Interaction logs and user studies help researchers analyze dialogue quality, strategies, feedback types, and user needs.Such datasets can also be annotated for model training and system development.
  • Lexicons and world knowledge: External resources such as Wikipedia, Wikitravel, and WordNet support entity-keyword mapping and semantic-distance calculations.These resources provide dictionaries, lexicons, and world knowledge for NLP-based CRS.

4.5 Discussion

CRS are knowledge- and data-intensive because conversational recommendation requires background information beyond a user-item rating matrix. The survey highlights constrained interaction coverage and substantial manual effort as continuing challenges.

  • Knowledge and data requirements: CRS often require item, domain, dialogue, and background knowledge in addition to user-item ratings, particularly for dialogue management.This distinguishes them from traditional relevance prediction for unseen items.
  • Pre-defined knowledge versus learning: Forms-based systems predefine dialogue states, supported intents, and user attributes, whereas NLP systems usually allow more dynamic flows using additional resources.These resources include dialogue corpora, answer templates, lexicons, and world knowledge.
  • Pre-defined knowledge versus learning: Pure end-to-end learning from recorded dialogues remains challenging because supported interaction patterns are usually implicitly or explicitly predefined.The common pattern is that the user provides preferences and the system recommends.
  • Pre-defined knowledge versus learning: Predefined or instruction-shaped dialogue data can narrow the range of utterances that a CRS supports.The cited example cannot handle a request for a good sci-fi movie.
  • Intent engineering and dialogue states: Defining supported user intents is often a central manual development task, and richer applications may require domain-specific or group-decision intents.Examples include style advice, collaborator invitations, group recommendations, preference-conflict resolution, and voting.
  • Intent engineering and dialogue states: The supported intent set determines how varied conversations can be, while missing appropriate responses can damage perceived system quality.Recommendation explanations are described as useful for decision-making and user trust.
  • Intent engineering and dialogue states: Anticipating or learning user intents can require substantial manual effort, including professional writing to achieve natural and rich conversations.Rule-based intent modeling remains a challenge for conversational systems.

5 COMPUTATIONAL TASKS

Conversational recommender systems perform recommendation-centered actions such as requesting preferences, recommending items, explaining suggestions, and responding to user utterances. The survey finds broad technical diversity across these tasks, but limited CRS-specific research and support for explanations.

  • 5.1 Computational Tasks: CRS carry out four general system actions: Request, Recommend, Explain, and Respond, although individual systems may implement only some of them.System-driven systems commonly request attribute preferences and process feedback, while user-driven systems mainly respond to user conversational acts.
  • 5.1.1 Request: Slot-filling CRS determine which item facet to ask about next, using attribute weights, entropy-based methods, or learned policies.Recent reinforcement-learning approaches can choose between requesting a predefined facet and presenting a recommendation from the current dialogue state.
  • 5.1.1 Request: Preference elicitation can also use feedback on individual items, item pairs, or item sets, with bandit methods selecting items or pairs for feedback.Feedback may express absolute preferences, such as like or dislike, or relative preferences between two items.
  • 5.1.2 Recommend: Recommendation is the core CRS task, supported by collaborative, content-based, knowledge-based, and hybrid approaches that mainly use short-term preferences.Some systems additionally use long-term preferences to speed preference elicitation, while others apply constraint-based filtering, machine learning, or visual-feedback methods.
  • 5.1.3 Explain: Few CRS-specific studies address explaining, and only a smaller set of proposed CRS support explanation functionality.The survey also reports few CRS-specific studies of intent detection and named entity recognition, possibly reflecting limited taxonomies and annotated recommendation-dialogue data.
  • 5.3 Discussion: Dialogue management is often implemented with static predefined transitions, implicitly through intent mapping, or through preference-acquisition strategies such as slot-filling.In some systems, dialogue states are limited to choosing between asking questions and providing recommendations.

6 EVALUATION OF CONVERSATIONAL RECOMMENDERS

The survey finds that CRS evaluations span task outcomes, efficiency, conversational quality, usability, and subtasks, using varied offline, user-study, and combined methods. Despite this breadth, field evaluations and standardized comparisons remain limited.

  • Evaluation dimensions: CRS evaluation emphasizes human-computer interaction and measures both recommendation outcomes and the efficiency or quality of the conversation.This differs from algorithm-oriented evaluation, which more often centers on task fulfillment alone.
  • Evaluation dimensions: CRS evaluations examine task effectiveness, task efficiency, conversation quality and usability, and the effectiveness of subtasks such as intent recognition.Task effectiveness may use accuracy, acceptance, rejection, satisfaction, or perceived recommendation quality; efficiency often uses interaction steps.
  • Methods and materials: Studies use offline experiments, user studies, or combinations of both, while reports on deployed systems and A/B tests are rare.Some evaluations are absent, qualitative, or anecdotal, and field-test reporting often provides limited detail.
  • Methods and materials: CRS experiments use prototype applications, item databases, logged human conversations, and explicit dialogue-related knowledge such as supported intents.The materials depend on the technical approach used by each system.
  • Measurement approaches: Offline evaluations commonly use withheld preferences, accuracy metrics, real or simulated user profiles, and ground-truth datasets to assess recommendation or component effectiveness.Examples include Average Precision, Hit Rate, RMSE, and evaluation after successive question-answering rounds.
  • Measurement approaches: Pure natural-language interfaces produced less efficient recommendation sessions in one comparison, partly because natural-language utterances were interpreted incorrectly.The study compared NLP-based, button-based, and mixed chatbot interaction modes using questions, interaction time, and time per question.
  • Discussion: A review of CRS research finds diverse methodologies and metrics, limited adoption of user-centric frameworks, no established standards, and difficult comparisons across systems.Systems vary in application domain, interaction strategy, and background knowledge.
  • Discussion: BLEU scores alone may poorly reflect users’ perceptions of generated utterances, so subjective evaluations should supplement automatic assessment.The survey notes that task-oriented CRS can be especially difficult to evaluate.

7 OUTLOOK

The outlook identifies open questions about interaction modalities, non-standard environments, theories of conversation, and end-to-end learning. It emphasizes that CRS usefulness, user expectations, explanations, adoption, and evaluation remain insufficiently understood.

  • Overview: Recent CRS approaches increasingly use machine learning, especially deep learning, and natural-language interaction, but several research questions remain open.The survey organizes its future outlook around four general research directions.
  • Interaction modalities: Research must clarify which interaction modality best supports users for particular tasks and situations, including whether alternatives should be offered.The outlook also highlights interpretation of non-verbal communication and limitations of entirely voice-based CRS.
  • Application environments: Most research focuses on web or mobile applications, leaving limited knowledge about CRS requirements in physical stores, cars, kiosks, and robots.These settings represent non-standard application environments discussed in the survey.
  • Theories of conversation: Few CRS studies draw on conversation theories, leaving unclear what makes systems helpful, what users expect, which intents to support, and why systems fail.The outlook also calls for more research on explanations, trust, intimacy, adoption, and adapting communication style to users.
  • End-to-end learning: The usefulness of pure end-to-end learning approaches remains an open technical and methodological question, closely tied to how such systems are evaluated.These approaches use an item database and a corpus of past conversations as inputs, while recent NLP advances have not resolved the question of practical usefulness.
Loading 2004.00646v2…