Source-linked AI summary
Investigating the Ability of Large Language Models to Analyze Recipes for Diabetes
Revathy Venkataramanan, Aditya Luthra, Venkatesan Nadimuthu, Amit Sheth
TL;DR
The paper asks whether LLMs can reliably determine whether recipes are suitable for diabetes, a task requiring guideline retrieval, dietary concept understanding, and recipe-level analysis. It evaluates three levels of prompting on a balanced 7,607-recipe benchmark and finds that guideline-based reasoning improves performance, while models tend to classify recipes as unsuitable. Mistral-7B and Llama3.1-70B show the strongest overall performance among the evaluated models.
Problem
Determining recipe suitability for diabetes requires comprehensive dietary-guideline retrieval, conceptual understanding, and analysis of ingredients and cooking methods.
Method
The study evaluates four classes of LLMs with Direct Query, Context-Guided, and Exemplary Context prompts on 7,607 recipes, including 3,807 suitable and 3,800 unsuitable recipes.
Results
Models that reason using dietary-guideline keywords perform better, with Mistral-7B and Llama3.1-70B showing superior and stable performance across prompts.
Takeaways & Limitations
Most models are cautious and tend to predict recipes as unsuitable, reflecting concern about false positives when harmful recipes are classified as safe.
Takeaways & Limitations
The approach requires expert validation, and unequal representation of dietary guidelines means models may reason from only a few categories despite many recommendation and avoidance categories.
Abstract
from arXiv · showhide
Several studies have evaluated the ability of Large Language Models (LLMs) for meal planning, yielding positive outcomes. These models can process natural language inputs and leverage learned knowledge from their pretraining to generate meal plans. In this work, we investigate the ability of LLMs to analyze the suitability of given recipes for diabetes. The primary challenge for LLMs is to retrieve relevant dietary guidelines for diabetes, decompose recipes into ingredients and cooking methods, and apply these guidelines to determine the recipe's suitability. To study these challenges, we employ three kinds of prompts namely, (i) Direct Query Prompt (ii) Context-Guided Prompt, and (iii) Exemplary Context Prompt that incorporate different levels of diabetes dietary guidelines from medical sources. We introduce a benchmark dataset curated for this investigation consisting of 7607 recipes that include 3807 recipes suitable for diabetes and 3800 recipes not suitable for diabetes. Our results demonstrate that most LLMs are cautious in predicting recipes as suitable to prevent detrimental outcomes. Further, the models that can reason using the dietary guidelines performed better in predicting the suitability of recipes for diabetes. Overall, Mistral-7B and Llama 70B showed superior performance to their counterparts.
1 Introduction
This study examines whether LLMs can assess recipe suitability for diabetes, a task requiring comprehensive dietary knowledge rather than cautious meal recommendation. It focuses on retrieving guidelines, understanding dietary concepts, and applying them to recipe ingredients and methods.
- Motivation: Dietary guidelines for diabetes are extensive, difficult to remember and apply, and do not exhaustively specify which foods should be avoided.For example, some fruits provide healthy carbohydrates and fiber, while glucose-rich fruits may be unsuitable.
- Related context: LLMs have shown positive results for meal planning, but studies also report limitations in diabetes nutrition management and recommend clinician augmentation.
- Research problem: Assessing recipe suitability requires comprehensive dietary-guideline knowledge rather than retrieving only a subset of guidelines used for cautious meal recommendations.
- Research problem: The task challenges LLMs to retrieve relevant medical knowledge, understand diabetes dietary concepts, and apply general rules to specific recipes.The study frames these as medical knowledge retrieval, conceptual understanding, and deductive analysis challenges.
- Study design: The evaluation uses Direct Query, Context-Guided, and Exemplary Context prompts, alongside four LLM classes and a 7,607-recipe balanced dataset.The dataset contains 3,807 diabetes-suitable recipes and 3,800 recipes judged unsuitable.
2 Related Work
Prior work has applied LLMs to food recommendation, personalized nutrition, and diabetes meal planning. These systems combine retrieval, domain-specific reasoning, causal methods, knowledge graphs, or multimodal inputs to support dietary recommendations.
- Research context: The related work positions LLMs as tools for generating personalized diets aligned with nutritional guidelines for conditions such as diabetes.
- Food recommendation: Recent food recommendation systems use LLMs with retrieval, food-specific reasoning, causal reasoning, knowledge graphs, and dietitian validation.
- Diabetes nutrition: LLM-based health agents and tools have been developed to integrate diabetes guidelines, personalize recommendations, and optimize glycemic control.
3 Approach
The study evaluates recipe analysis with a medically grounded, balanced dataset and three prompts that provide progressively richer dietary guidance. Models process recipes individually and, for contextual prompts, explain decisions using guideline concepts.
- Dataset creation: 3,807 diabetes-specific recipes were collected from MayoClinic, Diabetes UK, and Diabetes Hub to form the suitable class.
- Dataset creation: 3,800 unsuitable recipes were randomly selected from Recipe1M after keyword searches in recipe titles and ingredients, preserving class balance.
- Models: The evaluation compares Mistral, Llama, Gemma2, and ChatGPT models spanning small, medium, and large sizes.
- Prompt design: Three prompts test direct internal guideline retrieval, explicit context-guided reasoning, and exemplary context containing direct and indirect dietary examples.
- Prompt design: The direct query asks for a YES/NO suitability judgment from the recipe title, ingredients, and instructions without explicitly supplying dietary guidelines.
- Prompt design: Context-guided prompts provide recommended and avoided dietary categories, while exemplary prompts add examples such as vegetables, butter, beef, and sausages.
- Evaluation procedure: Each model received one recipe at a time for each prompt, and outputs were stored with explanations; contextual prompts specifically required reasoning with guideline concepts.
4 Result and Discussion
Across prompts and models, performance reflected both cautious classification and the ability to reason with diabetes dietary guidelines. Mistral 7B and Llama3.1 70B were especially strong, while retrieval, conceptual coverage, and ingredient-level deduction remained challenging.
- Model Performance: Mistral 7B and Llama3.1 70B showed the most consistent or strongest performance across the three prompt types, while larger models did not necessarily perform better.Mistral 7B performed well with Direct Query Prompt, whereas Llama3.1 70B performed well with Context Guided and Exemplary Context Prompts.
- Model Performance: Most models produced lower recall than precision, indicating caution toward labeling recipes as suitable for diabetes.This behavior may reduce false-positive errors, where a harmful recipe is incorrectly identified as safe.
- Reasoning and Performance: Models using more diverse dietary-guideline keywords generally achieved better F1-score and stability, with Mistral 7B producing the highest amount of diverse keywords.Figure 3 links keyword counts with average F1-score and stability, while Mistral 7B averaged 1200 keywords for prompt-2 and 1008 for prompt-3.
- Medical Knowledge Retrieval: External dietary context generally made medical knowledge easier to access: prompt-1 produced 765 average keywords, compared with 836 for prompt-2 and 789 for prompt-3.ChatGPT was an exception because prompt-2’s external knowledge conflicted with its internal knowledge, producing high precision but low recall.
- Conceptual Understanding: Models handled straightforward concepts such as healthy carbohydrates and fiber-rich foods more confidently than categories involving fats, cholesterol, processed foods, or cooking methods.The authors emphasize that recipe evaluation requires assessing both ingredients and cooking methods against dietary guidelines.
- Deductive Analysis: Models could reason over dietary concepts to some extent, but the study could not establish whether they applied detailed reasoning separately to each ingredient and cooking method.The authors identify ingredient and cooking-method decomposition as a target for future work.
5 Conclusion
The study finds that Mistral 7B and Llama3.1-70B performed best overall, while most models conservatively classified recipes as unsuitable for diabetes. Performance improved with external dietary knowledge and reasoning, although models still showed gaps in guideline coverage and category use.
- Mistral 7B and Llama3.1-70B showed superior performance overall to their counterparts.
- Most models conservatively predicted recipes as suitable for diabetes because of the domain’s high-stakes nature.
- External knowledge in prompts improved performance, but conflicted with ChatGPT’s internal knowledge in some cases.
- Models understood healthy carbohydrates, good fat, added sugar, and refined carbohydrates, but needed stronger use of protein, non-fat, low-fat, deep-fried, and ultra-processed food categories.