Source-linked AI summary
Large Language Model-Driven Context-Aware Eco-Feedback Generation and Evaluation
Wooyoung Jung, Prosper Babon-Ayeng
TL;DR
Existing eco-feedback often relies on energy-use patterns without adequately representing household characteristics, while LLM-based residential eco-feedback generation has not been systematically studied. This paper develops a contextual-engineering framework using structured energy analysis, utility information, household characteristics, and SC-CoT prompting, and evaluates its accuracy and adaptability. The framework aligned with reference solutions and changed targeted appliances and strategies across household contexts, while the study’s scope is limited by synthesized characteristics and summer data from Austin, Texas.
Problem
Existing eco-feedback often relies on energy-use patterns and insufficiently reflects household characteristics, while LLM-based residential eco-feedback generation remains underexplored.
Method
The framework combines structured household energy analysis, utility rate structures, household characteristics, and self-consistency with chain-of-thought prompting to generate eco-feedback.
Results
The framework generated accurate, data-driven, and context-sensitive eco-feedback, aligning with reference interventions and shifting targeted appliances and strategies across household contexts.
Takeaways & Limitations
Contextual engineering with an LLM can support more adaptive and feasible household energy guidance by combining quantitative energy analysis with qualitative household context.
Takeaways & Limitations
The study uses synthesized household characteristics and summer energy data from Austin, Texas, limiting representation of real household diversity and other seasons or climates.
Abstract
from arXiv · showhide
The objective of this study was to demonstrate the potential of generating eco-feedback that accounted for unique household contextual information, named as context-aware eco-feedback, through a large language model-integrated framework. Previous studies have introduced personalized eco-feedback, mostly relying on household energy use patterns; however, they frequently did not reflect distinct household characteristics, including their persona or non-negotiable routines, leaving eco-feedback ineffective and sometimes superficial. To address these limitations, we introduced a contextual engineering framework that generated eco-feedback using a self-consistency with chain-of-thought prompt that leveraged household energy analysis data, utility rate structures, and characteristic information. We conducted a rigorous empirical validation and a combinatorial evaluation analysis to assess this framework systematically. The former aimed to test the framework's ability to generate accurate and data-driven eco-feedback, customized to given contexts by comparing it with reference interventions. The latter aimed to reveal the framework's adaptability across diverse household contexts by investigating how context-aware eco-feedback changed. Key findings were the following: our proposed framework generated eco-feedback that aligned with reference solutions at a mean accuracy of 92.0% across different household configurations, accurately leveraging the provided household data for feedback generation (95.7% of data citation accuracy). Also, it was largely adaptive to diverse household contexts, significantly shifting targeted appliances and energy-saving strategies. Ultimately, this study contributes to realizing the next level of context-aware interactions between occupants and buildings which paves the way for higher occupant living quality and sustainability.
1. Introduction
Households have complex, diverse energy needs, while existing eco-feedback often provides limited personalization and insufficiently feasible guidance. The study proposes LLM-generated context-aware eco-feedback that combines energy patterns with household characteristics and other contextual data.
- Motivation: Homes support diverse activities and increasingly incorporate distributed energy resources and energy-intensive appliances, creating complex household energy-use patterns.Utilities are responding with demand-side strategies including time-of-use pricing and smart thermostat incentives.
- Limitations of existing feedback: Existing smart-meter profiles and monthly bills provide limited information and often lack personalized suggestions about when, where, and how energy was used.Smart-meter feedback may show accumulated daily or weekly usage, whereas bills generally lack detailed temporal and appliance-level information.
- Limitations of existing feedback: Personalized eco-feedback has mainly relied on charts, graphs, pre-framed texts, and energy-use patterns, limiting its depth and adaptability.Effective household guidance must account for differing preferences, constraints, and circumstances.
- Proposed direction: The study examines context-aware eco-feedback that incorporates household energy-use patterns and contextual characteristics through an LLM-integrated framework.The framework is intended to reflect distinct household characteristics beyond energy-use patterns alone.
2. Background
Prior eco-feedback and LLM research establishes the need for more adaptive residential energy guidance, but systematic LLM-based eco-feedback generation remains underexplored. The paper frames contextual engineering as a bridge between structured household information and household-specific recommendations.
- Eco-feedback evolution: Eco-feedback has evolved from basic consumption reporting toward metrics, valence, background information, and real-time sensing through connected devices.Research has examined energy consumption, environmental impact, costs, weather, energy sources, utility pricing, and timing-related feedback.
- Eco-feedback limitations: A longstanding limitation is insufficient adaptivity to household priorities, non-negotiable routines, and feasible interventions.Personalized feedback partially addresses this issue but remains dependent on predefined features and energy-use patterns.
- LLMs in BEM: LLM research in building energy management has explored load prediction, fault diagnosis, anomaly detection, data mining, assistants, and smart-building control.These studies highlighted the role of prompt design and LLM integration in addressing BEM tasks.
- Research gap: Prior research had not systematically investigated LLMs for generating eco-feedback for residential units.This study positions LLMs as agents capable of incorporating household characteristics into eco-feedback.
- Contextual engineering: Contextual engineering systematically organizes relevant information, specifies reasoning scope, and tailors prompts to optimize LLM performance.It structures household energy data and contextual descriptions so models can produce context-aware recommendations.
3. Methodology
The methodology combines structured household energy analysis, contextual household information, and utility conditions in a prompt-based framework, then validates generated feedback against reference interventions. It uses appliance-level data, synthesized household contexts, and metrics for identifying curtailment and shifting opportunities.
- Framework: The conceptual framework transforms energy use, generation, household characteristics, utility rates, and building information into structured prompts for LLM-based eco-feedback generation.The framework supports automated generation of context-aware feedback within building energy management systems.
- Validation design: The empirical validation compares generated context-aware eco-feedback with reference interventions using data utilization, error detection, and consistency metrics.The analysis tests whether household context is translated into feasible, high-impact recommendations consistently.
- Household data: Validation uses three sets of actual appliance-specific Austin energy data at 15-minute intervals, two utility rates, and synthesized personas representing household preferences or constraints.Homes were selected from a 50-household pool based on appliance count and solar-panel adoption.
- Energy analysis: Five metrics assess appliance intervention potential: total energy and mean power use, usage frequency, usage variability, solar power alignment, and solar power coverage.Standby energy is excluded, while solar coverage follows an aggregated-meter net-metering assumption.
- Intervention design: The reasoning process evaluates home type, appliance demand, usage patterns, solar conditions, curtailment or shifting opportunities, and contextual feasibility before finalizing strategies.Behavioral interventions consider preferences, routines, constraints, efficiency, substitution, and load timing.
- Reference construction: Reference interventions were established through systematic reasoning and external-expert adjudication, then used to evaluate the generated feedback.The process considered both load curtailment and load shifting, including simultaneous frequency and variability assessment under TOU rates.
3.3. Combinatorial Analysis
The combinatorial analysis tested framework adaptability across household energy data, utility rates, and synthesized personas, producing diverse scenarios and evaluating recommendation characteristics and persona fidelity.
- Scenario design: The analysis used 50 actual household energy-usage datasets and two utility rate structures, including standard flat and time-of-use rates.Four personas were developed from the literature to represent diverse household characteristics.
- Scenario design: 400 scenarios combined 50 households, two utility rate structures, and four personas, with three generated responses per scenario for 1,200 total responses.Each scenario represented a unique combination of appliance-level usage, utility conditions, solar availability, and contextual characteristics.
- Analysis procedure: Descriptive and inferential analyses compared intervention characteristics across persona types, utility rates, and their combinations.Means and standard deviations were calculated for strategy percentages, targeted appliances, and persona fidelity scores, with statistical significance assessed for observed differences.
- Evaluation measures: Responses were examined for recommended appliances, shifting, curtailment, and efficiency intervention types, plus persona fidelity.Efficiency interventions targeted appliance settings or upgrades rather than reduction or scheduling changes.
- Evaluation measures: Persona fidelity was scored against persona-specific criteria using GPT-4o-mini ratings of 0.0, 0.5, or 1.0, followed by manual validation.The criteria assessed whether generated interventions reflected behavioral characteristics defined for each assigned persona.
4. Results
The framework produced data-grounded interventions that closely matched reference solutions and adapted targeted appliances, strategies, and contextual framing to utility rates and personas. Results also showed strong but uneven citation of household data and distinct persona-conditioned behavioral signatures.
- Empirical validation: 92.0% overall reference alignment showed that generated interventions closely matched reference energy-saving solutions.Reference alignment rates ranged from 89.8–95.9%.
- Empirical validation: 96.6% overall appliance validity indicated that most interventions targeted appliances present in household data rather than hallucinated appliances.Appliance validity rates ranged from 94.3% to 100%, although some appliance associations were extrapolated beyond the provided data.
- Strategy adaptation: 77.1% curtailment dominated Home #1, whereas load shifting dominated Homes #2 and #3 at 67.6% and 77.6%, respectively.The framework varied strategies with utility rates, solar availability, and household contextual characteristics.
- Data grounding: 90.8–98.7% of appliance energy-use and usage-pattern citations were made, with 90.2–100% citation accuracy.Citation frequency and accuracy were strongest for appliance-level energy data.
- Data grounding: Utility-rate citation frequency ranged from 18.8% under flat rates to 84.5% under time-of-use rates, while cited-rate accuracy remained 91.7–100%.Solar data, available only for Home #3, was cited in 75.6% of interventions with 87.3% accuracy; household-context citation frequency ranged from 14.7–68.9%.
- Appliance targeting: Under time-of-use rates, refrigerators were targeted 32 times versus 172 times under flat rates, while washing machines and dishwashers were targeted more often because they were shiftable.The refrigerator difference was significant, with a chi-square statistic of 87.1 accounting for 85.3% of the total effect.
- Persona effects: Persona did not change which appliances were targeted, but it changed recommendation framing and emphasis, such as prioritizing cost savings for cost-minimizers.Strategy distributions nevertheless varied across personas in their secondary preferences.
- Persona effects: Persona self-alignment scores were 98.2% for Tech adopter, 92.8% for comfort maximizer, 88.8% for cost minimizer, and 79.4% for routine constrainer.The matrix showed differentiated behavioral signatures and asymmetric overlap between personas.
5. Discussion and Limitation
The framework combines LLM-based synthesis with externalized energy analysis to produce context-aware, feasible eco-feedback. The discussion identifies building information and synthesized household characteristics as important limitations affecting contextual grounding and generalizability.
- 5. Discussion and Limitation: The LLM translated quantitative energy metrics and qualitative contextual information into coherent, human-readable, context-aware, and feasible eco-feedback.It functioned as a reasoning and synthesis engine rather than a simple text generator.
- 5. Discussion and Limitation: Externalizing energy analysis reduced the LLM’s analytical burden while supporting rigor, consistency, and interpretability.The energy analysis component extracted quantitative insights from household energy data.
- 5. Discussion and Limitation: Building attributes could improve contextual grounding by distinguishing energy use driven by structural constraints from occupant behavior.Relevant attributes include floor area, construction vintage, envelope characteristics, insulation, windows, HVAC type, and thermal performance.
- 5. Discussion and Limitation: Synthesized household characteristics may not fully capture the complexity and diversity of real households.The study grounded these characteristics in appliance usage patterns and literature to approximate realistic contexts.
- 5. Discussion and Limitation: Summer-only energy data from Austin, Texas may limit generalizability to other seasons or climate conditions, especially heating-driven settings.The data covered June to August.
6. Conclusion
The study validated a context engineering-based framework for generating context-aware eco-feedback using systematic empirical validation and combinatorial analyses. It concludes that the framework supports adaptive household energy guidance while identifying future work on delivery and interaction.
- 6. Conclusion: The framework was validated for accuracy, data utilization, and adaptivity through systematic empirical validation and combinatorial analyses.The analyses assessed generated feedback across household energy and contextual characteristics.
- 6. Conclusion: Generated feedback identified high-impact appliances, proposed contextually reasonable interventions, and reflected household-specific characteristics.The evaluation integrated quantitative appliance-level energy analysis with qualitative personas and non-negotiable routines.
- 6. Conclusion: Future controlled experiments will examine how feedback modality, timing, and presentation affect engagement, comprehension, and sustained energy-saving behaviors.The planned evaluation focuses on different context-aware feedback delivery types.
- 6. Conclusion: The current implementation did not explicitly model bidirectional interaction, adaptive dialogue, or real-time user feedback.Planned agentic AI work includes continuous adaptation, human-in-the-loop learning, and automated controls.
A. Framework Components
The framework integrates household data, external energy data analysis, tailored prompts, and one or more LLMs. These components connect quantitative energy insights with prompts and model roles designed around household needs and constraints.
- Household Data: Household data provide contextual signals, including utility rate structures and high-granularity appliance- or circuit-level energy usage.Rate structures help determine whether strategies such as load shifting are financially beneficial.
- Energy Data Analysis: The energy data analysis component handles computationally intensive analysis outside the LLM.This separation reduces the data and token burden assigned to the model.
- Prompt: Tailored prompts guide the LLM toward more adoptable and contextually aligned eco-feedback.Prompts can emphasize affordability, comfort preservation, or simple explanations according to household characteristics.
- LLM: One or multiple LLMs can interpret contextual inputs, generate eco-feedback, verify reasoning, refine outputs, or ensure alignment with household constraints.Multiple models can provide complementary strengths such as reasoning and reliability.
B. Household Appliance-Level Energy Use Data
The dataset covers 50 households with diverse energy consumption, technology adoption, and appliance ownership profiles. Its statistics show substantial variation in technologies, monitoring duration, total consumption, and monitored appliance counts.
- Dataset Overview: 50 households comprise the appliance-level energy-use dataset, representing diverse consumption patterns, technology adoption, and appliance ownership configurations.The dataset captures a wide spectrum of residential energy behaviors.
- Technology Adoption: 58% of households had solar panels and 56% owned EV chargers.These percentages describe technology adoption across the dataset.
- Energy Consumption Profile: Total energy consumption ranged from 206.73 kWh to 8,959.14 kWh over the monitoring period.21 homes had three months of data and 29 had one month; mean consumption was 1,449.83 kWh and median consumption was 799.11 kWh.
- Energy Consumption Profile: Households with both solar PV and EVs had the highest mean consumption at 2,112.97 kWh.This is reported as a technology-adoption-category statistic.
- Appliance Distribution and Characteristics: Monitored appliances per household ranged from 2 to 21, with a mean of 9.3 and median of 10.The standard deviation was 5.1 appliances.
C. Energy Use Analysis Equations
The framework quantified appliance use through binary activity indicators and aggregated these data into frequency, variability, power, and solar-alignment metrics. These measures supported analysis of usage timing, consistency, energy demand, and compatibility with solar availability.
- Usage frequency: Appliance activity was determined at each 15-minute interval using a binary indicator with heuristically selected thresholds separating active and standby modes.The hourly use frequency then summed indicator values across the four intervals within each hour.
- Usage frequency: Hourly use frequency was aggregated across four 15-minute intervals and normalized as an average over D days.The timeframe-specific average use frequency was then computed from these interval-based measures.
- Usage variability: Usage variability was calculated from an appliance’s hourly frequency using its overall mean hourly frequency and standard deviation.The equations use ¯Fi for the overall mean hourly frequency and σi for its standard deviation.
- Power and energy: Mean power was computed as a separate metric to characterize appliance energy demand.The metric set also included total energy use and mean power use.
- Solar alignment: Solar power alignment measured how closely appliance use coincided with periods of solar power availability.Solar power coverage additionally quantified the proportion of an appliance’s total demand that solar power could supply, with PV denoting solar panels and Ptot total demand.
D. Household-Specific Reference Eco-Feedback
The reference eco-feedback translated household-specific energy patterns, rate structures, solar availability, and household conditions into appliance-targeted curtailment or load-shifting strategies. Recommendations differed across flat-rate, TOU, and solar-generating homes while excluding appliances below the stated consumption threshold.
- Home #1: Home #1 used a flat rate without on-site solar, so reference interventions focused on load curtailment despite the household persona.HVAC was the primary target, followed by water heating, cooking appliances, and dishwasher-use adjustments.
- Home #2: Home #2 used a TOU rate, creating incentives for both load curtailment and load shifting without behavioral or comfort constraints.The analysis identified five appliances with meaningful energy contributions for targeted interventions.
- Home #2: Home #2 recommendations targeted HVAC, water heating, clothes drying, and dishwashing according to peak-hour demand and appliance-specific operating patterns.Strategies included widening temperature setpoints, shifting water heating and dishwasher operation off-peak, and air-drying or limiting dryer use.
- Scope boundary: Each reference table excluded appliances whose total energy consumption was below 0.5%.This exclusion note accompanied the appliance-level analyses for the household cases.
- Home #3: Home #3 combined a flat rate with solar generation, shifting interventions toward periods of on-site solar availability to maximize solar self-consumption.The household-specific targets included HVAC, a variable pool pump, an EV charger, and an electric water heater.