Source-linked AI summary
GPT-3-driven pedagogical agents for training children's curious question-asking skills
Rania Abdelghani, Yen-Hsiang Wang, Xingdi Yuan, Tong Wang, Pauline Lucas, Hélène Sauzéon, Pierre-Yves Oudeyer
TL;DR
Manual generation makes curiosity-training content costly, so the paper tests whether GPT-3 can produce pedagogical cues through natural-language prompting. The generated content was relevant and pedagogically useful, with open cues supporting more efficient divergent question-asking than cues targeting predefined questions.
Problem
Manually producing pedagogical prompts for curiosity training is redundant, time-consuming, and costly, while alternative model-adaptation methods can require machine-learning expertise.
Method
The study uses GPT-3 prompt-based learning to generate linguistic and semantic cues for children’s divergent question-asking training, evaluating them against hand-generated content with expert annotations and field testing.
Results
Open GPT-3-generated cues elicited more efficient divergent question-asking than cues leading to one predefined question, while children asked significantly more syntactically correct divergent questions after training across all conditions.
Takeaways & Limitations
Prompt-based LLM content can support children’s divergent question-asking and may be usable in classroom and e-learning settings without specialist AI expertise.
Takeaways & Limitations
The training lasted only one short session, so engagement may have contributed to the positive results and its long-term impact remains to be investigated.
Abstract
from arXiv · showhide
In order to train children's ability to ask curiosity-driven questions, previous research has explored designing specific exercises relying on providing semantic and linguistic cues to help formulate such questions. But despite showing pedagogical efficiency, this method is still limited as it relies on generating the said cues by hand, which can be a very costly process. In this context, we propose to leverage advances in the natural language processing field (NLP) and investigate the efficiency of using a large language model (LLM) for automating the production of the pedagogical content of a curious question-asking (QA) training. We study generating the said content using the "prompt-based" method that consists of explaining the task to the LLM in natural text. We evaluate the output using human experts annotations and comparisons with hand-generated content. Results suggested indeed the relevance and usefulness of this content. We also conduct a field study in primary school (75 children aged 9-10), where we evaluate children's QA performance when having this training. We compare 3 types of content : 1) hand-generated content that proposes "closed" cues leading to predefined questions; 2) GPT-3-generated content that proposes the same type of cues; 3) GPT-3-generated content that proposes "open" cues leading to several possible questions. We see a similar QA performance between the two "closed" trainings (showing the scalability of the approach using GPT-3), and a better one for participants with the "open" training. These results suggest the efficiency of using LLMs to support children in generating more curious questions, using a natural language prompting approach that affords usability by teachers and other users not specialists of AI techniques. Furthermore, results also show that open-ended content may be more suitable for training curious question-asking skills.
1 Introduction
Curiosity supports learning and can be cultivated through classroom questioning, but manually producing pedagogical cues for conversational-agent training is costly. This study investigates GPT-3 prompting as a way to automate such content and evaluates it against hand-generated material.
- Curiosity supports learning and can be cultivated by encouraging children to question uncertainty.Research characterizes curiosity as a desire for new information and a malleable skill that can be elicited in classrooms.
- Conversational agents have used semantic and syntactic cues to improve children’s question-asking and domain-knowledge learning.These systems guide children toward constructing divergent questions.
- Manual production of large pedagogical-content datasets is redundant, time-consuming, and costly.
- The study tests whether GPT-3 prompting can partially automate divergent question-asking content while preserving pedagogical usefulness.It evaluates generated cues through expert annotations, comparisons with hand-generated content, and a field study of 75 children aged 9–10.
- The work contributes to applying LLMs to educational curiosity training.
2 Related work
Prior work links curiosity with divergent, information-seeking questions and shows benefits from conversational agents, while NLP and LLM advances offer ways to automate educational-content production. Prompt-based GPT-3 is especially accessible because it adapts tasks through natural-language instructions without fine-tuning.
- Curiosity-driven divergent questions require hypotheses, predictions, inferences, or connections that bring additional information.
- Conversational-agent training with hand-generated cues has improved children’s curiosity-driven behavior and subsequent domain-knowledge learning.
- NLP methods can automate educational tasks such as extracting key concepts and generating questions from learning resources.These applications may reduce teachers’ time spent creating assessment and evaluation materials.
- Fine-tuning and related adaptation methods can be costly, data-intensive, task-specific, or dependent on machine-learning expertise.
- GPT-3 prompt-based learning verbalizes a task in natural text without updating model parameters, enabling simple adaptation for educational users.The approach can also use zero-shot or one-shot examples.
3 Current study
The study compares GPT-3-generated and hand-generated cues for training children’s divergent question-asking. It evaluates cue quality with expert annotations and tests three cue conditions in a primary-school field study, followed by post-intervention assessments.
- GPT-3 cues are evaluated against hand-generated content for semantic relevance, divergence level, and syntactic quality.
- The study compares hand-generated closed cues, GPT-3-generated closed cues, and GPT-3-generated open cues in children’s divergent question-asking training.
- The field study includes 75 primary-school children aged 9–10 and compares their divergent question-asking performance across cue conditions.
- Post-intervention tests assess divergent-question learning progress, perceived competency, and attitudes toward curiosity and asking questions.The assessments follow a three-day intervention.
4 Study design
The study uses the Kids Ask platform to train divergent question-asking through conversational-agent cues, comparing hand-crafted and GPT-3-generated implementations. It examines closed cues leading to predefined questions versus open cues supporting multiple questions, while addressing usability and safety through prompting and human oversight.
- 4.2.1 Choice of the interaction workflow: Kids Ask combines a curiosity-elicitation workspace with a question-asking workspace during reading-comprehension activities.Children can reflect on confidence in a skippable general-knowledge quiz before selecting themed texts and receiving cues for divergent questions.
- 4.1 Experimental conditions: The study compares hand-crafted cues, GPT-3-generated closed cues, and GPT-3-generated open cues across three experimental conditions.Groups 1 and 2 receive a questioning word plus an answer to one possible divergent question, whereas Group 3 receives two related keywords supporting several possible questions.
- 4.1 Experimental conditions: Closed cues combine a questioning word with an answer to guide children toward one predefined divergent question about the text.For example, cues about vaccines and medicine are intended to elicit a question comparing them, even though the answer is not explicitly available in the text.
- 4.2.2 Choice of the cues: Open cues provide two related keywords and allow different question starters, enabling several divergent questions about the text.The design gives children more choice over the questions they formulate, consistent with the stated rationale concerning autonomy and curiosity.
- 4.2.3 Choice of the technology: GPT-3 is used with natural-language prompting to generate pedagogical cues without fine-tuning, supporting adaptation by teachers and other practitioners.The system uses an offline human-in-the-loop configuration to verify generated content, protect children’s data, and reduce risks from uncontrolled or biased outputs.
- 4.3 Ethical considerations: The study identifies reliance risks for both children and teachers, positioning LLMs as monitored complements to pedagogical intervention and human instruction.The authors recommend teacher monitoring, critical-thinking support, and avoiding replacement of teachers or other information sources.
5 Experimental procedure
The study used three one-hour sessions to collect baseline measures, train divergent question asking with agent cues, and assess post-intervention outcomes.
- Session 1: Children completed baseline profiling and pre-intervention measures before training, including curiosity, reading ability, and question-asking fluency assessments.
- Session 2: During training, children read or listened to three texts and generated six divergent-thinking questions per text using agent-provided cues.The session produced 18 questions per child without a separate time limit beyond the session length.
- Agent conditions: The hand-crafted and GPT-3 incentive agents used the same behavior, differing only in how their cues were generated.
- Agent conditions: The open condition used a different cueing approach from the two incentive agents during the question-generation interaction.
- Session 3: Children completed post-intervention motivation, workload, question-asking fluency, and curiosity-perception measures in the third session.
6 Data collection and measures
The study measured GPT-3 cue quality, children’s divergent question performance, participant profiles, learning experience, and pre-post changes in questioning and curiosity.
- Cue measures: Cue quality was assessed through human annotations of linguistic variety and complexity, semantic relatedness, and divergence relative to the source text.Linguistic cues were compared with hand-crafted cues, while semantic cues were evaluated for contextual relevance and divergence.
- Question measures: Children’s question quality was assessed with standardized syntactic and semantic annotation grids, with disagreements averaged between coders.
- Participants: The study recruited 75 fourth-grade students aged 9 to 10.5 and assigned them to hand-crafted incentive, GPT-3 incentive, or open GPT-3 groups.The groups contained 24, 26, and 25 children, respectively.
- Question measures: Divergent question performance counted correct questions and calculated the percentage classified as divergent, counting repeated questions once.A question was classified as divergent when its answer was not explicitly stated in the text.
- Learning experience and outcomes: The study also measured motivation, perceived workload, question-generation fluency, curiosity attitudes, and perceived competence in asking complex questions.Curiosity attitudes were assessed before and after training with the CIAC questionnaire.
7 Research goals and hypotheses
The study examined whether GPT-3 could generate effective cues for children’s divergent question asking and whether open cues improved outcomes relative to closed cues.
- Research goals: The study aimed to evaluate GPT-3-generated linguistic and semantic cues against hand-designed cues using expert annotations and children’s training outcomes.
- Hypotheses: The authors hypothesized that GPT-3 would generate correct and effective curiosity-eliciting cues comparable to hand-crafted cues.
- Hypotheses: They predicted similar divergent question-asking performance for the GPT-3 and hand-crafted incentive agents.
- Hypotheses: They predicted that the open GPT-3 agent would lead children to ask more divergent questions than either incentive agent.
- Hypotheses: Additional hypotheses concerned positive links between divergent question performance and curiosity traits, plus improvements in curiosity perception and divergent questioning after training.
8 Results
GPT-3 produced cues comparable to hand-crafted cues on several quality measures, while the open GPT-3 condition yielded stronger divergent-question performance and broader intervention gains.
- 8.1 Cue quality: 100% of GPT-3-generated content was rated non-offensive and relevant to the educational text.Human annotations used five-point Likert scales for offensiveness and contextual relevance.
- 8.1.2 Linguistic cues: Human and GPT-3-generated linguistic cues showed no significant differences in questioning-word variety or compound-word proportions.The reported tests found p-value=0.65 for variety and p-value=0.51 for compound-word proportions.
- 8.1.3 Semantic cues: Semantic relatedness did not differ significantly across the three conditions, and incentive-agent cue divergence also showed no significant difference.The relatedness comparison reported F(2,74)=0.41, p-value=0.66; the incentive divergence comparison reported t=0.9, p-value=0.37.
- 8.2 Cue use: Children used open-agent cues more often than cues from either incentive agent, with open use at M=91.78% versus M=77.28% and M=76.12%.The open condition differed significantly from the hand-crafted and GPT-3 incentive conditions.
- 8.2.1 Divergent QA skills: The open GPT-3 agent produced significantly more divergent questions than both the hand-crafted and GPT-3 incentive agents, while the two incentive groups did not differ.The overall condition effect was significant, F(2,72)=4.11, p-value=0.02; open versus hand-crafted yielded p-value=0.003, and open versus GPT-3 incentive yielded p=0.04.
- 8.4.1 Intervention effect: All groups improved in divergent question fluency after training, with significant effects of time and condition but no time-by-condition interaction.The repeated-measures ANOVA reported time F(2,72)=46.95, p-value<0.001, condition F(2,72)=3.9, p-value=0.002, and interaction p=0.11.
- 8.4.2 Curiosity perception: Curiosity-perception scores increased across participants, without a significant time-by-condition interaction.The time effect was F(1,72)=15.74, p-value<0.0001, while the interaction was nonsignificant at p=0.06.
9 Discussion
The study finds that GPT-3-generated cues can support children’s divergent question-asking, with open cues outperforming cues leading to predefined questions. The training also improved syntactic fluency and perceptions of curiosity across conditions.
- Open cues elicited more divergent questions than cues leading to one predefined question.The authors attribute this to open cues allowing several possible questions.
- The open-cue group showed a strong relationship between parent-reported curiosity trait and divergent question-asking performance.The authors suggest open activities may provide a favorable setting for expressing trait curiosity or transposing activities to children’s zones of proximal development.
- Children asked significantly more syntactically correct divergent questions after training in all three experimental conditions.
- Training positively improved children’s perceptions of curiosity across all experimental conditions.The measure included children’s fear of classmates’ negative judgment when asking questions.
10 Limitations and future directions
The authors identify limitations involving absent performance feedback, short training duration, researcher-led annotation, prompt-engineering demands, educational risks, and unaddressed metacognitive skills.
- The training provided no feedback on children’s question relevance, divergence, or syntactic construction.Future work proposes using LLMs to analyze questions and provide real-time feedback.
- Because children completed only one short session, engagement and positive outcomes may not generalize to longer-term curiosity-driven behavior.The authors propose longer training to assess long-term impact.
- Relevant GPT-3 outputs still require prompt-engineering knowledge, and less precise prompts or different settings may produce different results.The authors suggest teacher training to support good prompting practices.
- Using LLM systems without guidance in best practices and critical thinking may reduce children’s creativity and engagement.
- The study did not explicitly address metacognitive efficiency, a skill the authors identify as influencing curiosity-driven behavior.Future work targets metacognitive skills through dedicated pedagogical training.
11 Conclusion
The paper presents a GPT-3-driven, prompt-based system for generating pedagogical content to train children’s divergent question-asking. Field-study results indicate benefits for divergent QA, curiosity perception, and QA fluency.
- The study proposes a GPT-3-driven system that uses prompt-based methods to generate content for children’s divergent question-asking training.
- The field study found benefits from the tools in engaging children in divergent QA tasks and improving curiosity perception and QA fluency.
- Cue semantic relatedness was annotated on a five-point Likert scale from not related to super related to the text’s context.
- Cue divergence was scored from 1 to 3 according to whether it was explicitly stated, implied, or absent from the text.
- Questions counted when they were questions, related to the text, and not repeated, including relevant questions that did not use the agents’ cues.
- For a Big Bang text, accepted examples included questions about the universe’s initial temperature, the cause of the explosion, and the meaning of “microscopic.”
D Syntactic scores for evaluating children’s questions
The question-scoring grid evaluates syntactic construction, use of questioning words, and whether a question requires explanation of a mechanism or relationship. The study also reports that curiosity trait scores predicted divergent QA for the automated open agent.
- The scoring grid awards one point when a question requires explaining a mechanism or relationship rather than recalling a simple fact.
- Syntactic construction receives 1–4 points, from closed or declarative questions to correctly formed questions beginning with an interrogative word.
- Use of questioning words is separately scored from 1 to 3 points.
- For the automated open agent, children’s curiosity trait scores were strong predictors of divergent question-asking ability.
E Statistical tests for investigating the relationship between curiosity and divergent QA performance
The analysis tested whether children’s parent-reported curiosity scores were associated with divergent question-asking performance across the three experimental conditions.
- The ANCOVA examined divergent-question percentage as the dependent variable and parent-reported curiosity trait score as a covariate across all three conditions.
- The interaction between divergent-QA performance and curiosity scores was statistically significant across conditions, F(2,71) = 4.06, p-value=0.02.
- Bonferroni-corrected post-hoc comparisons found a statistically significant difference only for the open agent condition, p-value=0.006.
- Pearson-correlation comparisons showed similar relationships for the two incentive agents, z=1.7, p-value=0.14.