Source-linked AI summary

Investigating the use of ChatGPT for the scheduling of construction projects

Samuel A. Prieto, Eyob T. Mengiste, Borja García de Soto

arXiv:2302.02805v1cs.HCcs.AI

TL;DR

The paper investigates whether ChatGPT can assist construction scheduling, using a simple project and participant evaluation of generated schedules and interaction. ChatGPT generally produced logical, coherent schedules and a positive interaction experience, but omissions, irrelevant tasks, limited complexity, and few participants constrain adoption and generalization.

  • Problem

    Construction scheduling is a potential application for natural-language models, but the paper investigates their applicability and limitations in this setting.

  • Method

    The study used ChatGPT to generate a schedule for a simple partition-wall project and evaluated outputs and user interaction through six participants and baseline comparisons.

  • Results

    ChatGPT generally generated logical and coherent task breakdowns and sequences, while participants rated the output positively overall but identified omissions and irrelevant tasks that reduced accuracy and reliability.

  • Takeaways & Limitations

    The findings indicate potential for ChatGPT to support preliminary and time-consuming construction scheduling tasks, while tables and charts remain more intuitive final display formats.

  • Takeaways & Limitations

    The case study used a very limited project complexity and a limited participant pool, so the results should not be generalized without larger and more complex studies.

Abstract

from arXiv · show

Large language models such as ChatGPT have the potential to revolutionize the construction industry by automating repetitive and time-consuming tasks. This paper presents a study in which ChatGPT was used to generate a construction schedule for a simple construction project. The output from ChatGPT was evaluated by a pool of participants that provided feedback regarding their overall interaction experience and the quality of the output. The results show that ChatGPT can generate a coherent schedule that follows a logical approach to fulfill the requirements of the scope indicated. The participants had an overall positive interaction experience and indicated the great potential of such a tool to automate many preliminary and time-consuming tasks. However, the technology still has limitations, and further development is needed before it can be widely adopted in the industry. Overall, this study highlights the potential of using large language models in the construction industry and the need for further research.

1 Introduction

The paper examines ChatGPT as a natural-language tool for construction scheduling, aiming to assess its applications, limitations, and performance in a preliminary multi-user case study.

  • Natural language processing can improve construction efficiency, accuracy, and communication through applications such as document information extraction.
  • The study evaluates GPT for developing an automated construction schedule from natural-language prompts.
  • The paper explores possible applications and limitations of ChatGPT for construction scheduling and resource loading.
  • The preliminary case study involved multiple users generating a resource-loaded schedule for a simple project from a detailed natural-language description.
  • Participants evaluated outputs using accuracy, efficiency, clarity, coherence, reliability, relevance, consistency, scalability, and adaptability.

2 Literature review

The literature review surveys language models and automated construction scheduling, emphasizing that existing scheduling approaches can produce logical sequences but depend on data, models, or manual intervention.

  • GPT, BERT, and T5 represent major language-model families used for language understanding and generation tasks.
  • Construction schedule automation includes BIM-driven and machine-learning-based approaches.
  • BIM-based methods use model information and preset rules or historical knowledge to automate task dependency and sequencing.
  • BIM schedule generation often requires manual intervention because models may omit environmental factors, temporary structures, equipment, materials, and site-specific methods.
  • Prior language-based scheduling research used GPT, LDA, and LSA to interpret task descriptions, develop dependencies, and cluster tasks.
  • Graph-based scheduling approaches can recycle past best practices to optimize resource usage, task sequence, and duration.
  • Machine-learning approaches have produced promising construction schedules, but their performance depends on data availability and users’ technical and infrastructural capacity.

3 Methodology

The methodology uses a simple construction project and repeated participant interactions with ChatGPT to assess schedule generation, task assignment, output quality, and user experience.

  • The experiment evaluates ChatGPT as a project-management aid for producing a logical and accurate task breakdown for a simple construction project.
  • Different participants challenged and modified ChatGPT’s original plan to evaluate its responses.
  • Evaluation parameters included accuracy, efficiency, clarity, coherence, reliability, relevance, consistency, scalability, and adaptability.
  • Accuracy was measured against a human project manager’s baseline schedule and assignments, while efficiency included creation time and error-correction time.
  • The input described work details, dimensions, materials, completeness, and due dates, with expected outputs including task lists, project sequencing, and responses to user changes.
  • A survey assessed generated-output quality and participants’ interaction experience.

4 Case study

The case study applies ChatGPT to scheduling a partition-wall project, using a common scope prompt, a human baseline, structured output requests, and feedback from six participants.

  • The project involved adding a partition wall in an existing space, with all participants receiving the same initial scope information.
  • ChatGPT was asked to retain the provided information before generating a suitable project schedule.
  • The scope described a 4-by-4-meter rectangular room, 3-meter concrete masonry-unit walls, and a new concrete masonry-unit partition.
  • A typical schedule in Table 1 served as the baseline for comparing ChatGPT’s generated schedule.
  • The requested output structure included task name, priority, dependencies, required people, and expected duration, with an option to display the information in columns.
  • The survey was distributed to six participants with different construction and artificial-intelligence skills and qualifications.
  • The baseline listed a total duration of 15.5 days, including weekends, equivalent to 11.5 work days.

5 Results and discussion

Across six participant experiments, ChatGPT generated fast, mostly logical and coherent construction schedules, but scope mismatches and missing construction details reduced accuracy and reliability. Participants viewed the interaction and potential uses positively, while noting that the simple case and limited participant pool constrain generalization.

  • Schedule generation: ChatGPT produced a logical, coherent, though very linear sequence of construction tasks in all six participant experiments.The output extrapolated task breakdowns and dependencies from the prompt, often within seconds.
  • Scope alignment: Wooden-door tasks were omitted from all six schedules, while three responses added demolition and other responses added potentially unnecessary foundation or rebar work.The comparison identified missing door-frame installation and protection tasks, along with scope-inconsistent activities.
  • Schedule generation: ChatGPT’s proposed door sequence placed door installation after wall erection rather than after plastering, potentially exposing the door to damage from intermediate tasks.The sequence could be adequate for a door frame alone, but no clear frame-only breakdown was provided or inferred.
  • Duration and resources: ChatGPT did not seem to account for drying times, whereas the baseline included drying for plastering and painting.The baseline therefore assigned more than three days to each of those tasks.
  • Duration and resources: The highest deviation from the baseline worker estimate was two workers, based on comparisons of maximum and minimum possible differences.Worker estimates were compared task by task against baseline ranges across participant responses.
  • Participant evaluation: Participants rated the output positively overall, but accuracy and reliability received the lowest scores because some generated tasks did not match the project scope.Tables and charts remained easier to read than the dialogue format, although participants saw ChatGPT as useful for extracting information and supporting text-processing tasks.
  • Participant evaluation: Participants identified possible uses including initial consulting, safety and risk identification, basic design, cost estimation, and contract processing.These proposed applications extended beyond scheduling to several construction-management and text-processing activities.
  • Limitations: The case study supports preliminary evaluation, but its simple project scope and limited participant pool prevent generalization and require more complex future studies.The authors call for studies that better resemble actual construction projects and use a larger statistical pool.

6 Conclusions and future work

The study finds promising but limited applicability for ChatGPT-assisted construction scheduling: interaction was positive and performance was reasonable in a simple use case, while real-project adoption requires reliable performance and further research.

  • ChatGPT showed reasonable performance and positive interaction experience in a simple construction-scheduling use case.The authors note that ChatGPT was not specifically trained for this application.
  • Several significant flaws would limit applying the tool in a real construction project.
  • Field-specialized tools could automate repetitive and time-consuming scheduling tasks if they achieve consistent and reliable performance.
  • Future studies should test construction-trained GPT models in more complex scenarios and survey larger participant pools.

Appendix A

Appendix A presents data extracted from the first experiment by six participants, including a construction sequence and alternative duration totals. The schedule covers demolition through inspection, with tasks generally ordered from foundational work to finishing and cleanup.

  • Table A.1 is identified as data extracted from the first experiment by six participants.
  • Site cleanup and final inspection appear as concluding activities after the partition construction and finishing work.
  • The planned sequence begins with demolition, foundation excavation, steel reinforcement, concrete pouring, and erection of the new partition.
  • The schedule then covers door installation, two stucco layers, and two white latex paint layers on the new partition.
  • The participant-extracted data report total durations of 14 days and 15 days in two schedule entries.
Loading 2302.02805v1…