Source-linked AI summary
Large Models for Time Series and Spatio-Temporal Data: A Survey and Outlook
Ming Jin, Yaxuan Kong, Yuxuan Liang, Chaoli Zhang, Siqiao Xue, Xue Wang, James Zhang, Yi Wang, Haifeng Chen, Xiaoli Li, Vincent S. Tseng, Yu Zheng, Lei Chen, Hui Xiong, Shirui Pan, Qingsong Wen
TL;DR
Large models are being explored for analyzing temporal data, but the transfer of language-centered representations to temporal patterns remains theoretically limited. This survey organizes the field through a taxonomy of models for time series and spatio-temporal data, reviews techniques and their strengths and limitations, and identifies open research needs.
Problem
Large models are being explored for temporal-data analysis, yet it remains unclear when language-centered representations, tokenization, and attention mechanisms capture temporal patterns or fail under distribution shift.
Method
The paper presents an extensive survey organized around a taxonomy of large models for time series and spatio-temporal data, including reviewed techniques and their strengths and limitations.
Results
The survey consolidates reviewed large-model techniques for time series and spatio-temporal data and examines their strengths and limitations.
Takeaways & Limitations
The survey frames LLM4TS and LLM4STD as practical repurposings of language-centered foundation models and highlights the need for theoretical frameworks and interpretability tools.
Takeaways & Limitations
The theoretical basis for transferring language-centered mechanisms to temporal data remains limited, including uncertainty under distribution shift.
Abstract
from arXiv · showhide
Temporal data, including time series and spatio-temporal data, are pervasive in real-world applications. Generated in massive volumes by physical and virtual sensors, they record dynamic system behaviors and enable a wide range of downstream tasks. Effectively analyzing such data is crucial to unlocking their rich information content. Recent advances in large language models and other foundation models have accelerated their use in time series and spatio-temporal data mining. These approaches not only improve pattern recognition and reasoning across diverse domains but also support progress toward artificial general intelligence that can understand and process temporal data. In this survey, we present a comprehensive, up-to-date review of large models tailored or adapted for time series and spatio-temporal data along four dimensions: data types, model categories, model scopes, and application areas/tasks. We organize existing work into two main groups: large models for time series analysis (LM4TS) and for spatio-temporal data mining (LM4STD), and further distinguish general-purpose from domain-specific models. We also curate related resources, including datasets, model implementations, and tools, organized by major application areas. Overall, this survey consolidates recent advances and highlights foundations, applications, resources, and open research opportunities in large model-centric temporal data analysis.
1 Introduction
Large models have expanded from language-centered foundations into reasoning across modalities and domains, motivating their application to temporal data analysis.
- 1 Introduction: Large language models developed for natural-language tasks have demonstrated emergent semantic representations and reasoning abilities.GPT-3 exhibited few-shot and zero-shot learning capabilities that were not achieved by GPT-2.
- 1 Introduction: Vision-language models extend large-model reasoning to visual and textual data across tasks including image captioning and visual question answering.
- 1 Introduction: Large-model applications have broadened to audio and speech analysis alongside established language and vision domains.
- 1 Introduction: This cross-domain progress raises whether large models can be effectively employed to analyze temporal data.
time series and spatio-temporal data?
Time series and spatio-temporal data capture system dynamics but require models that address both temporal patterns and spatial dependencies. This survey organizes emerging large-model methods, resources, applications, and research opportunities across these data types.
- Motivation: Time series and spatio-temporal data are pervasive sources of dynamic information, with spatio-temporal analysis additionally requiring spatial-dependency modeling.
- Motivation: Most historical temporal-data models are relatively small and task-specific, limiting broad transfer, semantic representation, and multi-task reasoning.
- Challenges: High-quality, diverse large-scale datasets remain limited across many temporal-data domains.
- Applications: Large-model applications span climate modeling, transportation, video understanding, forecasting, representation learning, and medical event prediction.
- Survey scope: The survey reviews LLMs and PFMs for temporal data across data categories, model scopes, application domains, and tasks.
- Unified taxonomy: Its taxonomy separates LM4TS from LM4STD and further categorizes models by type, scope, domain, and task.
- Resources: The survey compiles datasets, open-source implementations, evaluation benchmarks, and practical applications as references for future research.
- Future directions: It identifies future opportunities involving data sources, architectures, training, inference, and other research perspectives.
2 Background
The survey distinguishes language-centered LLMs repurposed for temporal tasks from PFMs designed or adapted as general-purpose backbones, and reviews both across time series and spatio-temporal analysis. It also introduces temporal-data definitions and describes multimodal, reasoning, and interaction capabilities relevant to these models.
- LLMs and PFMs: LLMs are language-centered foundation models repurposed for time series or spatio-temporal tasks, whereas PFMs are designed or adapted as general-purpose backbones for temporal, spatial, relational, visual, or multimodal data.The distinction is based on each model’s role in existing studies rather than an ontological separation.
- LLM adaptation: LLM adaptation for temporal tasks commonly uses multimodal repurposing or API-based prompting, with examples including OFA, Time-LLM, LLMTime, PPT, and VideoChat.Multimodal repurposing aligns target and source-task modalities, while API prompting wraps target modalities into natural-language prompts.
- Applications: Both LLM adaptation paradigms have shown promising results across time-series and spatio-temporal tasks in domains including transportation, energy, healthcare, environment, and finance.The survey includes methods that use LLMs as backbones, interfaces, or reasoning modules without necessarily training standalone foundation models.
- PFM capabilities and scope: PFMs are characterized by modality bridging, reasoning and planning, and interaction, while temporal and spatio-temporal PFMs remain at an early development stage.The survey focuses on PFMs for time series and spatio-temporal data, which are described as far from fully realizing the latter capabilities.
- Temporal data definitions: A time series is an ordered sequence of data points indexed in time, with univariate series having one value per time step and multivariate series having multiple dimensions.The survey represents multivariate data as a sequence of vectors and distinguishes univariate from multivariate examples such as temperature alone versus temperature combined with humidity.
3 Overview and Categorization
The survey organizes the literature into LM4TS and LM4STD, then separates LLM- and PFM-based approaches and further classifies them by purpose, domain, modality, and task. Its taxonomy is explicitly practical, while coverage of time-series foundation models remains nascent.
- Top-level organization: The taxonomy divides large temporal-data models into LM4TS for time series and LM4STD for spatio-temporal data.The time-series branch contains LLM4TS and PFM4TS, while the spatio-temporal branch contains LLM4STD and PFM4STD.
- Time-series models: LLM4TS covers LLMs used for time-series tasks, whereas PFM4TS covers foundation models explicitly designed for those tasks, regardless of whether LLMs are fine-tuned or frozen.The LLM/PFM distinction is practical rather than ontological and emphasizes how models function in existing studies.
- Field development: PFM4TS is relatively nascent, and existing models may not fully capture the potential of general-purpose PFMs, although the survey includes them for their future-facing insights.The authors nevertheless categorize these models as PFM4TS.
- Model scope: Each subdivision is further classified as general-purpose or domain-specific according to whether it targets general time-series tasks or restricted domains such as transportation, finance, and healthcare.This purpose-based distinction is applied across the survey’s model subdivisions.
- Spatio-temporal models: For spatio-temporal data, the taxonomy explicitly organizes LLM4STD and PFM4STD by domains and focuses on spatio-temporal graphs, temporal knowledge graphs, and video data.Representative tasks appear as leaf nodes, while model-scope categorization is conditioned on modality-specific problem definitions.
- Field development: PFM4STD has developed more extensively than PFM4TS, with current work mainly targeting spatio-temporal graphs and video data and often emphasizing multimodal bridging.The survey presents these categories and related works through its taxonomy and subsequent sections.
4 Large Models for Time Series Data
This section reviews large language and foundation models for time series, distinguishing general-purpose and domain-specific approaches across forecasting, reasoning, and applications. Recent work expands from forecasting adaptation toward multimodal interaction, broader reasoning, and cross-domain temporal modeling.
- The survey organizes time-series large models by general-purpose versus domain-specific applications.
- Time series analysis supports forecasting, imputation, anomaly detection, and classification across diverse real-world domains.
- PromptCast frames forecasting as a natural-language input-output task, offering a “code less” alternative to increasingly complicated architectures.
- Time-LLM reprograms time series using prompting, tokenization, decomposition, or related mechanisms, achieving state-of-the-art forecasting and strong few-shot and zero-shot performance.
- Recent multimodal models unify numerical and textual modalities, support conversational analytics, and extend time-series systems toward perception, extrapolation, and decision-making.
- TimeOmni-1 introduces a reasoning suite and unified model, signaling movement from pattern matching toward general time-series reasoning.
- Domain-specific applications fuse temporal signals with contextual information in transportation, finance, healthcare, and clinical decision support.
5 Large Models for Spatio-Temporal Data
The survey reviews large models for spatio-temporal data across spatio-temporal graphs, temporal knowledge graphs, and videos, organizing the literature by model type and scope.
- Large models for spatio-temporal data are grouped into spatio-temporal graphs, temporal knowledge graphs, and videos.
- The literature is further organized by model type and scope.
5.1 Spatio-Temporal Graphs
Large models for spatio-temporal graphs combine language-level semantics with graph topology and temporal dynamics. The field is moving from task-specific graph models toward general-purpose and domain-specific foundation models.
- Spatio-temporal graph forecasting captures spatial correlations with graph networks and temporal dependencies across time steps.
- LLMs enrich spatio-temporal graphs by integrating textual, visual, and structured modalities into contextualized representations.
- Graph-aware LLM architectures combine graph encoders or tokenizers with language-model backbones for spatio-temporal prediction.
- These models integrate language semantics with graph topology, while practical limits include graph quality, unstable spatial dependencies, and adaptation cost for dynamic networks.
- Foundation-model approaches use contrastive pre-training, mixture-of-experts, and domain-specific architectures to improve transfer, generalization, and adaptability.
- Climate models demonstrate the domain-specific direction through fast medium-range forecasting, efficient fine-tuning, calibrated long-range prediction, and computational gains.
5.2 Temporal Knowledge Graphs
Temporal knowledge graph models use large language models to reason over evolving facts for forecasting and completion. Their methods progress from historical prompting toward retrieval, rule adaptation, and structure-aware inference.
- Temporal knowledge graphs extend entity-relation triples with timestamps, representing structural and temporal dependencies among evolving facts.
- LLM-based temporal knowledge graph methods address forecasting and completion tasks.
- Forecasting approaches incorporate historical chains, dynamically adapt temporal logic rules, combine structural and textual views, and replay analogous event sequences.
- Overall, these methods can infer missing links in dynamic environments by leveraging semantic understanding and structural evolution.
- Completion approaches formulate temporal link prediction as event generation, model relation dynamics by frequency components, and combine textual semantics with temporal graph context.
- Effectiveness remains constrained by incomplete facts, sparse timestamps, and difficulty aligning symbolic graph dynamics with continuous temporal changes.
5.3 Videos
Videos are spatio-temporal data represented as image sequences, and video understanding has progressed from frame-based and 3D convolutional methods toward multimodal, long-context, and reasoning-driven models.
- Video represents visual information as a sequence of images or frames that collectively convey motion and temporal changes.
- Conventional video understanding uses separate-frame 2D CNNs, spatio-temporal 3D CNNs, and Transformers for long-range dependency modeling.
- Large multimodal models jointly process visual and textual modalities to extract contextual information and support transfer across domains.
- Recent video-language systems extend capabilities through multimodal alignment, temporal reasoning, causal knowledge extraction, sparse memory, and programmatic reasoning.
- Long-form video research compresses or selects visual information to handle hour-scale inputs while retaining reasoning over extended temporal horizons.
- Domain-specific models adapt multimodal systems to traffic and sports, supporting captioning, event recognition, commentary generation, foul detection, and question answering.
5.4 Methodological Convergence and Cross-Domain Insights
LM4TS and LM4STD share challenges in tokenizing temporal information, modeling long-range dependencies, and aligning modalities, while their structures create complementary trade-offs and transfer opportunities.
- Both LM4TS and LM4STD must convert temporal observations into model-readable tokens while preserving temporal order, context, and domain semantics.
- Long-range dependencies are managed through patching, decomposition, lightweight adaptation, graph summarization, pooling, memory modules, and frame selection.
- LM4STD can augment temporal tokens with road networks, entity relations, trajectories, and video-language alignment beyond LM4TS representations.
- LM4TS is simple and scalable but may lose fine-grained numerical information, whereas graph and video methods trade structural or visual richness against data and memory constraints.
- Structured representations and compression from spatio-temporal methods can inform time-series analysis, while prompt- and adapter-based transfer can reduce adaptation costs for other modalities.
6 Resources and Applications
The survey compiles datasets, models, and tools for time-series and spatio-temporal applications, spanning traffic, healthcare, weather, finance, video, event sequences, and general benchmarks.
- Common resources are organized by application area, with datasets, models, and tools summarized in Table 3.
- Traffic: Traffic resources include road-network speed and flow datasets, traffic videos with question-answer pairs, large-scale sensor data, SUMO, and SafeGraph Data for Academics.
- Healthcare: Healthcare resources support disease progression, mortality, and time-dependent risk analysis using ECG, clinical notes, labeled readmission and mortality data, and clinical corpora.
- Weather: Weather resources cover hourly changes in weather parameters and global climate-model evaluation, alongside foundational models such as Pangu-Weather and GraphCast.
- Finance: Finance resources combine prices, corporate events, employment changes, policies, and tweets, with financial LLMs emerging for prediction and decision-making.
- Other applications: Video, event-sequence, and general benchmarks span question answering, captioning, irregular event modeling, anomaly detection, classification, and spatio-temporal modeling.
7 Outlook and Future Opportunities
The outlook identifies open challenges in theory, multimodal alignment, adaptation, interpretability, privacy, and safe decision-making for large temporal models.
- Theoretical understanding remains limited regarding when linguistic representations, tokenization, and attention capture temporal patterns or fail under shift and modality mismatch.
- Future work should analyze transfer conditions and task-specific limits across forecasting, anomaly detection, classification, and spatio-temporal reasoning.
- Models should align multimodal signals with different temporal resolutions to combine textual information with numerical temporal dynamics.
- Deployable systems must address edge-cloud collaboration, latency, privacy, robustness, and continuous adaptation under non-stationary patterns without catastrophic forgetting.
- Interpretability research should provide rationales and support counterfactual analysis, causal inference, root-cause analysis, scenario simulation, and intervention planning.
- Privacy-preserving learning and uncertainty-aware agent policies are needed because sensitive data may be memorized and distribution shifts can produce unsafe decisions.
8 Conclusion
The survey offers an extensive, up-to-date review of large models for time series and spatio-temporal data, introducing a taxonomy and synthesizing techniques, limitations, resources, applications, and future opportunities.
- The survey reviews large models for time series and spatio-temporal data from a fresh taxonomy-based perspective.
- It categorizes reviewed models and summarizes prominent techniques within each category.
- The survey examines the strengths and limitations of techniques across the reviewed categories.
- It highlights foundations, applications, resources, and open research opportunities in large model–centric temporal data analysis.