Source-linked AI summary

VLM-based automatic multi-granularity graph representation of building layouts for design informatics

Song Guo, Zhuoshi Chen, Maosu Li, Weimin Zhuang

arXiv:2608.24886v1cs.AI

TL;DR

Public-building floorplans lack task-adaptive, automatically extracted graph representations at multiple granularities. This paper introduces a zero-annotation VLM pipeline for multi-granularity graph construction and finds that optimal granularity depends on the downstream task.

  • Problem

    Automatic public-building layout graphs lack multi-level specifications and task-adaptive granularity, while existing workflows depend heavily on annotations and manual interventions.

  • Method

    The study defines fine-, meso-, and coarse-grained Level-of-Graphs and uses a four-step VLM pipeline for zero-annotation layout extraction.

  • Results

    Meso-grained graphs best support node-level zone prediction (Macro F1 = 0.647), whereas coarse-grained graphs best support graph-level layout evaluation (Spearman’s ρ = 0.610).

  • Takeaways & Limitations

    Granularity is an analytical choice that determines which spatial knowledge downstream models can access.

  • Takeaways & Limitations

    The dataset comprises 147 academic library floorplans, so generalizability to other building typologies remains to be verified.

Abstract

from arXiv · show

Architectural floorplan images encode rich relational knowledge among functional spaces, which underpins design retrieval, knowledge-based reasoning, and BIM enrichment through the building lifecycle. However, it remains challenging to automatically construct task-adaptive graph representations for public buildings. To address this gap, we first define a multi-granularity Level-of-Graphs (LoGs) for public building layouts. Methodologically, we present a Vision-Language Model (VLM)-based automatic LoG construction through node identification, edge inference, text parsing, and graph coarsening. VLM-generated representations are systematically evaluated and tested in real-world tasks, using 147 academic library floorplans worldwide as a case study. Experiments showed VLM-generated graphs were broadly consistent with human-labeled graphs (matched node ratio >= 92%; 509.3 s per floor plan for three-LoG graph generation). Meso-grained graphs yield the best node-level zone prediction (Macro F1 = 0.647, at 65% of fine-grained complexity), while coarse-grained graphs are most effective for graph-level layout quality evaluation (Spearman's \r{ho} = 0.610, at 16% of fine-grained complexity). By enabling scalable, annotation-free extraction of structured layout information from floorplan images, this study advances design informatics by converting plan images into knowledge representations, thereby enhancing the utilization of design information across the building life cycle.

1. Introduction

Building layouts encode discrete relational knowledge that graphs can support for design informatics, yet automatically extracting suitable representations for complex public buildings remains challenging. This study introduces a multi-granularity LoG framework and an annotation-free VLM pipeline to construct task-adaptive layout graphs from images.

  • Motivation: Building layouts encode relational information such as adjacency and connectivity, making them a distinctive form of engineering knowledge in the AEC domain.
  • Motivation: Graphs naturally represent building-layout structure and support downstream tasks including design retrieval, knowledge-based reasoning, and BIM enrichment.
  • Problem: Public-building complexity makes a single graph representation unsuitable for all tasks, because representation choice affects machine-learning performance and constrains GNN expressiveness.
  • Contributions: The study introduces fine-, meso-, and coarse-grained Level-of-Graphs representations for academic-library layouts at different levels of detail.
  • Contributions: A four-step VLM pipeline—node identification, edge inference, text parsing, and bottom-up coarsening—constructs multi-granularity graphs directly from layout images without manual annotation.
  • Contributions: Optimal graph granularity differs across architectural analysis tasks, establishing granularity–task alignment as an empirical contribution to design informatics.

2. Related Works

Prior work shows that architectural analyses require task-specific graph granularity, yet layout representations lack standardized multi-level specifications and remain difficult to automate. Existing methods depend largely on manual annotation or fragmented pipelines, motivating a zero-shot VLM-based multi-granularity approach for public-building floor plans.

  • Graph granularity: Coarse, meso, and fine-grained representations support whole-layout, connectivity, and detailed-perception analyses, respectively.Coarse graphs capture major spatial divisions efficiently, meso graphs preserve functional-area connectivity, and fine graphs support detailed perceptions.
  • Graph granularity: No standardized layout-detail levels exist, and graph granularity is usually dictated by annotation protocols rather than downstream task requirements.This limits graph reusability across purposes and leaves empirical task–granularity matching underexplored.
  • Automatic representation: Existing graph-construction methods rely mainly on manual annotation or partially automatic deep-learning pipelines, while fully automatic VLM approaches remain scarce.Manual workflows trace contours, assign labels, and mark connections or relationships; partial automation extracts regions but does not provide a fully integrated solution.
  • Automatic representation: Four persistent gaps include limited public-building dataset coverage, dependence on manual annotation, fragmented subtasks, and insufficient end-to-end automation.These limitations hinder computational convenience and broader reuse of layout representations.
  • VLM-based automation: The proposed framework integrates four subtasks into one VLM-based zero-shot pipeline requiring no task-specific training data and operating directly on complex public-building floor plans.The framework is positioned as a response to the annotation, integration, and public-building coverage gaps identified in prior work.
  • VLM-based automation: Prior automatic layout representation focuses mainly on residential buildings, depends on manual annotation, and underexplores VLM-based fine-grained spatial relationship extraction.The paper responds by introducing multi-granularity graph extraction and empirically linking graph granularity to architectural analysis tasks.

3. Multi-granularity Graph Representation

The paper introduces a three-level Level-of-Graphs representation for architectural floor plans and an automatic VLM pipeline that constructs fine-, meso-, and coarse-grained graphs. The framework also defines graph-fidelity and computational-complexity measures for evaluating the resulting representations.

  • Level-of-Graphs definition: Three LoG levels represent floor plans at coarse, meso, and fine granularities, respectively capturing rooms, functional components, and furniture groupings.LoG 1 is relatively stable, LoG 2 changes with spatial use, and LoG 3 changes with furniture layout.
  • Level-of-Graphs applications: LoG 1 captures physical partitioning, LoG 2 functional distribution and spatial relationships, and LoG 3 furniture-layout effects on user behavior and perception.The levels support applications ranging from footprint and room-area analysis to functional programming, accessibility, occupancy, and user experience studies.
  • Automatic graph construction: The VLM pipeline uses node identification, edge inference, text parsing, and bottom-up coarsening to construct multi-granularity graphs from floor plan images.The first three stages directly produce LoG 3, while coarsening derives LoG 2 and LoG 1.
  • Automatic graph construction: Node identification detects fine-grained functional spaces through four spatial passes, assigns predefined functional categories, and merges spatially proximate duplicate detections.The VLM infers furniture-group boundaries from symbols, zoning, labels, partitions, and global layout rather than treating walled rooms as nodes.
  • Graph coarsening: Bottom-up coarsening merges adjacent same-category LoG 3 nodes into LoG 2 nodes and all adjacent LoG 3 nodes into LoG 1 nodes, while preserving function composition semantics.The coarsening rules progressively abstract functional and physical structure from the fine-grained graph.
  • Evaluation framework: Graph evaluation measures fidelity through node and edge properties, matching alignment, semantic consistency, and topological consistency, alongside computational-complexity measures.Fidelity comparison requires matching VLM-generated and reference nodes and edges before similarity is computed.

4. Experiment Settings

Experiments apply the multi-granularity graph pipeline to a globally collected academic library floorplan dataset, comparing VLM-generated and human-labeled graphs across three LoG levels. Graph representations are evaluated for fidelity and downstream effectiveness through layout-quality regression and semi-supervised functional-zone classification.

  • Dataset and graph setup: The dataset contains globally available academic library cases from ArchDaily, excluding interior design, temporary design, and excessively small cases.Its expansive open areas, intricate circulation spaces, and diverse adaptive functions make it a testbed for evaluation.
  • Dataset and graph setup: For each case, the image-only pipeline produces three LoG graphs, while the human-labeled LoG 3 graph is coarsened to LoG 2 and LoG 1 for comparison.This yields six graphs per case across two sources and three granularity levels.
  • Graph alignment: Before fine-grained node matching, human-labeled graphs are aligned to VLM coordinates using control points and an affine transformation estimated via RANSAC least-squares.The VLM graph uses normalized 0–1000 coordinates, whereas the human-labeled graph remains in original image coordinates.
  • Downstream tasks: Layout-quality evaluation is formulated as regression using expert-assessed scores and a 42-dimensional graph-level feature vector for each source–granularity combination.SVR-RBF, Ridge regression, and XGBoost are tested with hyperparameter selection by Leave-One-Out Cross-Validation and repeated 5-fold cross-validation.
  • Downstream tasks: Functional allocation is formulated as semi-supervised node classification, inferring target-zone categories from partial knowledge of surrounding node functions.Targets include collections, reading areas, lounge, group study areas, and self study areas; each node uses a 25-dimensional feature vector and four representative GNN architectures are evaluated.

5. Results

VLM-generated multi-granularity graphs were structurally faithful to human-labeled graphs while preserving finer local detail and exhibiting some cross-zone connectivity omissions. LoG 1 was strongest for layout-quality evaluation and computational efficiency, whereas LoG 2 retained competitive zone-function prediction performance.

  • Graph fidelity: VLM graphs preserve individual furniture groups as separate nodes, producing denser local representations but slightly underrepresenting corridor-mediated cross-zone connectivity.Human-labeled graphs may instead merge adjacent same-function furniture groups, reducing local spatial resolution in dense functional regions.
  • Graph fidelity: Matched node ratios exceeded 0.92 across granularity levels, while node function, density, and path similarities remained high despite lower matched edge ratios.The results support fidelity in graph scale, functional composition, and topological reachability rather than exact node-by-node reproduction.
  • Downstream tasks and efficiency: LoG 1 achieved the strongest layout-quality performance, with Ridge regression reaching Spearman ρ = 0.610; LoG 3 VLM graphs closely matched human-labeled graphs at ρ = 0.510 versus ρ = 0.508.LoG 1 also reduced GNN complexity to 40% and feature-extraction cost to 16% of the LoG 3 baseline.
  • Downstream tasks and efficiency: LoG 2 reduced GNN complexity by approximately 35% and feature-extraction cost to 42% of the LoG 3 baseline while maintaining competitive zone-function prediction performance.LoG 1 provided greater computational savings and stronger layout-quality regression, making granularity selection task-dependent.
  • Feature analysis: 13 of 42 features showed significant correlations with consistent directions across human-labeled and VLM-generated graphs for layout-quality evaluation.Graph scale dominated; diameter and average shortest path were positively associated with quality, while density and closeness measures were negatively associated.

6. Discussion

The discussion presents LoGs and the VLM pipeline as advances that make public-building layouts computationally operable and scalable for design informatics. It emphasizes task-dependent graph granularity, human–VLM complementarity, practical applications, and limitations concerning attributes, evaluation scope, and generalizability.

  • Contributions: LoGs bridge visually rich but computationally inaccessible floor plans with data-driven design workflows by converting public-building layouts into computationally operable representations.The representations support downstream design and engineering tasks.
  • Contributions: The VLM-based pipeline reduces the manual annotation bottleneck and enables large-scale construction of layout graphs for public buildings.This supports scalable extraction of structured layout knowledge without relying on extensive manual graph annotation.
  • Applications: Multi-granularity graphs support rapid case retrieval, precedent-based design, and targeted analysis of building footprints, functional allocation, and user experience.They provide a unified framework that can enhance collaboration between urban planners and architects.
  • Findings: VLMs are stronger at local functional recognition than relational spatial reasoning, whereas human-labeled graphs better integrate architectural knowledge and three-dimensional reasoning.VLM-generated graphs nevertheless provide reliable functional recognition and consistent compositional representation, with comparable prediction performance at fine granularity suggesting knowledge can span formats.
  • Findings: Graph granularity is task-dependent: coarse- and meso-grained graphs perform best on different analyses, while fine-grained graphs can introduce local redundancies.Coarsening functions as feature selection by discarding locally redundant information.
  • Limitations: Limitations include unevaluated geometric attributes, architect-defined feature sets, and uncertain generalizability beyond 147 academic library floor plans.Future work should assess geometric estimation, broader downstream and AEC workflows, and additional public-building typologies.

7. Conclusion

The study introduces an annotation-free VLM pipeline for constructing multi-granularity Level-of-Graphs representations from public-building floorplans. The resulting graphs are broadly structurally faithful, while task-dependent granularity, collaborative inference, and future geometric reasoning shape the method’s implications and limitations.

  • Contributions: The study proposes Level-of-Graphs and a VLM-based pipeline that constructs multi-granularity layout representations from floorplan images without manual annotation.Academic libraries serve as the case study, with evaluation organized around structural fidelity and computational complexity.
  • Empirical findings: Matched node ratio >= 92%; 509.3 s per floor plan for three-LoG graph generation, while relational cross-zone connectivity remains limited.The pipeline generally matches human-labeled graphs in functional composition and global configuration.
  • Empirical findings: Optimal graph granularity is task-dependent across representative downstream tasks.The conclusion identifies meso-grained (LoG 2) graphs as best for one downstream task, though the supplied passage truncates the remaining result.
  • Implications: A human–VLM collaborative workflow could combine VLM-based element recognition with human-based relational inference as substitutable modules.The asymmetry between VLM-generated and human-labeled graphs motivates decomposing architectural spatial cognition into complementary components.
  • Implications: Coarsening acts as feature selection by discarding locally redundant detail and concentrating task-relevant spatial signals.Granularity therefore determines which spatial knowledge a model can access.
  • Limitations and future work: Future work should evaluate geometric attributes, incorporate three-dimensional spatial understanding, test generative tasks, and integrate representations with BIM enrichment.These directions aim to address relational reasoning gaps and extend downstream use beyond domain-knowledge-derived feature sets.

CRediT authorship contribution statement

The contribution statement identifies each author’s roles across writing, visualization, software, methodology, validation, analysis, data curation, investigation, and conceptualization.

  • Author contributions: Song Guo contributed across all listed areas, including writing, visualization, validation, software, methodology, investigation, formal analysis, data curation, and conceptualization.Guo is credited with both original-draft and review-and-editing writing.
  • Author contributions: Zhuoshi Chen contributed to original-draft writing, visualization, software, formal analysis, and methodology.
  • Author contributions: Maosu Li contributed to review-and-editing writing, validation, methodology, and conceptualization.
  • Author contributions: Weimin Zhuang contributed to review-and-editing writing and data curation.

Appendix A. VLM graph construction experiments · Appendix A.1. Node identification prompt

Appendix A presents the prompts for VLM-based graph construction and compares four state-of-the-art VLMs for node identification, informing the selection of Gemini 3 Pro. The node-identification prompt directs comprehensive extraction of fine-grained functional sub-space nodes from floorplan images using normalized coordinates and plain-text output.

  • Appendix A. VLM graph construction experiments: Four SOTA VLMs were compared on node identification, informing the choice of Gemini 3 Pro as the primary model.The comparison is presented alongside the graph-construction prompts in Figure A.1.
  • Appendix A.1. Node identification prompt: The prompt assigns the model the role of an expert in architectural floorplan semantic graph annotation.
  • Appendix A.1. Node identification prompt: Its goal is to extract all fine-grained functional sub-space nodes from the full floorplan image.
  • Appendix A.1. Node identification prompt: The required output format is plain text only, without explanations or markdown.
  • Appendix A.1. Node identification prompt: Before producing output, the model must scan the entire floorplan, identify functional furniture clusters and enclosed rooms, and then apply the region constraint.
  • Appendix A.1. Node identification prompt: Coordinates use a normalized 0..1000 system with the top-left at (0,0) and bottom-right at (1000,1000).

[REGION FILTER - MUST FOLLOW]

This section defines region-constrained node extraction from the full floorplan, requiring nodes to represent coherent functional sub-spaces and include qualifying enclosed rooms. Nodes must be placed at activity or geometric centers and classified using the prescribed category rules.

  • Region Filter: Only output nodes whose centers satisfy both the specified X and Y bounds.The region filter requires X ∈ [{x0}, {x1}] and Y ∈ [{y0}, {y1}].
  • Region Filter: Use the full floorplan context to infer each space’s true activity center.This applies even when only a restricted region is being reported.
  • Node Definition: Represent each distinct coherent functional sub-space as one node, using furniture, zoning, labels, partitions, and layout as evidence.Continuous identical furniture patterns within one coherent space should not create one node per row.
  • Empty Enclosed Rooms: Enclosed empty rooms bounded by continuous walls must be detected and output as nodes, even without an explicit door symbol.Place these nodes at the geometric center of the enclosed area.
  • Category Distinction: Distinguish categories by spatial and furniture cues, including teaching orientation for classrooms, individual focus for self-study, discussion layouts for group study, reading-hall patterns, and entrance counters for reception.The cited rules differentiate these functional categories using enclosure, desk orientation, table arrangement, and location.

Appendix A.2. Edge inference prompt

The edge-inference prompt directs the VLM to identify only the center node’s directly adjacent neighbors by cross-checking node-reference and original floorplan images. It distinguishes hard and soft connections and enforces a strict one-line output format.

  • Inputs and task: The VLM uses a node-labeled reference image and the original floorplan to identify the center node’s direct neighbors.The images are cross-checked to interpret node IDs, walls, doors, partitions, and functional boundaries.
  • Output format: The required output is exactly one plain-text line in the format CenterID|NeighborID1:Type,NeighborID2:Type,..., or CenterID| when no neighbors exist.Explanations and markdown are prohibited.
  • Adjacency criteria: Direct adjacency requires immediate local proximity and only a very short transition that is not a distinct functional sub-space.Long-range connections, links across intervening functional nodes, and nonlocal same-room relationships are excluded.
  • Edge types: Edges are classified as H for explicit door or hard connections and S for immediate open adjacency without a separating wall or partition.The prompt defines H and S as the only edge types.

Appendix B. Downstream experiments

The downstream experiments use separate graph-level and node-level feature sets for multi-granularity layout graphs. Features are standardized before modeling, with graph-level path features excluded from node-level inputs; class-wise results show Lounge and Self-study zones are essentially unpredictable in the five-class setting.

  • Feature sets: Graph-level and node-level features are defined separately for downstream tasks on multi-granularity layout graphs.The appendix presents the two feature sets in Table B.1 and Table B.2.
  • Feature sets: Graph-level features include connectivity, path, centrality, function-distribution, multifunctionality, and edge-type measures.Examples include largest connected-component node count, LCC diameter, mean betweenness, function class ratio, multifunctional node ratio, and adjacent edge ratio.
  • Implementation: Features are standardized before model fitting or input, while path features are not applicable to node-level prediction.For LoG 1 graphs, aggregated nodes use the dominant function label of their constituent fine-level elements for function class ratio and entropy.
  • Zone function prediction: Categories 4 (Lounge) and 6 (Self-study) are essentially unpredictable in the five-class zone function prediction setting.Figure B.1 reports class-wise F1 scores for the GINE-based zone function prediction task.
Loading 2608.24886v1…