Source-linked AI summary

Public transport networks: empirical analysis and modeling

C. von Ferber, T. Holovatch, Yu. Holovatch, V. Palchykov

arXiv:0803.3514v1physics.soc-phphysics.data-an

TL;DR

The paper addresses limited comparative evidence on the statistical properties of large urban public transport networks and the need to analyze complete multimodal systems. It surveys fourteen cities across multiple network representations, then proposes a route-growth model based on self-avoiding walks. The networks show correlated small-world structure and the model reproduces many essential PTN features, while route-based harness behavior requires information beyond the simple graph representations.

  • Problem

    The paper seeks comprehensive comparative evidence on PTN properties across large cities and across complete multimodal networks rather than isolated transport sub-networks.

  • Method

    The authors analyze fourteen-city PTN data in multiple representations and use the comparative empirical results to construct a route-growth model based on self-avoiding walks.

  • Results

    The networks are strongly correlated small-world structures with high clustering and comparatively low mean shortest paths, and the model captures many essential real-world PTN features.

  • Takeaways & Limitations

    Comparing L, P, C, and B-spaces provides passenger-relevant network measures and enables more adequate PTN modeling.

Abstract

from arXiv · show

We use complex network concepts to analyze statistical properties of urban public transport networks (PTN). To this end, we present a comprehensive survey of the statistical properties of PTNs based on the data of fourteen cities of so far unexplored network size. Especially helpful in our analysis are different network representations. Within a comprehensive approach we calculate PTN characteristics in all of these representations and perform a comparative analysis. The standard network characteristics obtained in this way often correspond to features that are of practical importance to a passenger using public traffic in a given city. Specific features are addressed that are unique to PTNs and networks with similar transport functions (such as networks of neurons, cables, pipes, vessels embedded in 2D or 3D space). Based on the empirical survey, we propose a model that albeit being simple enough is capable of reproducing many of the identified PTN properties. A central ingredient of this model is a growth dynamics in terms of routes represented by self-avoiding walks.

I. INTRODUCTION

The paper provides a comprehensive comparative survey of large urban public transport networks, combining all transport modes and multiple network representations. It uses the empirical findings to propose a growth model intended to reproduce many PTN properties.

  • Research objectives: The study surveys statistical properties of public transport networks in major cities using data from networks of previously unexplored size.The cities were selected to represent different geographical, cultural, and economic backgrounds, extending earlier studies of smaller systems.
  • Data scope: The PTNs include 1,881 routes and 44,629 stations in the Los Angeles example, illustrating the large network sizes considered.The database includes buses, electric trolleybuses, ferries, subways, trams, and urban trains.
  • Modeling objective: The empirical survey motivates a simple growth model designed to capture many characteristic statistical features of real-world PTNs.The paper also considers scale-free behavior in relation to network evolution and possible operational optimization.
  • Research objectives: Unlike studies restricted to buses, trams, or subways, the analysis treats all public transport modes in each city as one complete network.The authors argue that individual systems are subgraphs of a wider urban transportation system and that combining modes can substantially change network properties.
  • Analytical framework: The study systematically compares PTN characteristics across different representations, including successor-based and common-route neighborhood relations.The primary topology is specified by routes serving ordered station sequences, while alternative representations define station neighborhoods differently.

II. PT NETWORK TOPOLOGY

The paper represents public transport networks in several graph spaces, each encoding different station, route, or adjacency relations. These representations make standard network measures interpretable in passenger-relevant terms and support comparative analysis.

  • Network representations: L-space represents stations as nodes, linking two stations when at least one route services them consecutively.The underlying route data are ordered station lists, including the stations served between terminals or along a round trip.
  • Network representations: L′-space extends L-space by retaining multiple links or weights according to the number of services between consecutive stations.Unlike simple L-space, this representation preserves service multiplicity between neighboring stations.
  • Network representations: P-space links stations that share a route, so each node’s neighborhood represents destinations reachable without changing transport.Multiple or weighted links can also be retained in an extended P-space representation.
  • Network representations: B-space is bipartite: route nodes connect only to the station nodes they service, with no links between nodes of the same type.Projecting B-space onto station nodes yields P-space, while projecting onto route nodes yields C-space.
  • Interpretation and data: In passenger terms, mean shortest paths in L-space measure stops traveled, whereas those in P-space measure transport changes between stations.C-space shortest paths analogously describe changes between routes, and the characteristics are calculated from publicly available transport data.

III. LOCAL NETWORK CHARACTERISTICS

The paper compares local PTN structure across L-, P-, and C-space representations, focusing on node-degree distributions and clustering. These properties vary across cities and representations, revealing organization linked to transport-network construction.

  • Node-degree distributions: PTN node-degree distributions differ across L-, P-, and C-space representations and among cities.The analysis uses distributions and cumulative distributions to compare these representations across urban networks.
  • Node-degree distributions: Almost half of the analyzed cities show power-law decay in L-space, whereas P-space can exhibit varied behavior rather than a universal exponential form.The study reports that different combinations of distributions in L- and P-space may occur across the less homogeneous city sample.
  • Node-degree distributions: C-space degree distributions decay exponentially or faster; Berlin, London, and Los Angeles are approximated by exponential decay.The C-space result is based on the distributions shown for the cities in Figs. 4c and 5c.
  • Clustering coefficients: ⟨CP⟩ = 0.7 ÷ 0.9 in P-space, compared with ⟨CL⟩ = 0.02 ÷ 0.1 in L-space.P-space clustering is high because each route forms a complete graph; values are normalized against a random graph for comparison.
  • Clustering coefficients: Clustering coefficients in L- and P-space exceed random-graph values by orders of magnitude, while the difference is less pronounced in C-space.Values vary strongly across the fourteen-city sample, indicating measurable organizational differences among PTNs.
  • Clustering coefficients: In P-space, ⟨CP(k)⟩ decreases with node degree k because higher-degree stations are more likely to belong to multiple routes.The change is expected near a k corresponding to the mean number of route stops, and the decay follows a power law.
  • Clustering coefficients: The exponent β ranges from 0.65 for São Paolo to 0.96 for Los Angeles, with a mean value of 0.82.The paper compares these values with the β = 1 exponent found in a simple star-like network model.

A. Path length distribution

The paper analyzes shortest-path distributions and their dependence on node degree in multiple PTN representations. Path lengths generally follow fitted forms, decrease with node degree, and show representation- and city-specific variation.

  • Path length definition: Shortest-path averages are restricted to pairs within the same connected component and are calculated for the largest connected component.Accordingly, N denotes the number of nodes in the giant connected component.
  • Path length distribution: Mean shortest-path distributions are generally reproduced by an asymmetric unimodal fit with parameters A, B, and C.Los Angeles in L-space shows a second local maximum, qualitatively associated with more than one network community.
  • Degree dependence: Mean shortest paths in L-space decrease with node degree k and are approximated by a power law for most analyzed cities.The exponent αL ranges from 0.17 to 0.27, with examples including Berlin αL = 0.23 and Paris αL = 0.15.
  • Degree dependence: In P-space, shortest-path length counts the minimal number of routes needed to reach another station.Higher-degree stations provide easier access to other routes, so ℓP(k) is expected to decrease as k increases.
  • Degree dependence: αP = 0.09 for Sydney to αP = 0.17 for Dallas, centered around αP = 0.12 ÷ 0.13.The paper reports this range for the power-law dependence of mean P-space path length on node degree.
  • Degree-product dependence: Mean P-space path length tends to decrease as the product kq increases, although the data scatter is more pronounced than in L-space.The trend is shown for several cities, with guide lines characterizing the decay.

B. Centralities

The paper evaluates node importance through path-length and path-mediation centralities across multiple PTN representations, finding representation-dependent relationships and scaling patterns.

  • Closeness and graph centralities measure node importance using shortest-path lengths to other nodes.
  • Mean closeness and graph centralities correspond similarly with mean shortest path length across L, P, and C spaces.Because these centralities use inverse path length, the correspondence has a negative slope.
  • In Paris, mean betweenness centrality follows an approximate power law with degree in L- and C-spaces, with exponent η = 2 ÷ 3.The L-space relationship is especially clear, while the C-space relationship is less precise.
  • Paris’s L-space betweenness distribution has an exponent δ ≈1.5, consistent within accuracy with the resulting scaling relation.
  • B- and P-space betweenness-degree plots show two regimes associated with distinct node roles or route-membership patterns.In B-space, low-degree behavior describes station nodes and high-degree behavior route nodes; in P-space, low-degree stations often belong to one route and have low betweenness.

C. Harness

The harness effect captures parallel route sequences, a route-level feature that requires representations preserving route continuity and has implications for network organization and spatial optimization.

  • The harness effect describes routes proceeding in parallel along sequences of stations sharing the same street or track grid.It is also relevant to spatially embedded networks such as cables, neurons, pipes, and vessels.
  • Some harness distributions follow power laws, whereas Berlin, Dallas, Düsseldorf, London, and Moscow are better fitted by exponential decay.
  • Los Angeles shows a crossover from power-law behavior at small sequence lengths to an exponential tendency at larger lengths.
  • Observed power laws have γs = 2 ÷ 4, while exponential distributions have characteristic scales r̂s = 1.5 ÷ 4.
  • Harness behavior is invisible in the paper’s simple graph representations and requires route-continuity information beyond the listed multigraph representations.The distribution is related to network flow and load and may inform spatial optimization of space-consuming links.

V. GENERALIZED ASSORTATIVITIES

The paper generalizes assortativity to correlations between multiple node characteristics and shows that mixing patterns depend strongly on the network representation and mean path length.

  • Generalized assortativities measure correlations between observables assigned to the endpoints of network links, including degree, clustering, path-based centralities, stress, and betweenness.
  • In L-space, PTNs separate into groups with finite positive degree assortativity r(1)L = 0.1 ÷ 0.3 or near-neutral values r(1)L = −0.02 ÷ 0.08.Both large and medium networks occur in each group, indicating no correlation between network size and degree assortativity.
  • In P-space, almost all cities have very small positive or negative degree assortativity values, unlike the clearer assortative mixing observed in C-space.
  • Second-nearest-neighbor assortativity is non-negative and generally stronger than nearest-neighbor assortativity when the latter is significant.PTNs neutral with respect to nearest neighbors also remain neutral for second nearest neighbors.
  • Stress, betweenness, and clustering assortativities generally show no correlation, except for small positive stress and betweenness assortativities in L-space.The paper relates the L-space effect to relatively long paths and small node degrees.

A. Motivation and description of the model

The model treats routes as the essential growing elements of PTNs, using self-avoiding walks on a two-dimensional lattice with preferential attachment to previously served stations.

  • Route-based growth is central because routes preserve the PTN’s finite station sequences and bipartite structure.
  • Random-network or preferential-attachment models alone would reproduce degree distributions but not the route-based structure required for PTNs.
  • The model includes route attraction because observed harness distributions indicate that routes prefer to service common stations.
  • The model represents each route as a self-avoiding walk of adjacent stations on a square lattice.This choice reflects the near absence of loops and the observed fractal dimension df = 4/3 of PTN routes.
  • Subsequent routes are generated as self-avoiding walks using preferential attachment rules based on previously visited lattice sites.

B. Global topology of model PTN

The model’s global topology is controlled mainly by parameters a and b, which regulate disconnected-route initiation and spatial route spreading. Small or vanishing a with b ≥ 0.5 produces realistic-looking PTN maps.

  • Parameters a and b remain after fixing the number of routes, stations per route, and lattice size, while dependencies on R, S, and X are studied separately.
  • Increasing b broadens route coverage from near-overlapping paths toward wider spatial distribution while retaining a densely covered center.
  • Parameter a controls whether new routes start outside the existing network: a = 0 yields one connected component, whereas finite a can create disconnected components.
  • For b = 0.5, increasing a beyond 15 produces a sharp transition from one dominant cluster to multiple clusters, eventually covering the lattice more homogeneously.

C. Statistical characteristics of model PTN

Across multiple network representations, the model reproduces many statistical properties of real PTNs, including path lengths, degree distributions, centrality correlations, and harness-distribution regimes. Its agreement is strongest in P-, C-, and B-spaces, while square-lattice geometry limits L-space behavior.

  • Most probable simulated path lengths are ˆℓP ∼3 and ˆℓC ∼2, while L-space values are of the order of route length S.
  • Simulated C-space degree distributions tend toward exponential behavior with scales increasing with R and decreasing with S.
  • The model reproduces the global decay properties of B-space station degree distributions, although individual simulated distributions can be non-monotonic.
  • Mean betweenness-degree correlations are qualitatively reproduced in C-, P-, and B-spaces, but L-space behavior is not reproduced because of square-lattice geometry.
  • Changing b tunes harness distributions from power-law decay at b = 0.2 toward exponential decay at b = 1.0 by reducing parallel-route hubs.
  • The model captures many essential real-world PTN features, especially when multiple network representations are compared.

VII. CONCLUSIONS

The paper surveys statistical properties of large urban PTNs and uses comparative analysis across network representations to construct a model reproducing many observed properties. Its conclusions also connect network statistics to passenger-relevant characteristics and PTN-specific harness structure.

  • The study surveys PTNs from cities with previously unexplored network sizes and proposes a model reproducing a majority of their statistical properties.
  • Comparing L-, P-, C-, and B-space representations enabled the PTN modeling presented in the paper.
  • The analyzed networks are strongly correlated small-world structures with high clustering and comparatively low mean shortest-path values.
  • For Paris, average separation is 5.4 stations and travel requires 1.7 changes between stops.
  • Power-law degree distributions occur in many L-space and some P-space networks, while other surveyed networks show exponential decay.
  • Harness distributions quantify routes proceeding in parallel across sequences of stations, a feature shared by PTNs and related networks.
Loading 0803.3514v1…