Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
72,870 characters · 32 sections · 31 citation commands
Diagnosing Urban Street Vitality: A Visual-Semantic and Spatiotemporal Framework for Street-Level Economics
\email{[email removed]} \orcid{0009-0001-6642-5357} \orcid{0009-0009-9269-9560} \email{[email removed]}
\orcid{0009-0001-2511-8692} \email{[email removed]}
\orcid{0000-0003-3895-4689} \email{[email removed]} \authornote{Corresponding author.} \email{[email removed]} \orcid{0000-0002-5271-0472}
\ccsdesc[500]{Applied computing} \ccsdesc[500]{Applied computing}
A persistent puzzle in urban economics is why streets with seemingly identical physical infrastructure, geographical advantages, and high establishment densities often exhibit drastically divergent economic trajectories. Urban policymakers and investors increasingly confront the "illusion of prosperity"—situations where high nominal commercial density masks underlying vulnerability, low economic yield, and hidden recession. Understanding and resolving this divergence is fundamental to deciphering the microscale dynamics of urban life and correcting the misallocation of spatial resources during urban renewal Li2022_IJGI, Ying2019, He2024, Yang2024, Chen2022.
Traditional urban economic models and conventional metrics struggle to explain this micro-level divergence. Metrics aggregated at district or grid levels often obscure the fine-grained interactions among commercial establishments and pedestrian flows KOO2023, Chen2019, Zhang2021_EPB. While the advent of big data has introduced Point of Interest (POI) density as a proxy for vitality, this approach suffers from critical information asymmetry. It frequently identifies “zombie stores” (establishments that remain registered but are operationally closed) as active participants, thereby yielding false positives and failing to capture actual market exit behaviors. To address this data-driven opacity, recent scholarship has employed Street View Imagery (SVI) Jiang2022, Li2022, Zhang2019_SocialSensing to visually audit store closures. However, these visual approaches remain reductionist, treating all active storefronts as homogenous entities based on a binary “open vs. closed” status.
Consequently, existing literature overlooks two decisive micro-mechanisms driving this economic divergence: the quality heterogeneity of commercial presence (e.g., the economic disparity between a local survival-oriented shop and a global premium chain) and the spatial externalities of large commercial anchors. This semantic opacity prevents a true understanding of whether a street's vitality stems from endogenous high-quality commercial agglomeration or exogenous reliance on a nearby commercial anchor. Furthermore, the inherent complexity of storefront signage represents unstructured market signals that traditional Optical Character Recognition (OCR) fails to decode, leaving the brand-driven spillover effects unquantified.
To untangle this economic puzzle, we propose a visual-semantic and spatiotemporal framework for assessing street-level economic vitality, moving beyond simple enumeration toward a diagnostic evaluation of urban commercial health. We construct a Street Economic Vitality Index (SEVI) and demonstrate its applicability through an empirical study of central Nanjing. We conceptualize street vitality as the result of a dynamic synergy between internal commercial quality and external radiative fields. Specifically, we leverage Vision-Language Models (VLMs) coupled with Large Language Model (LLM) reasoning to extract and refine brand semantics from signboards. This allows for the precise calculation of a Weighted Brand Ratio Index, which differentiates between local and international commercial tiers, effectively pricing the quality heterogeneity of street-level retail.
Crucially, we extend this framework by rigorously modeling Mall Spillover Vitality ($MV_i$) to capture the spatial externalities of large commercial anchors. While foundational studies have significantly advanced commercial agglomeration analysis Sevtsuk2014, Yue2017, they typically operationalize influence through uniform gravity metrics. These approaches homogenize the spatial externalities of diverse commercial entities. By calibrating specific Gaussian decay rates ($\sigma$) contingent on semantic categorization within a defined spatial boundary (e.g., a 2,000-meter localized threshold), we shift from threshold-based to field-based assessment. This dual-model approach enables us to distinguish areas where vitality is driven by structural advantages (field energy) versus those suffering from hidden recession. Furthermore, to overcome the temporal limitations of static visual audits, we incorporate dynamic Location-Based Services (LBS) data to capture the true tidal rhythms of street-level human activity.
Nanjing’s central districts, designated as the primary development axis in the Territorial and Spatial Master Plan (2021--2035) NanjingPlan2024, serve as an ideal case study. The area faces the dual challenge of aging infrastructure and fragmented commercial fabrics. By operationalizing SEVI, we provide a tool that not only measures current vitality but also unveils the latent mechanisms of quality-driven growth and spatial energy essential for precision spatial governance.
This study advances the discourse on urban economic vitality through four main contributions:
Research on urban economic vitality has undergone a paradigm shift from macroscopic statistical analysis to microscopic street-level sensing. In this section, we review the evolution of these measurement methodologies, specifically critiquing current limitations regarding semantic depth (commercial heterogeneity) and spatial modeling (agglomeration externalities), which motivate our methodological interventions.
Early assessments of urban vitality relied heavily on macro-level socio-economic indicators, such as regional GDP and employment density Wang2022_JUM. Although effective for regional comparisons, these metrics lack the granular resolution required for community-level spatial governance. The advent of Big Data introduced meso-scale proxies, such as POI density and LBS positioning data Lan2020, Chen2019. However, these datasets are inherently prone to an observational lag in market exit—they record registered entities but fail to capture the friction of business collapses in real-time, often painting an overly optimistic picture of declining neighborhoods and contributing to the illusion of prosperity.
Recent studies have attempted to resolve this temporal latency by utilizing Street View Imagery (SVI) to capture the physical reality of streets Jiang2022, Zhang2019_SocialSensing, Xu2024, Ling2025. Researchers have demonstrated that physical features—such as sky view factors and greenery—significantly correlate with vitality Long2016_Green. Furthermore, the link between micro-scale walkability and human experience has been emphasized as a key driver for healthy urban living Liao2025, KOO2023. More importantly, recent works have begun to use computer vision to identify attrition indicators, such as closed stores or roller shutters, providing a more realistic net vitality assessment. However, while the visual dimension has been significantly advanced, the semantic dimension of commercial quality remains fundamentally under-explored.
Current SVI-based vitality assessments predominantly operate on a binary logic: distinguishing solely between open and closed storefronts. While detecting closures is a significant methodological stride, it simplifies all active stores into homogenous units. This binary approach ignores the substantial economic heterogeneity and varying resilience within active commerce.
For instance, a street lined with generic, subsistence-level retail shops may exhibit the same nominal “openness ratio” as a street dominated by high-end chain brands. Yet, their economic yield and risk-resistance capabilities are vastly different. Existing studies on storefront interfaces focus largely on physical transparency or signage density Mehta2009, rarely penetrating the semantic content of the signage itself. Crucially, traditional Optical Character Recognition (OCR) methods fall short in this domain. They are often strictly literal, struggling with stylized logos and failing to resolve semantic ambiguities—such as linking multilingual aliases (e.g., mapping localized logograms to global brand names). The lack of semantic reasoning means that traditional models can quantify the mere presence of commerce but cannot evaluate its brand capital or market hierarchy.
To address these semantic bottlenecks, the field has seen a rapid adoption of Vision-Language Models (VLMs) for unstructured visual data processing. Recent systematic reviews confirm that VLMs have emerged as a robust tool for street view analytics and socioeconomic perception VLMReview2025. State-of-the-art frameworks like UrbanVLP have successfully utilized multi-granularity VLM pretraining to predict broad urban indicators UrbanVLP2025, while models such as MINGLE have extended VLM capabilities to detect semantically complex social regions MINGLE2025. However, while these advanced models excel in general scene perception, they lack a dedicated economic mechanism for deciphering the precise brand hierarchy of commercial storefronts and coupling this intangible asset with continuous spatial externalities.
Our study bridges this gap by integrating VLMs with LLM-based semantic rectification. By doing so, we transition vitality assessment from basic visual enumeration to a robust qualitative diagnosis, establishing a systematic integration of hierarchical brand semantics and mall-induced spatial spillovers.
In addition to internal storefront attributes, street vitality is profoundly shaped by exogenous anchors, particularly large-scale shopping malls. Existing literature acknowledges malls as vitality hubs but typically models their influence using simplistic spatial metrics, such as Euclidean distance to the nearest mall or rigid buffer zones Sevtsuk2014, Yue2017.
These geometric approaches assume that commercial influence is uniform within a radius or decays linearly. They fail to capture the true agglomeration externalities of commercial centers—where the intensity of spatial spillover follows a non-linear distance decay law and varies significantly by the anchor's market position (e.g., a massive mixed-use complex radiates much further than a local supermarket). By ignoring these continuous spatial externalities, traditional metrics often underestimate or misjudge the economic potential of streets located in the spillover zones of major commercial centers. This study addresses this limitation by introducing a Gaussian decay function, modeling mall spillover as a continuous, typology-weighted energy field. This approach differentiates attenuation intensities based on fine-grained POI sub-categories, ensuring a realistic representation of commercial gravity and spatial resource distribution.
Building on the visual-semantic framework, we formalize the Street Economic Vitality Index (SEVI) not merely as a statistical aggregation, but as a diagnostic feature matrix that captures the interplay between internal quality and external field effects. This section operationalizes the three dimensions—Commercial Activity, Spatial Utilization, and Physical Environment—into computable vectors, explicitly incorporating attrition filters (closure detection) and field-based interactions (spillover modeling) to provide the independent variables ($X$) for our subsequent spatiotemporal dynamic analysis.
To ensure cross-indicator comparability, all raw variables are first directionally aligned (e.g., transforming the Closure Ratio into an openness proxy) and standardized using Min-Max normalization. For street segment $i$, we define the standardized indicator vector as:
Specifically, the components correspond to:
These indicators are structured into three dimensions using a block-diagonal weight matrix, ensuring that indicators within each dimension are aggregated independently before final synthesis:
where $M_A \in \mathbb{R}^{1\times 4}$, $M_U \in \mathbb{R}^{1\times 3}$, and $M_P \in \mathbb{R}^{1\times 2}$ are the row vectors representing the weights calculated via the Entropy Weight Method (EWM) Zhu2020_EWM, employed to avoid subjective bias and objectively capture the spatial information variance. Applying $M$ to $\mathbf{x}_i$ yields the three-dimensional feature vector $\mathbf{z}_i = [A_i, U_i, P_i]^{\mathsf{T}}$.
Finally, to generate a comprehensive diagnostic score for static spatial quality, the composite SEVI is derived via the TOPSIS method Hwang1981. This ranks street segments by their relative proximity to the Ideal State ($\mathbf{z}^+$) and Negative State ($\mathbf{z}^-$):
This formulation ensures that the nine-dimensional diagnostic feature matrix accurately captures the built environment's vitality mechanisms, which are then individually regressed against real-time LBS crowd intensity to decode the temporal drivers of urban life. Meanwhile, the composite SEVI is utilized for macro-scale spatial diagnosis.
The formulation rests on specific theoretical assumptions regarding urban mechanisms:
We established a comprehensive analytical framework to diagnose the economic vitality of urban streets. Moving beyond traditional single-source metrics, this study constructs a multimodal data fusion pipeline that integrates high-resolution Street View Imagery (SVI) with semantic geospatial data. This section details the study area, the visual-semantic interpretation pipeline, and the operationalization of vitality indicators.
The study focuses on the central metropolitan area of Nanjing, Jiangsu Province (Figure (ref)). As a city with over 2,500 years of history, Nanjing exemplifies the typical challenges of modern urban renewal: a complex mosaic of historical preservation zones, aging residential neighborhoods, and high-density commercial cores. While the master plan identifies this area as the primary axis for vitality, it faces significant spatial fragmentation NanjingPlan2024. The coexistence of thriving commercial hubs and declining, hollowing-out streets makes it an ideal testbed for our micro-scale diagnostic framework. Validating SEVI in such a heterogeneous environment ensures its transferability to other dense urban contexts facing similar regeneration pressures.
To capture the fine-grained texture of street life, we implemented a systematic sampling and acquisition workflow (Figure (ref)). We first constructed the foundational road network by extracting pedestrian-accessible segments from OpenStreetMap, excluding highways and tunnels to yield 7,153 valid segments. To ensure a continuous visual narrative, we adopted a high-frequency sampling strategy at 20-meter intervals—a significant resolution upgrade compared to the sparse 50m or 100m sampling intervals commonly used in previous studies Li2022, Jiang2022. This density is critical for capturing micro-variations in storefront continuity and detecting localized dead spots that coarser sampling might miss. Consequently, using the Baidu Maps API, we retrieved 557,672 panoramic images (1024 $\times$ 800 pixels) across 278,836 unique points, with each point capturing dual-directional views ($90^\circ$ and $270^\circ$ relative to the road heading) to cover the street interface comprehensively.
Crucially, to avoid simultaneity bias and explicitly introduce causal awareness, we implemented a strict temporal lag design between the built environment features and vitality outcomes. The explanatory variables ($X_{i, t-1}$) are derived from a synthesized historical baseline comprising the SVI data (captured in or before 2022) and finely curated Point of Interest (POI) data from 2023. This integration is methodologically intentional: SVI effectively records "slow variables" such as building facades, structural closures, and green infrastructure, while the 2023 POI data provides a necessary semantic update to capture the "fast variables" of commercial turnover. Together, they constitute the accumulated physical and commercial baseline of the street at time $t-1$.
Furthermore, to overcome the snapshot limitation inherent in static visual audits, this study incorporates dynamic Location-Based Services (LBS) data from October 2025 as the primary proxy for actual street vitality (the dependent variable, $Y_{i, t}$). This natural temporal lag ($t-1$ preceding $t$) logically isolates the pre-existing spatial supply from subsequent human behavior, ensuring that our model captures the quasi-causal driving effects of historical street morphologies rather than purely descriptive co-occurrences. Sourced from mobile signaling and location heatmaps, the LBS dataset captures real-time active user volumes (UV), representing the revealed preference and realized spatial demand of urban residents. To map the spatiotemporal rhythms of urban life, the raw data was temporally aggregated into four distinct tidal periods: Morning Peak (07:00--09:00), Midday (11:00--14:00), Evening Peak (17:00--19:00), and Night Economy (20:00--22:00), covering both a typical weekday (October 15) and a weekend (October 19). Spatially, the raw positioning data was cleaned to remove signal drift noise and spatially joined to the 7,153 street segments. As illustrated in Figure (ref), the dynamic crowd intensity exhibits profound spatiotemporal heterogeneity, revealing a distinct tidal pattern that necessitates a time-sliced modeling approach to decouple the underlying spatial drivers across different periods.
A core innovation of this study is the shift from visual detection to semantic understanding. We developed a two-stage deep learning pipeline to extract both the physical status and semantic quality of street entities.
To quantify indicators, we defined a taxonomy of visual markers specifically relevant to urban renewal (Table (ref)) and employed YOLOv5-seg Redmon2016_YOLO for instance segmentation. The model was trained to identify and segment three distinct classes: signboards, glass interfaces, and closed stores. Unlike simple object detection, instance segmentation generates pixel-level masks, enabling the precise separation of these elements from complex backgrounds. These segmented markers serve specific downstream analytical functions: signboards are quantified to measure Shop Density, while glass interfaces are used to calculate Shopfront Glazing Density. Crucially, the joint detection of closed stores and signboards allows for the computation of the Closure Ratio, thereby capturing the negative commercial dynamics often missed by traditional POI data. Furthermore, the extracted signboard regions serve as the direct input for our subsequent VLM-LLM Brand Decoding Pipeline.
The model was trained on a manually annotated dataset of 642 images using the LabelMe tool Wada2021 (parameters in Table (ref)). As illustrated in the training performance curves (Figure (ref)), the model exhibited stable convergence characteristics. Ultimately, the detection of Signboards achieved high precision ($[email removed] = 0.774$), providing a reliable foundation for the subsequent semantic analysis.
To address the limitation that all active stores look alike in traditional metrics, we developed a Visual-Language Model (VLM) driven Brand Decoding module. As illustrated in Figure (ref), the pipeline leverages a dual-stage recognition-rectification framework to process street-view imagery:
(1) Stage 1: VLM Extraction (S1). We employ a pre-trained VLM (Qwen2-VL-7B-Instruct) to transcribe textual information and identify commercial logos from the input images. To minimize model hallucination and ensure deterministic outputs, generation hyperparameters were strictly constrained (temperature = 0.0, top_p = 1.0, max_new_tokens = 64.0, and do_sample = False). The VLM is guided by a carefully engineered zero-shot prompt instructing it to act as a street-view signage recognition assistant. To explicitly suppress hallucination, the prompt enforces strict constraints (i.e., extracting only visible text/logos without fabrication). The output is mandated as a structured JSON object containing brands_found and a descriptive \texttt{summary}. If no clear brands are present, it returns an empty list, effectively filtering out environmental noise.
(2) Stage 2: LLM Rectification & Classification (S2). The raw, potentially noisy text list extracted from S1 is fed into an LLM agent powered by DeepSeek-Chat (temperature = 0.0, top_p = 1.0, max_new_tokens = 128.0, do_sample = False). To standardize heterogeneous inputs, the LLM acts as a professional commercial classification assistant, bridging local observations with global semantic hierarchies. The prompt dynamically injects both a customized reference database and the raw VLM outputs, enforcing a strict reasoning hierarchy to classify brands into International Brand, Local Brand, or Ordinary Brand tiers. The final output is strictly constrained to a key-value JSON format, ensuring seamless programmatic integration for the subsequent calculation of the Weighted Brand Ratio ($BR_i$). Finally, to mitigate spatial discontinuity inherent in discrete sampling points, we apply a spatial sliding window (size=5) to smooth the $BR_i$ index, ensuring it reflects the continuous commercial atmosphere.
To validate the efficacy of the proposed dual-stage visual-semantic pipeline and quantify the necessity of LLM rectification, we conducted an ablation study. We constructed a manually annotated ground truth (GT) dataset comprising 200 street view images, categorized into 99 International Brands, 54 Ordinary Brands, and 47 Local Brands. We evaluated three progressive pipeline configurations: Traditional OCR (EasyOCR + RapidFuzz), VLM Only (Qwen2-VL), and VLM+LLM (Ours).
As presented in Table (ref), the traditional OCR approach struggles significantly with the stylized fonts, occlusions, and complex lighting inherent in streetscapes, yielding an overall $F1$-score of only 0.412. Upgrading the visual backbone to a VLM markedly improves text recognition ($F1 = 0.652$). Crucially, the introduction of the LLM rectification agent standardizes localized semantic variants and corrects extraction noise, elevating the overall $F1$-score to 0.821. Notably, the VLM+LLM pipeline achieves an $F1$-score of 0.920 for Ordinary Brands and 0.768 for International Brands, representing a 99.2% relative improvement in overall accuracy compared to the OCR baseline. This quantitative leap justifies our claim that high-fidelity brand hierarchy extraction requires deep semantic reasoning beyond primitive OCR, laying a highly accurate data foundation for subsequent spatial vitality modeling.
To complement the storefront analysis, we employed a pre-trained YOLOv8-m model Jocher2023 to detect dynamic mobility flows. The widely-used COCO dataset classes Lin2014 were mapped to our mobility indicators. The model achieved robust detection performance ($[email removed] \approx 0.64$ for mobility classes), ensuring reliable estimation of street-level utilization intensity.
Based on the visual and semantic data extracted above, we operationalized the nine indicators defined in the Formulation section. Table (ref) summarizes the calculation logic. The operationalization of Mall Spillover Vitality ($MV_i$) directly follows the empirical Average Nearest Neighbor (ANN) calibration defined in Section 2, ensuring that the continuous spatial externalities are rigorously constrained within the 2,000-meter threshold. Ultimately, these nine carefully engineered static spatial features will be regressed against the dynamic LBS crowd intensity ($Y$) to decode the temporal heterogeneity of street vitality.
To systematically diagnose the multifaceted nature of street-level economics and explain the spatiotemporal heterogeneity of the dynamic LBS crowd intensity (the dependent variable), we operationalized nine independent spatial features into a tri-dimensional diagnostic matrix (Table (ref)). The statistical distributions of these static indicators at the sampling-point level are visualized in Figure (ref), revealing significant micro-scale spatial heterogeneity across the study area. Unlike traditional metrics that focus solely on volume, our indicator system is designed to capture the nuance among quantity (density), quality (brand and environment), and status (market exit vs. active).
This dimension serves as the core supply-side driver of street vitality; however, we posit that mere establishment density is insufficient to reflect true economic prosperity. To construct a robust assessment, we synthesize four complementary indicators (individual spatial distributions detailed in Figure (ref)). First, Shop Density ($SD_{i}$) measures the absolute concentration of commercial entities per unit length, providing a baseline assessment of commercial agglomeration. Second, acting as a critical attrition filter, the Closure Ratio ($CR_{i}$) visually identifies closed storefronts. This variable explicitly corrects the survivorship bias inherent in traditional POI data, distinguishing between streets that are functionally active and those suffering from economic distress; thus, $1-CR_{i}$ represents the effective functional survival rate. Third, the Weighted Brand Ratio ($BR_{i}$) transcends simple binary enumeration to capture the intangible brand premium of commerce. Derived from the VLM-LLM semantic pipeline, this metric assigns differential weights to international chains and local brands, effectively distinguishing modernized commercial hubs from generic, subsistence-level retail clustering. Finally, Mall Spillover Vitality ($MV_{i}$) quantifies the continuous spatial externalities of large commercial anchors. As detailed in the methodology, this category-weighted Gaussian field captures how large complexes export commercial gravity to surrounding streets. The aggregated spatial distribution of this dimension (Figure (ref)) clearly highlights the structural dependency of local street vitality on regional hubs, forming continuous high-activity corridors in the urban core.
Urban vitality is sustained by the continuous flow of economic agents. We measure the dynamic pulse of the street through three mobility indicators. Pedestrian Presence ($PP_{i}$) reflects the intensity of human-scale engagement and serves as the most direct proxy for revealed consumer demand, echoing classic theories of life between buildings Gehl1987. Complementing this, Motor Vehicle Density ($MD_{i}$) and Non-motor Vehicle Density ($ND_{i}$) capture spatial accessibility and transit-oriented activity. These indicators reveal the critical trade-offs between motorized transit efficiency and pedestrian-friendly vibrancy—a spatial friction often observed in the design of complete streets. Mapping these flows (Figure (ref)) aids in identifying spatial bottlenecks where commercial supply mismatches mobility demand.
The built environment determines the comfort, amenity value, and interaction potential of the street canyon. We focus on two key interface qualities. Shopfront Glazing Density ($GD_{i}$) measures the visual permeability of the interface. High transparency facilitates visual interaction between indoor commercial supply and outdoor pedestrian demand, effectively reducing information asymmetry—a key design principle aligned with the hedonic value of walkability Jacobs1961, Mehta2009. Meanwhile, the Green Coverage Ratio ($GR_{i}$) reflects the provision of ecological amenities and visual comfort. While occasionally exhibiting a negative spatial correlation with extreme commercial density, environmental amenity remains an indispensable factor for long-term urban resilience and livability Long2016_Green, Ye2019_LUP. Visualizing this dimension (Figure (ref)) highlights the structural disparity between hard commercial streets and soft livable neighborhoods.
A distinguishing feature of this framework is its scalar flexibility, bridging macro-level resource allocation with micro-level spatial intervention Huang2025. At the street-segment level (macro-diagnosis), aggregating scores provides a macroscopic view suitable for district-level master planning and identifying regional corridors of economic decline. Conversely, at the sampling-point level (micro-surgery), preserving data at 20-meter intervals captures fine-grained spatial heterogeneity. This ultra-high resolution is critical for precision urban renewal, enabling policymakers to pinpoint localized market failures (e.g., a 100-meter stretch of closed shops or low-quality facades) within an otherwise healthy economic corridor, thereby guiding targeted spatial investments.
Spearman rank correlation analysis unveils the internal coupling mechanisms among street vitality components, as illustrated in Figure (ref).
The analysis first highlights a robust Traffic-Commerce Nexus, where Motor Vehicle Density ($MD_i$) and Shop Density ($SD_i$) exhibit a strong positive correlation ($r=0.79$). This empirical evidence corroborates the classic urban theory that physical accessibility is a fundamental prerequisite for commercial agglomeration Sevtsuk2014. In parallel, the data confirms a significant Crowd-Mall Synergy, evidenced by the correlation ($r=0.62$) between Pedestrian Presence ($PP_i$) and Mall Spillover Vitality ($MV_i$). This finding strongly validates the Gaussian field assumption employed in our model, suggesting that large commercial anchors actively generate pedestrian flows that diffuse into the surrounding street network.
Crucially, the introduction of the Weighted Brand Ratio ($BR_i$) reveals the mechanics of quality agglomeration and provides robust validation for our semantic pipeline. To validate the economic interpretability of this VLM-derived indicator, we first examine its relationship with human mobility metrics. In street-level urban economics, sustained pedestrian footfall is a direct manifestation of revealed consumer demand. As demonstrated in Figure (ref), $BR_i$ exhibits strong positive correlations with Pedestrian Presence ($r = 0.59$) and Motor Vehicle Density ($r = 0.62$), indicating that segments with premium brand concentrations consistently draw higher consumer traffic.
To further cross-validate these findings with objective ground-truth data, we benchmarked $BR_i$ against an independent, exhaustive POI dataset. As detailed in Appendix Table (ref), our results reveal a striking monotonic correspondence between visual-semantic quality and physical business density. Street segments identified as "High-tier Brand Corridors" by our VLM-LLM pipeline exhibit a mean POI density 70.7% higher (12.94) and a premium/discretionary amenity count 73.1% higher (1.61) than "Low-tier" segments (7.58 and 0.93, respectively). This step-wise alignment with retail location theory provides compelling evidence that $BR_i$ is not merely a descriptive textual label, but a highly effective, empirically validated proxy for street-level economic quality and consumption tier.
Finally, and perhaps most intriguingly, the analysis exposes a Density-Competition Paradox. We observed a moderate positive correlation between the Closure Ratio ($CR_i$) and both Shop Density ($r=0.58$) and Brand Presence ($r=0.52$). This counter-intuitive finding reveals that high-density, high-quality commercial zones are not intrinsically stable; instead, they often experience intense competition and rapid tenant turnover. By detecting this correlation, our model successfully captures the “high metabolic rate” characteristic of prime locations—a phenomenon completely invisible to traditional POI-based metrics that would erroneously interpret high nominal density as absolute stability.
Principal Component Analysis (PCA) was employed to further decompose the indicator system into orthogonal latent factors. The dominant components collectively explain 72.9% of the total variance (Figure (ref)), validating the structural independence of our Diagnostic Matrix. To ensure clarity and focus on the primary economic drivers, Table (ref) presents the varimax-rotated factor loadings for the first four principal components.
The first component (PC1), accounting for 41.2% of the variance, functions as the Commercial-Mobility Engine. Loading heavily on Shop Density, Motor Vehicle Density, Pedestrian Presence, and Weighted Brand Ratio, this component embodies the traditional definition of urban prosperity driven by agglomeration, accessibility, and brand quality. It represents the magnitude of street vitality.
In contrast, the second component (PC2, 11.4%) emerges as the Ecological-Interface Trade-off Factor. Dominated by Green Coverage ($GR_i$, loading 0.745), PC2 stands in negative opposition to Shopfront Glazing Density (loading -0.375) and Shop Density (loading -0.290). This structural opposition reveals a fundamental dilemma within the historic core: streets with high ecological comfort often sacrifice commercial transparency and density. This quantitative evidence highlights a central challenge for urban renewal: how to introduce ecological interventions without dampening the visual permeability of established commercial interfaces Ye2019_LUP, Long2016_Green.
Perhaps the most novel insight stems from the third component (PC3, 10.6%), identified here as the Structural Recession Factor. Driven almost exclusively by the Closure Ratio ($CR_i$, loading 0.971), the emergence of decline as an orthogonal statistical dimension—distinct from low density or low brand quality—demonstrates that a dying street is fundamentally different from a quiet one. Empirically, a street can be large and dense (High PC1) yet structurally failing (High PC3). This finding confirms that without the explicit visual detection of store closures, any vitality assessment remains dimensionally incomplete.
Synthesizing the multi-dimensional indicators via the EWM-TOPSIS framework, we generated a spatial distribution map of SEVI to perform a comprehensive diagnosis of Nanjing’s central area (Figure (ref)). The results uncover a distinct Core-Periphery structure, nuanced by significant localized heterogeneity.
Validating the Field Effect, the Xinjiekou district—the city's commercial core—exhibits the highest SEVI scores, forming a continuous high-vitality cluster. Crucially, the vitality scores do not terminate abruptly at the boundaries of large commercial complexes; rather, they decay gradually along the connecting arteries. This observed gradient closely fits our typology-weighted Gaussian attenuation model, empirically demonstrating that the field effect of top-tier malls successfully activates the micro-circulation of the surrounding block network, physically manifesting the Crowd-Mall Synergy identified in our correlation analysis.
Beyond confirming the core, the model proves critical in distinguishing between Healthy Density and Hollow Density. Specifically, certain street segments in the historic southern districts register high traditional POI counts yet receive low SEVI scores. This discrepancy arises from the visual-semantic pipeline's dual capability: it simultaneously detects high Structural Recession (via Closure Ratio) and low Economic Quality (via Weighted Brand Ratio). By identifying streets that are nominally dense but dominated by low-tier generic shops or vacancies, the framework effectively filters out false positives in traditional data. Such diagnostic precision enables planners to prioritize these specific at-risk segments for urgent intervention, rather than misallocating resources to stable, albeit quieter, residential streets.
Furthermore, a spatial mismatch—or Green-Vitality Gap—is evident, where high-vitality zones (Red) largely coincide with low-greenery areas. This spatial pattern aligns perfectly with the Ecological-Interface Trade-off identified in PC2, highlighting the scarcity of Garden Streets in the dense core. This points to a bifurcated renewal strategy: environmental micro-intervention is prioritized for Red zones, while functional activation is required for Green/Yellow zones. Ultimately, the SEVI map functions not merely as a status visualization, but as a Computed Tomography (CT) scan of the urban tissue, revealing the hidden lesions (closures) and energy flows (spillovers) that drive the city's evolution.
While global correlation and PCA models unveil the general structural logic of street vitality, they rely on a static premise, obscuring the dynamic interactions between the built environment and human activity throughout the day. To address this, we operationalized a time-lagged Time-Sliced Geographically Weighted Regression (GWR). By assigning the static built environment features (derived from SVI and POI) to the historical baseline period ($t-1$) and the dynamic LBS crowd intensity to the subsequent observation period ($t$), we formulated the empirical model as:
where $(u_i, v_i)$ represents the geographic coordinates of street segment $i$. This temporal lag structure allows us to explicitly evaluate the quasi-causal, lagged driving effects of historical street morphologies ($X_{i,t-1}$) on real-time pedestrian presence ($Y_{i,t}$), stepping beyond purely descriptive co-occurrences. Overall, this lagged GWR model demonstrates robust predictive power, achieving an average Adjusted $R^2$ of 0.66 across the observed temporal spectrum.
The temporal evolution of the models' explanatory power (Adjusted $R^2$) reveals a profound tidal effect dictated by travel purposes (Figure (ref)). During the morning peak (07:00--09:00), the model's explanatory power drops to its lowest point (Adjusted $R^2 \approx 0.57-0.59$). This suggests that morning commuting is characterized by rigid mobility; pedestrians prioritize efficiency and shortest-path routing, rendering the historical micro-scale commercial baseline ($t-1$) largely ineffective in driving subsequent foot traffic ($t$). Conversely, the model achieves its highest predictive power during the midday (11:00--14:00) and evening (17:00--19:00) periods (Adjusted $R^2$ peaking at 0.71). During these windows of discretionary activity, pedestrians have the flexibility to engage in dining and leisure, allowing the high-quality commercial interfaces of the streetscape to exert a strong, quasi-causal attractive effect on their spatial choices.
Beyond global performance, analyzing the localized coefficients of specific variables uncovers the precise quasi-causal mechanics of commercial spillovers and urban decay (Figure (ref)). The lagged positive influence of Mall Spillover Vitality (top panel) demonstrates significant temporal elasticity. During morning hours, the boxplot is compressed near zero, indicating minimal radiative power. However, during the midday and night economy periods, the interquartile range expands upward, confirming that shopping malls act as time-activated energy pumps, quasi-causally inducing crowd aggregation and actively exporting pedestrian flows to the surrounding 2,000-meter street network primarily during non-commuting hours. Furthermore, this spillover effect exhibits strong long-tail resilience during the weekend night economy compared to weekdays.
Equally critical is the temporal dynamic of the structural recession factor, captured by the Closure Ratio (bottom panel). While the median coefficient remains relatively stable, the lower whisker and negative outliers extend dramatically during the evening and night periods. This provides quantitative validation for Jacobs' eyes on the street theory; while closed storefronts may simply represent inactive facades during the day, at night, continuous roller shutters drastically reduce street illumination and perceived safety, exerting a severe, lagged repulsion effect that actively drives pedestrians away.
To ensure that our findings are not artifacts of the selected spatial constraints, we performed a series of robustness checks by re-estimating the time-lagged GWR model with alternative parameters. First, we varied the maximum spatial threshold ($D$) for the Mall Spillover Vitality ($MV_i$) from the baseline 2,000 meters to 1,000 meters and 3,000 meters. As detailed in Appendix Table A1, the model's explanatory power (Adjusted $R^2$) demonstrates remarkable stability across all thresholds. Crucially, the fundamental spatiotemporal "tidal pattern"—characterized by lowest predictive power during morning commutes and peaking during midday and evening discretionary periods—remains completely intact, confirming the robustness of the dynamic decoupling results.
Furthermore, an examination of the localized GWR coefficients corroborates the stability of our mechanistic findings. As illustrated in Appendix Figure A1, the distribution of the Mall Spillover Vitality coefficients remains structurally consistent across the 1000m, 2000m, and 3000m thresholds. In all scenarios, the median coefficients stay robustly positive, confirming the enduring positive spatial externalities of commercial anchors. More importantly, the pronounced temporal elasticity—characterized by compressed radiative power during morning commutes and significant upward expansion during midday and evening discretionary periods—is perfectly preserved regardless of the spatial constraints applied. The slight upward shift in the interquartile ranges at the 3000m threshold realistically reflects the capture of broader, long-tail radiative fields generated by top-tier mega-malls, further validating our continuous field-based assumption over rigid binary boundaries.
In addition, to address concerns regarding the functional form of spatial attenuation, we replaced the baseline Gaussian decay function with alternative Exponential and Linear decay algorithms while keeping the baseline threshold constant. As shown in Appendix Table A2, the Adjusted $R^2$ values remain remarkably stable across all three mathematical specifications. The preservation of the core spatiotemporal tidal pattern confirms that the quantified mall spillover vitality reflects an objective urban economic mechanism, rather than a mathematical artifact of a specific distance-decay formula.
Finally, to verify that the spatial diagnosis of street vitality is not mechanically driven by the chosen EWM-TOPSIS algorithm, we constructed two alternative composite indices using the Equal-Weight method ($SEVI_{EQ}$) and Principal Component Analysis ($SEVI_{PCA}$). As detailed in Appendix Table (ref), correlation analysis reveals that the original $SEVI$ is highly consistent with both alternative indices, yielding phenomenal Spearman's rank correlation coefficients of $\rho = 0.994$ for $SEVI_{EQ}$ and $\rho = 0.977$ for $SEVI_{PCA}$ (all $p < 0.001$). This confirms that our multi-dimensional spatial diagnostic framework inherently captures the true structural variations of urban vitality, completely independent of the mathematical aggregation techniques employed.
This study addresses critical methodological gaps in urban vitality assessment—specifically, the semantic blindness regarding commercial quality and the spatial oversimplification of agglomeration externalities. By shifting the analytical paradigm from static POI enumeration to a visual-semantic and spatiotemporal framework, we developed the Street Economic Vitality Index (SEVI) to decode street-level economics. In doing so, we directly resolve the central economic puzzle introduced at the outset: unmasking the “illusion of prosperity” and providing a rigorous micro-evidence base to correct the misallocation of spatial resources during urban renewal.
Our empirical application in the complex urban tissue of Nanjing yields four fundamental insights that refine existing urban economic theories.
First, by transitioning from binary existence to semantic quality, we effectively corrected the observational lag in market exit. The introduction of the Closure Ratio unveiled a distinct class of pseudo-dense streets that appear thriving in nominal databases but are suffering from hidden physical recession. Crucially, the integration of the Weighted Brand Ratio—powered by our VLM-LLM pipeline—allowed us to quantify the intangible brand premium of street interfaces for the first time. We demonstrated that true economic vitality is defined not merely by the absolute density of establishments, but by the weighted agglomeration of high-tier commercial capital.
Second, our results structurally validate the spatial externalities of commercial agglomeration. The category-weighted Gaussian spillover model successfully quantified the invisible energy field of commercial anchors, proving empirically that top-tier malls act as vitality engines with predictable, non-linear spatial decay rather than uniform buffer zones.
Third, we revealed the profound temporal heterogeneity of vitality mechanisms. By deploying a Time-Sliced GWR against dynamic LBS data, we demonstrated that the attractive power of commercial semantics and mall spillovers is highly elastic, peaking during periods of discretionary midday and evening activity. Conversely, structural recession factors, such as shop closures, severely dampen vitality specifically during the night economy, quantitatively validating the nocturnal importance of Jacobs' “eyes on the street” theory regarding perceived safety and spatial friction.
Finally, a structural tension was uncovered between commercial intensity and environmental amenity. This quantifiable prosperity-ecology trade-off highlights the inherent market failure in achieving garden-city qualities within high-density commercial cores without deliberate public intervention.
Beyond methodological innovation, SEVI serves as a practical diagnostic tool for precision urban governance. Unlike regional macro-metrics, our dual-scale framework enables surgical, evidence-based spatial interventions. For policymakers, the implications dictate a shift toward targeted resource allocation: high-vitality zones facing ecological deficits require targeted public goods provision (e.g., environmental micro-regeneration) to mitigate the negative externalities of extreme density. Conversely, zones characterized by hollow density require functional replacement and brand upgrading strategies rather than redundant physical infrastructure investments. Furthermore, recognizing temporal vulnerabilities allows for temporally differentiated governance; for instance, areas suffering from high nocturnal closure rates require targeted nighttime lighting and pop-up activations to offset the severe drop in pedestrian flow.
While this study successfully overcomes the daily snapshot limitation by integrating tidal LBS data to capture intra-day dynamics, a notable limitation remains regarding the temporal mismatch between our multi-source datasets. Due to the updating constraints of large-scale street view platforms, the SVI dataset comprises static visual slices captured in or before 2022. Although we strategically integrated this with 2023 POI data to construct a comprehensive historical baseline ($t-1$), this macro-level approach inevitably introduces a temporal resolution mismatch when regressed against the highly dynamic 2025 LBS outcomes ($t$). Specifically, this static visual baseline may fail to capture the immediate impacts of short-term, localized street micro-renewals (e.g., rapid retail turnover or sudden facade upgrades) that could instantly stimulate contemporary pedestrian flows. Future research should integrate longitudinal, time-series SVI datasets to strictly align visual variables with mobility outcomes. This would not only resolve the temporal mismatch but also enable more rigorous panel data causal inference, effectively distinguishing between short-term volatile fluctuations and permanent inter-annual economic decline. Furthermore, enriching the semantic layer by incorporating multi-sensory data (e.g., soundscapes, social media sentiment) could further deepen the human-centric resolution of the model Xu2024.
In summary, by endowing machines with the ability to decode the semantic hierarchy and physical attrition of the street, and coupling these insights with dynamic human mobility, this research steers urban renewal away from blind, aggregate expansion toward quality-oriented, precision spatial governance.