EconBase
← Back to paper

Mapping Firms' Locations in Technological Space: A Topological Analysis of Patent Statistics

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

78,227 characters · 17 sections · 22 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Mapping Firms' Locations in Technological Space: A Topological Analysis of Patent Statistics

\end{tabular} \hskip 0.7em \@plus.17fil\relax

tabular[tabular omitted — 201 chars of source]

\hskip 0.7em \@plus.17fil\relax

tabular[tabular omitted — 101 chars of source]

\hskip 0.7em \@plus.17fil\relax

tabular[tabular omitted — 1,269 chars of source]

Introduction

The “rate and direction of inventive activity” have been recognized as one of the main themes in economics since at least the conference of the same title in 1960 (Nelson1962, LernerStern2012). Whereas the rate of innovation has been studied extensively, research on its direction has seen much less progress. Nevertheless, recent studies suggest the direction of scientific change is both an important choice for individual researchers and a critical outcome for scientific communities (AzoulayEtAl2019, Myers2020). These observations, along with the central role of product differentiation in the theory of industrial organization (IO), suggest the direction of inventive activity is important for firms and industries as well.

Mapping the locations and directions of firms’ research and development (R&D) activities is a challenging problem because technological space has many dimensions, unlike physical/geographical space.\footnote{Whereas a large literature exists on the geography of innovation (pioneered by JaffeTrajtenbergHenderson1993), relatively few papers explore technological space, because of methodological challenges.} Even a relatively “coarse” classification system by the US Patent and Trademark Office (USPTO) uses more than 400 categories (patent classes), and large firms frequently conduct R&D in more than 100 classes, obtaining thousands of patents each year. As a result, the dimensionality of the action/state space is extremely high, and infinitely many directions of inventive activity are possible in principle. Studying something we cannot even visualize and describe is difficult. Hence, developing a method for faithfully mapping their technological positions and documenting empirical regularities (i.e., measurement and exploratory data analysis) would be a crucial step.

Given the high dimensionality of the problem, some dimensionality reduction seems warranted. Commonly used methods include principal component analysis (PCA), multi-dimensional scaling (MDS), and various algorithms for clustering (e.g., k-means clustering). However, even though these existing methods provide some simplified visualization and description, fundamental issues remain unresolved: collapsing data would eliminate useful information about the direction of inventive activity. For example, Figure (ref) (a) shows a PCA that projects onto a two-dimensional plane 333 major firms’ patent portfolios (vectors of logged patent counts across 430 USPTO classes) in 1976--2005. Huge clusters of points on the left side would seem to suggest many firms conduct R&D in close proximity, but this “densely populated area” could partly be an artifact of collapsing the other 428 dimensions. Similar issues arise in other existing methods, due to information loss (see section 4.4 for an example of clustering). Thus, a faithful representation of the positions and directions of R&D requires new descriptive tools that avoid arbitrarily collapsing data, provide intuitive visualizations of how firms’ patent portfolios evolve over time, and permit quantification of these dynamics.

figure[figure omitted — 1,010 chars of source]

This paper presents such a new method to represent firms’ locations as a combinatorial/topological object (shape graph), which can be easily visualized and quantified in a variety of ways using graph theory. We adapt and extend a tool from computational topology called the Mapper procedure singh2007topological. This algorithm is well founded on mathematical concepts from computational topology and geometry, such as the Reeb graph, and aims to preserve the topological and geometric information of the original data, in two steps. First, it clusters data points in each local neighborhood based on a distance metric of one’s choice (e.g., cosine distance). Second, it connects clusters with edges if a pair of clusters shares at least one data point. Hence, even though the resulting graph might appear to visualize data on a two-dimensional plane--—see Figure (ref) (b)---as in the PCA plot, the shape graph retains the notions of proximity and continuity (in the original space) with edges between neighboring nodes.

We apply this method to the dynamic evolution of the 333 major firms’ patent portfolios across 430 USPTO classes in 1976--2005, and report three sets of results. First, we visualize these firms' technological positions and trajectories over the three decades. (Whereas “data visualization” plays only a minor role in most empirical studies, it embodies one of the main results in our context, because the systematic mapping of technological space is the central empirical problem that this paper addresses.) We find many engineering firms remain undifferentiated and cluster together in the densely populated “trunk” or the “continental” part of the map. However, a few dozen firms, primarily in the information technology (IT) sector, start differentiating from the rest in the 1980s and the 1990s, developing unique portfolios and exhibiting distinctive trajectories, as represented by long “branches” or “flares” that spike out of the main trunk. In the topological space, which is coordinate free, these shapes provide explicit signatures of the unique “directions” of inventive activity.

Second, we propose a formal definition of such flares based on graph theory, as well as a computational method to measure their length, and find 40.3 % of the firms exhibit some flares. We assess the empirical relevance of this new measure by evaluating its statistical relationships with the firms’ financial performances (revenue, profit, and market value). Regression results suggest positive correlations between the flare length and the performance metrics. This association is statistically significant at conventional levels, and economically significant in magnitude (e.g., an extra length of flare in 1976--2005 is associated with 31%--40% higher performances as of 2005). Moreover, these patterns continue to hold after controlling for (i) portfolio size, (ii) firm survivorship, (iii) industry classification, and (iv) firm fixed effects.

Third, we show how our method and results compare with JAFFE198987, which is based on k-means clustering and is one of the most prominent methods to study firms’ technological locations. The scope of Jaffe’s clustering is global, which makes it suitable for splitting firms into industries. But JAFFE198987 struggles to track firm-level trajectories and fails to find any statistically significant relationship between their moves and performances. By contrast, our scope of clustering is only local, which allows us to preserve details at the firm-year level. Moreover, the whole procedure is designed to retain and recover the continuum of firms and industries in the original data, and allows us to characterize firm-level trajectories. Our discovery of statistically significant relationships between the firms' financial performances and their length of unique technological trajectories (flares) demonstrates the benefit of this high-precision approach.

Thus, our approach is complementary to the existing methods and can generate new insights that are difficult to obtain otherwise. It helps us answer some of the most basic questions, including where firms innovate, how their technological trajectories are related to their product-market performances, and how industries and technologies evolve over time.

We organize the rest of the paper as follows. Section 2 presents a model of competition and innovation in a high-dimensional space. Section 3 explains the data. Section 4 introduces our topological method and presents a historical map of firms' inventive activities. Section 5 explains our method to measure flare length and assesses its correlation with firms' performances. Section 6 concludes. The Online Appendix contains (A) the details of our economic model, (B) raw-data patterns, (C) an introduction to TDA, and formal definitions and proofs, (D) the details of Jaffe-style clustering, (E) sensitivity analysis, (F) panel-data regressions and out-of-sample predictions, (G) comparison with network-centrality measures and Jaffe's distance measure, and (H) additional exhibits.

Conceptual Framework

We propose an economic model of firms' competition and innovation to (i) highlight key economic forces that affect firms' behaviors and market outcomes, (ii) guide our exploratory data analysis, and (iii) facilitate the interpretation of our empirical findings.

Competition and Innovation in High-Dimensional Space

We combine elements of the workhorse IO models of BLP1995 (BLP) and EP1995 (EP) in the presence of many product markets that are embedded in the space of technologies.

\paragraph{Markets and Technologies.}

Consider many product markets indexed by $m=1,2,...,\left \vert \mathcal{M }\right \vert $, each of which is populated by $M_{m,t}$ consumers and $ N_{m,t} $ firms in period $t$. They are independent of each other. Their main difference from geographical markets---whose physical locations can be characterized by only two numbers, longitude and latitude---is that we characterize their \textquotedblleft locations\textquotedblright \ from the viewpoint of technologies that are required to serve them. Let $l\left( m\right) \equiv \left( l_{1}\left( m\right) ,l_{2}\left( m\right) ,...,l_{K}\left( m\right) \right) $ denote the location of market $m$ in the $K$-dimensional space, where $l_{k}\left( m\right) \geq 0$ is its $k$th coordinate.\footnote{ We abstract from the distinction between product space and technology space because we use only patent statistics and financial data in our empirical analysis. See BloomVanReenenSchankerman2013 for an example that makes this distinction.}

\paragraph{Period Profit.}

Each of the $N_{m,t}$ firms earns period profit,

equation[equation omitted — 111 chars of source]

where $\xi _{i,t}$ is product quality (we assume single-product firms) and $ c_{i,t}$ is constant marginal cost of production. This reduced-form profit function encapsulates a BLP-style model of a differentiated-product demand system and Bertrand competition (see Appendix A.1). Hence, $\pi _{i,t}$ is increasing in $M_{m,t}$ and $\xi _{i,t}$ but decreasing in $N_{m,t}$ and $ c_{i,t}$. These four objects are determined by the history of (all) firms' actions, $h_{t}\equiv \left( h_{i,t}\right) _{i=1}^{N_{t}}$, where $ h_{i,t}\equiv \left( a_{i,\tau }\right) _{\tau =0}^{t-1}$ is firm $i$'s actions up to period $t-1$, and $N_{t}$ denotes the total number of firms that have operated in any of the $\left \vert \mathcal{M}\right \vert $ markets in any period up to $t$.

\paragraph{Market Size.}

Each market $m$'s size is realized at $t=0$ following some distribution $ F_{M}$ with spatial correlations, $M_{m,0}\sim F_{M}$. In any subsequent period $t>0$, its effective size is the portion of consumers that have not purchased anything yet,

equation[equation omitted — 148 chars of source]

where $\mathcal{H}_{m.t-1}$ is the set of remaining consumers in market $m$ at $t-1$, $\mathbb{I}\left \{ \cdot \right \} $\ is an indicator function, $ d_{h,t-1}$ is the discrete choice of consumer $h$ at $t-1$, and $ d_{h,t-1}\neq 0$ means the consumer bought something.

\paragraph{Number of Firms.}

The number of active firms in market $m$ at time $t$ is the sum of firms whose technological locations $l_{i,t}\equiv \left( l_{i,t,1},l_{i,t,2},...,l_{i,t,K}\right) $\ are in the neighborhood of $ l\left( m\right) $:

equation[equation omitted — 158 chars of source]

where $\mathcal{N}\left( \cdot \right) $ is the set of neighborhood locations (specified in section 4). Thus, firms can serve market $m$ only when they possess \textquotedblleft relevant\textquotedblright \ technologies $l_{i,t}\in \mathcal{N}\left( l\left( m\right) \right)$.

\paragraph{R&D Investments.}

Each firm's location is determined by $l_{i,t}=f^{l}\left( x_{i,t}\right) $, where $f^{l}$ is an increasing function (specified in section 4) and $x_{i,t}\equiv \left( x_{i,t,1},x_{i,t,2},...,x_{i,t,K}\right) $ is the amount of successful R&D investment in each of the $K$ technological areas at time $t$. Not all R&D investments are successful, and firms could be heterogeneous in their R&D productivity. We encapsulate these notions in a stochastic R&D-production function,

equation[equation omitted — 143 chars of source]

where $f^{x}$ is an increasing function of $b_{i,t-1,k}^{x}$ ($i$'s R&D budget in the previous period in area $k$), $\omega _{i,t-1,k}^{x}$\ is its area-specific R&D productivity that follows some exogenous Markov process, and $\varepsilon _{i,t,k}^{x}$ is an i.i.d. shock. Let $b_{i,t}^{x}\equiv \sum_{k=1}^{K}b_{i,t,k}^{x}$ denote the total R&D\ expenditure across all areas, and $\mathbf{\omega}_{i,t}^{x}\equiv \left( \omega_{i,t,k}^{x}\right) _{k=1}^{K}$ the vector of area-specific R&D productivity.

\paragraph{Other Investments.}

Firms can engage in two other categories of investments---marketing and operations---which determine the firm's product quality $\xi _{i,t}$ and production cost $c_{i,t}$, respectively. These state variables evolve according to some controlled\ Markov processes, $\xi _{i,t}=f^{\xi }\left( \xi _{i,t-1},b_{i,t-1}^{\xi };\omega _{i,t-1}^{\xi }\right) $ and $ c_{i,t}=f^{c}\left( c_{i,t-1},b_{i,t-1}^{c};\omega _{i,t-1}^{c}\right) $, where $b_{i,t}^{\xi }$ and $b_{i,t}^{c}$ are $i$'s budgets for marketing and operations, respectively, and $\omega _{i,t}^{\xi }$ and $ \omega _{i,t}^{c}$ are $i$'s productivity in these activities, which follow some exogenous Markov processes as well.

\paragraph{Budget.}

The firm's total budget is constrained by the amount of available cash,

equation[equation omitted — 115 chars of source]

which is determined by the following accounting rule,

equation[equation omitted — 107 chars of source]

where the first three terms on the right-hand side (RHS) reflect cash holding, expenditure, and profits in the previous period, respectively, and $fin_{i,t-1}\gtrless 0$ is the cashflow from financing activities.\footnote{ We assume $fin_{i,t}$ follows some exogenous Markov process and do not model the underlying financial markets. We include it to incorporate the possibility that retained earnings are not the only source of cash and that a firm can go bankrupt (see Appendix A.2 for entry and exit).}

\paragraph{Dynamic Optimization.}

Each firm allocates its budget to R&D $\mathbf{b} _{i,t}^{x}\equiv \left( b_{i,t,k}^{x}\right) _{k=1}^{K}$, marketing $b_{i,t}^{\xi }$, and operations $b_{i,t}^{c}$, to maximize the discounted present value of its current and future profits,

equation[equation omitted — 217 chars of source]

subject to the budget constraint ((ref)). $\beta _{i}\in \left( 0,1\right) $ is $i$'s discount factor. $E_{i,t}$ is the expectation operator given its information set and beliefs at $t$. We do not fully specify these objects because computing equilibria of this dynamic game is outside the scope of this paper, but we intend our framework as a model of the EP class (i.e., strategic industry dynamics with Markov-perfect equilibrium).

Implications for the Analysis of Technological Space

Five features of the model are particularly relevant for the analysis of firms' technologies:

enumerate• {Profit }$\pi _{i,t}$ is increasing in $M_{m,t}$ and $\xi _{i,t}$ but decreasing in $N_{m,t}$ and $c_{i,t}$; • These four objects are determined by the history $h_{t}$ of (all) firms' actions $a_{i,t}$; • Firms are heterogeneous in their productivity, $\omega _{i,t}\equiv \left( \mathbf{\omega}_{i,t}^{x},\omega _{i,t}^{\xi },\omega _{i,t}^{c}\right) $; • The size $M_{m,t}$ of each market is finite and could only decrease over time; and • Current profit $\pi_{i,t}$ could increase future R&D budget $\mathbf{b} _{i,t+1}^{x}$ via ((ref)) and ((ref)).

A direct implication of Features 1 and 2 is that firms would try to operate in markets with high $M_{m,t}$ and low $N_{m,t}$. Thus, the realized profile of locations, $l_{t}\equiv \left( l_{i,t}\right) _{i}$ will reflect firms' tradeoff between \textquotedblleft chasing consumers\textquotedblright \ and \textquotedblleft avoiding competitors.\textquotedblright \ Feature 3 suggests firms with comparative advantage in R&D (i.e., relatively high $\mathbf{\omega}_{i,t}$) would move away from crowded markets and try to carve out their own niches. The high dimensionality $K$ of the technological space, combined with firms' heterogeneous R&D capabilities across $K$ areas, offers ample room for such differentiation. Feature 4 limits the extent to which firms can \textquotedblleft rest on their laurels\textquotedblright \ (i.e., remain profitable in the same locations). Because potential demand in any given market is like an oil reserve that becomes increasingly difficult to extract, firms have to either constantly explore and conquer new markets or keep investing in $\xi _{i,t}$ and $c_{i,t}$ to dig deeper. Finally, Feature 5 highlights the possibility of a virtuous cycle in which “the rich gets richer.” That is, those who succeed in developing unique technologies earn extra profits, which can be reinvested in future innovations to pursue further growth opportunities.

These considerations suggest the locations of firms relative to each other $\left \{ l_{i,t}\right \} $ could exhibit rich variation and contain relevant information about their performances and underlying capabilities. In particular, a string of unique positions occupied by a firm may be indicative of its long track record of successful innovations and sustained profitability. We present our method for describing $\left \{ l_{i,t}\right \} $ in section 4, and formalize the measurement of firms' unique technological trajectories in section 5.

Data

\paragraph{ Patents.} We use Ozcan's (2015) data on patents that are granted by the USPTO between 1976 and 2010.\footnote{Ozcan2015 uses the USPTO's Patent Data Files, which contain raw assignee names at the individual patent level. By contrast, the NBER Patent Data File (another commonly used source of patent data) records standardized assignee names at the “pdpass” (unique firm identifier) level, which is less granular than the original assignee name.} We use their application years (instead of years in which they are granted) in our analysis, because the former is closer than the latter to the time of actual invention. We focus on patents that are applied through 2005, because a substantial fraction of later applications would still be under review as of 2010, which raises concerns about sample selection. We sometimes call these patents “R&D patents” to distinguish them from “M&A patents” (see below).

\paragraph{ Mergers and Acquisitions (M&As).} Aside from conducting in-house R&D and applying for patent protection, firms often obtain patents by acquiring firms that have their own portfolios of patents. Ozcan's (2015) dataset links the USPTO data to the Securities Data Company's M&A data module. This part of the dataset contains M&A deals between 1979 and 2010 in which both the acquiring firm and the target firm have at least one patent between 1976 and 2010.\footnote{The data include merger, acquisition, acquisition of majority interest, acquisition of assets, and acquisition of certain assets, but exclude incomplete deals, rumors, and repurchases. We use data on these transactions through 2005.}

\paragraph{ Financial Performances.} We use Compustat data on the firms' revenues, EBIT (earnings before interest and taxes), and stock-market capitalization in 2005 (or the last available fiscal year if the firm disappears before 2005). Our purpose is to assess the relevance of our topological measures in terms of their correlations with the firms' eventual financial performances (in section 5).

\paragraph{ Descriptive Statistics.} To keep the sample size suitable for visual inspection and detailed exploratory analysis, we focus on firms that acquired at least four firms with patents between 1976 and 2005. This criterion keeps 333 major firms that conduct nontrivial amount of both R&D and M&A. Table (ref) reports their descriptive statistics. The average patent count (2,081 for R&D and 268 for M&A) is much higher than the median, which suggests relatively few firms have disproportionately large portfolios even within our selective sample. The three financial-performance metrics exhibit similar skewness. Consequently, we use the natural logarithm of these variables to mitigate heteroskedasticity in our subsequent analysis.

table[table omitted — 1,761 chars of source]

\paragraph{ Where Do Firms Patent?} Panel (c) of Table (ref) counts the number of USPTO classes in which the firms have patents. The median firm conducts R&D in 34.5 classes, whereas the mean is 65. The most diversified portfolio (Mitsubishi Electric) covers 358 of the 430 classes, followed by General Electric's 347. Hence, the portfolio aspect of innovation is highly heterogeneous. Appendix B illustrates what these portfolios look like in raw data.

Mapping Firms' Locations Over Time

We explain our method to study firms' locations in technological space in sections 4.1 and 4.2, and investigate its output---a shape graph---in section 4.3. Section 4.4 compares Mapper with Jaffe's (1989) clustering method.

The Mapper Algorithm

We propose patents as a measure of successful R&D\ investment. For each firm $i=1,2,...,333$, each year $t=1976,1977,...,2005$, and each patent class $c=1,2,...,430$, we count the number of patent applications, $ p_{i,t,c} $. Hence, each firm-year observation is a 430-dimensional vector $ p_{i,t}\in \mathbb{R}^{430}$ (i.e., we use patent class $c$ as an empirical analog of technological area $k$ in our theoretical model and assume $K=430$).

\paragraph{Preprocessing.}

Because firms' patent applications in any single year tend to be volatile and may not be representative of their underlying R&D activities, we follow BennerWaldfogel2008 to smooth out yearly fluctuations by aggregating them in a five-year moving window: $\tilde{p} _{i,t}=\sum_{\tau =t}^{t+4}p_{i,\tau }$. We take its natural logarithm to accommodate the highly skewed distribution of patent count (see section 3), \footnote{ This equation is our main specification of $f^{l}\left( \cdot \right) $ in section 2. We also use an alternative transformation (calculating shares of classes within each firm-year) due to Jaffe (1989) in Appendix D.}

equation[equation omitted — 90 chars of source]

Let $L=\left \{ l_{i,t}\right \} $ denote the entire panel dataset of firms' locations.

We propose mapping the entire $L$ in a single graph, instead of creating a map for each $i$ or $t$ (see Appendix H for such plots), for two reasons. First, our model in section 2 suggests firms' locations relative to each other determine the number of competitors $N_{m,t}$ in each market $m$, which in turn affects profits. Second, the model also suggests their historical trajectories contain relevant information about firms' R&D capabilities and profitability: dynamics matter. Fortunately, our topological method works well with such a dataset (i.e., many data points, or a \textquotedblleft point cloud,\textquotedblright \ with many dimensions).

\paragraph{Mapper.}

We first present the Mapper procedure in purely mathematical terms, and then provide more intuitive explanations. The procedure creates a simplified representation of complicated data in a graph (\textquotedblleft shape graph\textquotedblright \ or \textquotedblleft Mapper graph\textquotedblright ) that captures topological features such as branching, flares, and islands. Mathematically, this shape graph $G\left( L\right) $ is constructed in four steps.

enumerate• Project $L$ into $\mathbb{R}^{d}$ by some filter function $ f:L\rightarrow \mathbb{R}^{d}$, where $d<K$ is the dimensionality of a lower-dimensional space. • Cover the image $f(L)$ using an overlapping cover $\mathcal{C} =\{C_{j}\}_{j=1}^{J}$. • For each cover element $C_{j}$, apply some clustering algorithm to its pre-image $f^{-1}(C_{j})$ based on the dissimilarity function $\delta $ to obtain a partition of $f^{-1}(C_{j})$ into $Q_{j}$ clusters, $V_{j,q}$ ($q=1, \hdots,Q_{j}$): \begin{equation*} f^{-1}(C_{j})=\bigsqcup_{q=1}^{Q_{j}}V_{j,q}, \end{equation*} where the notation $\sqcup $ represents a disjoint union. • Construct the graph $G$ with nodes (vertices) consisting of all $ V_{j,q}$s. Connect two nodes, $V_{j,q}$ and $V_{j^{\prime },q^{\prime }}$, by an edge if $V_{j,q}\cap V_{j^{\prime },q^{\prime }}\neq \emptyset $.

Conceptually, the idea is to simplify the raw data $L$ by clustering data points within each local region (in steps 1, 2, and 3, which define a set $V$ of vertices or nodes) but make sure to preserve the sense of continuity across regions (in step 4, which defines a set $E$ of edges), so that the resulting graph $G=\left( V,E\right) $ retains the topology of the data on a global scale. Appendix C.1 offers a brief introduction to TDA. Appendix C.2 features an illustrated example (with $K=2$, $d=1$ , and $J=4$) to help the reader develop a more concrete understanding.

\paragraph{Connections to the Economic Model.}

The graph $G(L)$ provides a topological map of firms' technological locations $L$. The set of nodes $V$ is an empirical analog of the set of product markets $ \mathcal{M}$ that have ever been visited by any of the firms in our data. Hence, the local clustering in step 3 empirically determines the neighborhood $\mathcal{N}$ in equation ((ref)). The set of edges $E$ preserves their relative positions by indicating for each market which other markets are adjacent to it.

Practical Considerations

The Mapper procedure offers a \textquotedblleft telescope\textquotedblright \ to directly look at data points---even when they reside in a high-dimensional space---by focusing on a coordinate-free representation of the underlying data in terms of a graph. This graph preserves the relative positions of the original data points as long as they form a continuum. Hence, it is suitable for visualizing any high-dimensional data points that exhibit some sort of continuity.

As is the case with a real telescope, its practical usefulness depends on properly tuning its \textquotedblleft parameters\textquotedblright : (i) the filter function $f$, (ii) the number of cover elements $J$, (iii) the clustering method, (iv) the dissimilarity function $\delta $, and (v) the degree of overlap $o$ between cover elements. We explain the role of each parameter and our baseline specification.

\paragraph{Filter.}

The choice of $f$ in step 1 determines the \textquotedblleft angle\textquotedblright \ at which we look at the data. Some angles allow us to see richer patterns than others because they expose greater variation. A typical choice is PCA or MDS, but any other \textquotedblleft off-the-shelf\textquotedblright \ technique for dimensionality reduction can be used in principle. We use two-dimensional PCA as our baseline $f$ (i.e., we project $L$ to its first two principal axes, $ f:L\rightarrow \mathbb{R}^{2}$) because PCA is fast, deterministic, and well-understood, and preserves the largest variation in data by definition. As a sensitivity analysis, we also use MDS and three-dimensional PCA in section 5.4.

\paragraph{Resolution.}

In step 2, $J$ determines the resolution of the graph. The higher the resolution, the more details are revealed. But a fundamental limit exists. An arbitrarily high $J$ would result in a degenerate graph with as many nodes as data points but no edges. Because data points are discrete objects, we cannot preserve the sense of continuity between them if our scope is narrower than the distance between them. We set $J=400$ because it reveals sufficiently detailed patterns at the individual-firm level without losing their historical trajectories. We assess sensitivity with $225$ and $ 625$ as well.\footnote{ We use the Python implementation, KeplerMapper, by KeplerMapper2019, in which this parameter is operationalized as the \textquotedblleft number of cubes,\textquotedblright \ $n$, in each of the $d$ dimensions (e.g., $J=n^{2}$ when $d=2$). Thus, we implement $J=225$, $400$, and $625$ by setting $n=15$, $20$, and $25$, respectively.}

\paragraph{Clustering.}

Step 3 performs the main simplification task: clustering nearby data points. Conceptually, the most important point of Mapper is not the choice of clustering algorithm but the idea that this operation is performed only on a specific subset of data points (i.e., those within each $f^{-1}\left( C_{j}\right) $) at a time. Hence, any reasonable clustering method may be used. We use hierarchical clustering with single-linkage method, and follow Sing, M\'{e}moli, and Carlsson's (2007) heuristic for choosing the number of clusters. We assess sensitivity with five other specifications.

\paragraph{Dissimilarity.}

Clustering requires a measure of (dis)similarity between a given pair of firm-year observations, say $\left( i,t\right) $ and $\left( i^{\prime },t^{\prime }\right) $. We use the cosine distance,

equation*[equation* omitted — 198 chars of source]

because it has been commonly used since Jaffe1986. We also use Euclidean, correlation, min-complement (BarLeiponen2012), and Mahalanobis distances.

\paragraph{Overlap.}

Step 4 completes the graph representation by adding an edge to any pair of clusters (nodes) that share at least one observation. This \textquotedblleft sharing\textquotedblright \ of observations requires an overlapping region between adjacent cover elements. The degree of overlap $o\in \left( 0,1\right) $ governs the tolerance for detecting continuity, with values close to $0$ generating almost no edges and values close to $1$ detecting continuity almost everywhere. Such extreme values defeat the purpose of capturing the shape of the data; we set $o=0.5$ (i.e., 50% of a cover element's \textquotedblleft area\textquotedblright \ overlaps with each of its neighbors), and assess sensitivity with $0.3$ and $0.7$.

A Topological Map of the Technological Space, 1976--2005

The shape graph of Figure (ref) (b) embodies our first main result: a faithful representation of the 333 firms’ inventive activities across 430 technological areas. Pooling all 30 years of panel data allows us to track their movements within a single map, including many unique trajectories. Appendix H reports alternative results based on year-by-year Mapper graphs.

\paragraph{IT.}

Figure (ref) reproduces the northern half of Figure 1 (b) with greater detail. The main trunk consists of large nodes containing hundreds of firm-years (see the lower-middle part labeled “many engineering firms”). Their patents are relatively few and undifferentiated. Even famous IT firms started from this densely populated “heartland” of electronics in the 1970s, but their inventive activities diverged from the rest in the 1980s and evolved into unique trajectories in the 1990s and the 2000s. These dynamics coincide with the macroeconomic trend in which IT emerged as a dominant sector with new technological opportunities in many directions. To demonstrate the authenticity of our map more concretely, we investigate five historically important cases.

figure[figure omitted — 430 chars of source]

First, the patenting activities of Intel---a leading chip maker---used to be indistinguishable from the rest. Between 1976 and 1988, it moved around but was always surrounded by many other firms. In 1989--1990, however, it started marching in a new direction, and established a clearly unique track record by 1995. This timing coincides with Intel’s “near-death experience” in the mid 1980s, in which Japanese rivals squeezed it out of the memory market, and its subsequent shift to microprocessors (see grove1996). During the 1990s, it invested heavily in new microprocessor designs and became a household name (“intel inside”) as personal computers (PCs) became popular. Our map successfully captures these developments as an outward flare, because the underlying patent data distinguishes between “memory” (class 711) and “processors” (712), and Mapper handles all of the 430 dimensions equally well, including the ones for classes 711 and 712.

Second, HP is recognized as the symbolic founder of Silicon Valley because it produced the world’s first PC in 1968.\footnote{“The First PC” (https://www.wired.com/2000/12/the-first-pc/). Wired. December 1, 2000.} In 1984, HP introduced inkjet and laser printers for desktop computers, and retained focus on computers and printers through the 1990s, while its older business in test and measurement instruments was spun off into Agilent Technologies in 1999. Figure (ref) summarizes this history well. HP operated in the middle of the electronics heartland in 1976–1980 alongside many other device makers and defense firms. But its unique direction became clearly visible by 1984, as it started breaking new grounds with patents in class 347 (incremental printing of symbolic information). This path continued to grow into one of the longest flares in our graph. HP briefly “touched” IBM in 1999 (see below), before the Agilent deal made HP unique again.

Third, IBM generated more US patents than any other businesses. Its patenting activities are “off the chart” in both scale and scope, which our map visualizes as an “island” detached from all other firms. Nevertheless, IBM in 2001--2005 was sufficiently similar to HP in 1999--2003, and the two firms were briefly collocated near the end of HP’s flare. This rendezvous is not a coincidence: IBM went through major restructuring in 1993–2002 (see gerstner2002). Thus, this collocation reflects IBM's downsizing as well as HP's growth.

Fourth, Cisco became a poster child of the Internet age, as the world adopted the Internet Protocol (IP) in the mid-to-late 1990s. Founded in 1984, Cisco makes networking hardware and software. Its first patent was filed in as late as 1993. But its focus on classes 370 (multiplex communications) and 709 (multicomputer data transferring), which together account for 60% of its patents in our data, was so unique that its trajectory quickly evolved into a flare in the mid 1990s. Thus, a firm does not have to be patenting a lot to develop a flare as long as its direction is unique. Note Cisco's flare touches Microsoft's at two points in the 1990s, when the latter began to expand into networking (see below). This episode highlights another key aspect of competition and innovation: uniqueness is a relative concept. A firm’s flare length is based on the entire graph. Hence, it is determined not only by its own innovations but also by all other firms’.

Fifth, Microsoft dominated the PC operating system (OS) market, first with MS-DOS and then with Windows, which was released in 1985. Since the 1990s, Microsoft has increasingly diversified from the OS market. It introduced the Office suite in 1990, Internet Explorer in 1995, and Xbox in 2001. Hence, Microsoft’s patent portfolio is more diversified than Cisco’s, but their overall trajectories are similar: both of them were close to other IT firms until the late 1980s (Microsoft) or the early 1990s (Cisco) and then grew into individual flares. Their paths crossed again in the mid-to-late 1990s as Microsoft expanded into computer networking in 1995.

\paragraph{Engineering Conglomerates.}

Engineering giants cluster together and constitute a large island in Figure (ref) (a). General Electric (GE), an archetypical conglomerate, holds one of the most diversified portfolios in our data. Its only peers are similarly diversified manufacturers of electronic and capital goods, such as Siemens, Philips, and Mitsubishi Electric.

figure[figure omitted — 821 chars of source]

\paragraph{Pharmaceuticals and Chemicals.}

Health care is another R&D-intensive sector, and patent protection is crucial for its business model. Unlike IT firms, however, pharmaceutical firms do not appear in flares or islands. Large drug makers, such as Pfizer, Merck, and Eli Lilly, are clustered in the southern “peninsula,” as Figure (ref) (b) shows, because most of the drug patents are in either class 424 or 514 (both are labeled “drug, bio-affecting, and body-treating compositions”), which limits the extent to which their patent portfolios could differ from each other. Further investigations into drugs would require subclass-level data.

Household chemicals firms appear near drug makers because some of their products are based on similar materials. Johnson and Johnson (J&J), Unilever, Procter and Gamble (P&G), and Kimberly-Clark hold patents in not only classes such as 510 (cleaning compositions), but also 424 (drugs) and 604 (surgery).

Whereas most of the flares that we have scrutinized so far represented firms’ outward movements, the chemicals industry features a few counterexamples, that is, firms whose technological trajectories are centripetal (i.e., moving inward) rather than centrifugal (i.e., moving outward). Monsanto was famous for Roundup, a herbicide developed in the 1970s, but became an agri-biotech business in the 1980s and a major producer of genetically engineered crops. In 1997–2002, it divested most of agrochemical businesses and focused on biotechnology, adopting the R&D/patent-intensive business model of biotech drug companies. This novel strategy shows up as a long march inward, from the periphery to one of the core drugs clusters.

Imperial Chemical Industries (ICI) forms another centripetal flare. ICI used to be one of the largest British firms, but divested most of its bulk chemicals businesses in 1991–2007 to focus on specialty chemicals. One of its spin-offs, Zeneca, merged with Astra to form AstraZeneca, a drugs company, in 1999.

Finally, conglomerates in general chemistry (DuPont, 3M, and Dow) form their own long flares together, not unlike the engineering conglomerates’ island. Dow connects with the rest of the chemicals firms via its long centripetal flare, because it has been increasingly focusing on specialty chemicals, including materials for pharmaceuticals, paper coatings, and advanced electronics. Seeds from genetically modified plants also play an important role in its agri-business. Hence, its strategy is broadly similar to ICI's and Monsanto's.

Whereas most of the IT success stories are associated with long, centrifugal flares, some of the most interesting chemicals firms appear in centripetal flares. The reason is that many of them had already become big conglomerates by 1976 and were ripe for restructuring and divestiture, which tend to generate centripetal movements due to downsizing (recall the path of IBM). Thus, the contrast between IT and chemicals reflects their historical differences.

\paragraph{Summary.}

These examples demonstrate close connections between firms' locations on the map and their actual histories of R&D (we also investigate M&A patents in Appendix E.4). The ability to accurately track the trajectories of individual firms, as well as their collective patterns at the industry and sector levels, is Mapper's advantage over existing methods, such as PCA and clustering.

Comparison with Jaffe's (1989) Clustering Method

How do our results differ from Jaffe’s (1989)?\footnote{Appendix D explains their methodological differences in detail and presents an alternative Mapper graph based on Jaffe’s data-transformation convention.} Table (ref) shows a list of clusters that global clustering \`a la Jaffe generates. The grouping seems intuitive, with clusters of firms in engineering (cluster 1), telecommunications (2), materials (3), medical devices (4), pharmaceuticals (5), and so on. Jaffe studies firms that “move” over time, which he defines as firms that belong to multiple clusters over the years. For example, clusters 7 (computers), 10 (semiconductors), and 11 (electronics) commonly feature Intel and HP. Monsanto appears in both clusters 6 (chemicals) and 15 (genomics). Classifying them as “movers” is consistent with their long flares in our Mapper graph (see section 4.3).

table[table omitted — 2,626 chars of source]

However, Jaffe-style clustering misclassify many other firms. The following firms exhibit flares---and therefore clearly move---in our Mapper graph but do not “move” between the Jaffe clusters in Table (ref): Bosch (cluster 1), Ericsson (2), Kimberly-Clark (4), P&G (4), Dow (6), IBM (7), Lockheed Martin (9), National Semiconductor (10), Corning (17), and Applied Materials (AMAT, 18). They happen to be near the centers of their respective clusters. By contrast, Alcoa (clusters 3 and 17) and Roper Industries (14 and 20) appear in multiple clusters and would be classified as “movers” by Jaffe even though they hardly show any flares in our graph. They appear to “move” only because the clustering algorithm happens to draw boundaries in the middle of their data points (and not because they actually traveled long distances).

These “false negatives” and “false positives” highlight the arbitrariness of cluster boundaries. Jaffe's clusters do contain similar firms on average, but their boundaries are ultimately an artifact of discretization and add too much noise at the firm level. This lack of precision is consequential: Jaffe tried but failed to find statistically significant relationships between firms’ performances and whether they “moved” in the technological space. We tackle the same problem and find statistically significant relationships in the next section.

Measuring Unique Technological Trajectories

Given the prominence of flares and islands in the shape graph of our data, as well as their apparent connections to the firms' R&D strategies, their systematic measurement seems desirable. We formalize the notion of \textquotedblleft firms' unique technological trajectories\textquotedblright \ and propose a method to measure their lengths in section 5.1. We then establish their empirical relevance in terms of correlations with the firms' financial performances in section 5.2. Sections 5.3--5.5 present their economic interpretations, sensitivity analysis, and comparisons with other measures, respectively.

Definition and Measurement of Flares

We use graph theory to formalize the notion of firms' unique technological trajectories. Our exposition here is brief and intuitive; see Appendix C.3 for proofs and computational details.

We aim to define each firm's unique trajectory as a flare and measure its length in the graph $G=\left( V,E\right) $ of our data, which requires several auxiliary concepts. Let us focus on a subgraph $G_{i}$ of $G$ that consists of nodes that contain firm $i$ and the edges among them. We decompose $G_{i}$ into \textquotedblleft interior\textquotedblright \ and \textquotedblleft boundary.\textquotedblright \ The interior $F_{i}$ is the nodes in $G_{i}$ whose immediate neighbors also contain firm $i$, whereas the boundary $G_{i}\setminus F_{i}$ (i.e., the rest of $F_{i}$) consists of the nodes in $G_{i}$ that connect with nodes not containing firm $i$. Appendix C.3 features a pictured example.

We further decompose $F_{i}$ into \textquotedblleft isolated pieces\textquotedblright \ (connected components, formally) as $ F_{i}=R_{1}\sqcup R_{2}\sqcup ...\sqcup R_{S}$,\footnote{ In graph theory, a (connected) component of an undirected graph is a connected subgraph that is not part of any larger connected subgraph.} and classify each $R_{s}$ as either a \textquotedblleft flare\textquotedblright \ or an \textquotedblleft island.\textquotedblright \ If $R_{s}$ is also a connected component of $G$ (i.e., if it is \textquotedblleft isolated\textquotedblright \ in the context of the full graph), we call $ R_{s} $ an island of firm $i$. Otherwise, we call it a flare of firm $i$.

To introduce the notion of length, we define an exit distance for each node $u$ in $ F_{i}$ as

equation[equation omitted — 132 chars of source]

where $d\left( u,v\right) $ is the distance between nodes $u$ and $v$ in $G$. \footnote{ In graph theory, distance $d_{G}\left( u,v\right) $ is defined as the minimum length of paths in $G$ from $u$ to $v$, which we write $d\left( u,v\right) $ for short. We assume a unit weight on every edge when we calculate path lengths, but our method can be extended to handle any positive weights.} In words, the exit distance is the shortest length of path to get out of firm $i$'s interior. Thus, $e_{i}\left( u\right) $ represents the extent to which technological location $u$ (or all firm-year observations $l_{i,t}$ that constitute cluster $u$) is differentiated from the nearest rival's subgraph. In the case of islands, we set $e_{i}\left( u\right) =\infty $ because no such path exists.

Computing $e_{i}\left( u\right) $ based on its definition ((ref)) is costly because it requires information on the length of all paths in $G$. Fortunately, we can show that

equation[equation omitted — 146 chars of source]

where $d_{G_{i}}\left( u,w\right) $ is the distance between $u$ and $w$ in $ G_{i}$ (see Appendix C.3 for the proof). Thus, we can compute $e_{i}\left( u\right) $ based only on firm $i$'s subgraph $G_{i}$, not the entirety of $G$.

Next, we characterize each connected component (i.e., flare or island) $R_{s} $ of $F_{i}$ based on the longest exit distance of its constituent nodes,

equation*[equation* omitted — 89 chars of source]

and call it the flare index of $R_{s}$. In other words, we aggregate the node-level information about exit distances at the level of connected components. We further aggregate $\lambda _{i}\left( R_{s}\right) $ at the firm level by defining the flare signature of firm $i$ as the multiset\footnote{ A multiset is a modification of the concept of a set that, unlike a set, allows for multiple instances for each of its elements. We denote it by double braces $\{ \{,\} \}$ to distinguish it from a set.}

equation*[equation* omitted — 116 chars of source]

Four cases are possible. First, if $F_{i}$ is empty (i.e., no interior exists in $G_{i}$), no flares or islands exist, and we define $\vec{\lambda}_{i}$\ as an empty multiset. Second, if only flares exist in $F_{i}$, $\vec{\lambda} _{i}$ contains only finite elements. Third, if only islands exist in $F_{i}$ , $\vec{\lambda}_{i}$\ contains only copies of $\infty $. Fourth, if both flares and islands exist in $F_{i}$, $\vec{\lambda}_{i}$\ contains both finite elements and copies of $\infty $.

Finally, we define the flare length of firm $i$ as

equation*[equation* omitted — 367 chars of source]

where $\mathop{\mathrm{finmax}}(\vec{\lambda}_{i})$ is the maximum among all finite elements of $\vec{\lambda}_{i}$. Thus, we propose to measure the length of firm $i$'s unique technological trajectory by its longest flare.

Flares and Firms' Performances

These formal definitions help us detect all firms' flares, including those that are located within the densely populated areas. Table (ref) shows that, whereas our visual inspection in section 4 identified only a few dozen flares and islands, this systematic examination reveals the existence of many more: 40.3 % of our sample (133 firms) shows some flares.

table[table omitted — 831 chars of source]

What makes portfolios “unique”? Raw data at the firm level suggest both the quantity and variety of patents help make their portfolios unique. For example, HP has a massive portfolio and a flare of length 6, whereas Dell's portfolio is much smaller and its flare length is 1 (see Appendix B for further details on HP, Dell, and Qualcomm). However, these conditions are not sufficient for long flares, because uniqueness is a relative concept. Our definition of flare is based on $G$, the graph of all firms in all years. Hence, the firm's flare length depends on not only its own activities but also all other firms'.

In the remainder of this section, we investigate whether flares contain any “relevant” information. Following a common practice in the patent statistics literature (e.g., PakesGriliches1984, JAFFE198987, and hall2005market), we look for correlations between these topological characteristics and the firms' performance metrics, including revenue, profit, and stock market value.

Let us study their correlations by running regressions of the following form:

equation[equation omitted — 165 chars of source]

where $y_{i}$ is firm $i$'s revenue (or other performance metrics) in 2005, $\lambda_{i}$ is the flare length of its patent portfolio's evolution in 1976--2005, $\mathbb{I}\left \{ \lambda_{i}=\infty \right \} $ is a dummy variable indicating the islands-only type, $p_{i}$ is the total count of firm $i$'s patents in 1976--2005 (i.e., $p_{i}=\sum_{t} \sum_{c}p_{i,t,c}$), $\alpha $s are their coefficients, and $\varepsilon _{i}$ is an error term.\footnote{ Note we do not intend to prove causal relationships or their specific channels. Our purpose is to assess the extent to which our topological measures predict these performance metrics.} We include $\ln({p_{i}}) $\ to control for the size of the firm's inventive activities.

table[table omitted — 2,934 chars of source]

Table (ref) shows flare length is positively correlated with the firm's revenue, EBIT, and market value in 2005. Columns 1, 4, and 7 use the flare variables alone; columns 2, 5, and 8 use $\ln (p_{i})$ alone; and columns 3, 6, and 9 use both. The purpose of comparison is to assess whether our topological characteristics convey additional information above and beyond what patent count alone could predict. The differences between the adjusted $R^{2}$s suggest they do. More formally, the F-tests of a linear restriction, $\alpha _{2}=\alpha _{3}=0$, reject the null hypothesis at the 0.01%, 0.1%, and 1% levels for the revenue, EBIT, and market-value regressions, respectively.\footnote{We calculate $F=[(R^{2}_{ur} - R^{2}_{r})/2] / [(1 - R^{2}_{ur})/(\#obs - 4)]$, where $R^{2}_{ur}$ is the $R^{2}$ of the unrestricted model in column 3 (6 or 9), $R^{2}_{r}$ is the $R^{2}$ of the restricted model in column 2 (5 or 8), and $\#obs$ is the number of observations (328, 301, or 325). We reject the null hypothesis, $\alpha_{2}=\alpha_{3}=0$, if $F$ is greater than the corresponding critical value of the F distribution.} Hence, the incremental contribution of the flare-and-island variables is statistically highly significant.

What about their economic significance? The estimates of $\alpha _{2}$ are 0.34, 0.33, and 0.27 in columns 3, 6, and 9 (i.e., after controlling for $p_{i}$), respectively, which imply an extra length of flare is associated with 40%, 39%, and 31% higher performances in terms of revenue, EBIT, and market value, respectively.\footnote{Likewise, the estimates of $\alpha _{3}$ (0.95, 0.94, and 0.70 in the same three columns) suggest islands-only firms tend to outperform no-flare firms by 159%, 156%, and 101% in these measures, respectively. However, their standard errors are large. Only three firms belong to this category, and all of them have relatively large patent portfolios, which makes $\alpha _{3}$ difficult to isolate from $\alpha _{4}$. Nevertheless, we keep $\mathbb{I}\left \{ \lambda_{i}=\infty \right \} $\ in these columns, because dropping it (and thereby grouping them with no-flare firms) would be unwise given the results on columns 1, 4, and 7.}

Economic Interpretations

Why do flares predict firms' success? Let us interpret these findings based on our model in section 2. First, flares reflect unique technological trajectories. Unique technologies permit product differentiation, which softens price competition (or avoid competition altogether) and increases profits. Specifically, unique technological location $l_{i,t}$ allows the firm to enter a new product market with low $N_{m,t}$. This mechanism directly connects $l_{i,t}$ with $\pi_{i,t}$.

Second, these extra profits could help finance subsequent R&D expenditure $b^{x}_{i,t}$, thereby reinforcing the firm's technological differentiation and conquest of new markets: a virtuous cycle. The length of flare reflects a string of unique $l_{i,t}$s and a track record of successful technological development in a unique direction. Hence, it is a good proxy for the duration of such virtuous cycles. These dynamics imply positive correlations between $\lambda_{i}$ and $\pi_{i}$.\footnote{One might wonder how our definition of flare length---which does not explicitly incorporate the time dimension---can capture the firm's actual duration of travel without bumping into its rivals in real time. We discuss this issue in Appendix C.4.}

Third, the fact that $\lambda_{i}$ conveys information above and beyond what $p_{i,t}$ predicts---which is known to be strongly correlated with firm size and R&D expenditure (e.g., Cohen2010)---suggests $\lambda_{i}$ captures more than just budget size $b_{i,t}$. Our model predicts connections between $\lambda_{i}$, $l_{i,t}$, and technological capabilities $\mathbf{\omega}^{x}_{i,t}$; our findings from panel-data regressions (in section 5.4) confirm the presence of persistent firm heterogeneity and its correlation with $\lambda_{i}$.

Thus, our empirical results---interpreted in the context of our model of competition and innovation---highlight the importance of the direction of innovation. Unique technological positions directly contribute to profits, which reinforces subsequent innovations and long track records. These dynamics reflect the firms' desire to avoid competition, conquer new markets, and exploit their idiosyncratic technological capabilities.

Finally, why are some firms profitable despite showing short or no flares? Our model permits two firm-level characteristics other than technologies: quality $\xi_{i,t}$ and cost $c_{i,t}$. Those who have comparative advantage in marketing or operations (i.e., high $\omega^{\xi}_{i,t}$ or $\omega^{c}_{i,t}$) would keep exploiting the existing markets by investing in $\xi_{i,t}$ or $c_{i,t}$ instead of technologies.

Sensitivity Analysis

This section assesses the sensitivity of our results to (i) the specification of the Mapper procedure, (ii) subsampling based on firms’ survival, (iii) subsampling based on industry classification, and (iv) panel-data regressions.

table[table omitted — 2,586 chars of source]

\paragraph{Mapper Specification.} Table (ref) reports descriptive statistics of the Mapper graphs under 16 different specifications. Our baseline Specification 1 (S1) generates a graph with 1,214 nodes, 2,926 edges, the average degree of 4.82 (edges per node), 27 connected components, 8.70 nodes per firm, and the average flare length of 0.77. Most of the alternative specifications lead to changes that are either small (S6--S14) or in directions that are consistent with Mapper’s mechanism (S2, S4--S5, and S15--S16). S3’s direction of change is less obvious because it is the only one that uses a non-PCA filter (i.e., takes a different “angle” at the data). Nevertheless, its descriptive statistics are comparable to others. Given the diverse set of specifications, perhaps the most surprising finding is that their regression results are remarkably similar to the baseline. Appendix E.1 explains S2--S16 in detail and shows the correlations between firms’ performances and flare length (based on the 16 different graphs) are always positive and statistically significant, with comparable magnitudes.

\paragraph{Survivorship.} Appendix E.2 shows the results are robust to (i) the elimination of firms that exited our sample before 2005 and (ii) conditioning on the balanced panel.

\paragraph{Subsampling by Sector and Industry.} These findings are not an artifact of aggregation or driven by a few specific sectors and industries. Appendix E.3 plots revenues and flares by economic sector defined by Standard and Poor's (S&P), a credit-rating agency. Appendix E.3 also studies the technology sector more deeply at the SIC-code level, with a focus on computers and semiconductor industries. The positive correlations are preserved within each sector and industry.

\paragraph{Panel Data Regressions}

Whereas our analysis in section 5.2 focuses on the relationships between the firms' flares in the whole graph for 1976--2005 and their eventual performances in 2005, Appendix F shows our findings hold more generally---at different points in time, with many years of lags, and in terms of out-of-sample predictions.

Comparison with Other Measures

This section compares flare length with other measures, including more conventional network-centrality measures and the Jaffe measure of technological distance.

\paragraph{Centrality Measures.} Flare length is the focus of our quantitative analysis because (i) long flares are the most salient feature of our Mapper graph and (ii) our model suggests the length of unique technological trajectories may reflect the firms’ profitability and capabilities. Nevertheless, flare length is not the only way to measure locations on a graph. Measures of network centrality offer more conventional alternatives. Appendix G.1 shows five centrality measures (degree, closeness, harmonic, betweenness, and eigenvector centralities) correlate with the firms' financial performances less strongly than our flare-based measures.

\paragraph{Jaffe's Technological Distance.}

Both our Mapper graph and Jaffe's (1989) measure of technological distance use patent count and almost identical dissimilarity functions. Hence, one might expect Jaffe's measure to produce similar results. When we regress revenue, EBIT, and market value on the Jaffe distance, however, the fit is nearly zero in many cases (columns 1, 4, and 7 of the table in Appendix G.2). It achieves a reasonable fit when patent count is also included (columns 2, 5, and 8), but its coefficient estimate is statistically insignificant and difficult to interpret (i.e., negative) in most cases. Finally, the inclusion of our flares and islands further improves the adjusted $R^2$, but the coefficient on Jaffe's measure remains insignificant and lacks cohesive patterns.

Conclusion

This paper proposes a new method to map, describe, and characterize firms' inventive activities. The shape graph from the Mapper procedure helps us understand where firms and industries are located, how they connect with each other (or not), and how their innovative activities evolve over time. In the past, economists' ability to answer these basic, descriptive questions---and hence the ability to ask and answer deeper, causal/policy questions that presuppose reliable descriptions or stylized facts---have been constrained by the “curse of dimensionality” of the technological space. With the new tool, we can start revisiting and answering some of the long-standing questions in economics, including the rate and direction of inventive activity. Because its underlying mathematics is general, we believe this method is potentially useful for describing and characterizing other high-dimensional data in economics as well, such as product characteristics and international trade.

\pagenumbering{arabic}