EconBase
← Back to paper

Shapes as Product Differentiation: Neural Network Embedding in the Analysis of Markets for Fonts

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

81,069 characters · 29 sections · 79 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Shapes as Product Differentiation

abstractMany differentiated products have key attributes that are unstructured and thus high-dimensional (e.g., design, text). Instead of treating unstructured attributes as unobservables in economic models, quantifying them can be important to answer interesting economic questions. To propose an analytical framework for these types of products, this paper considers one of the simplest design products---fonts---and investigates merger and product differentiation using an original dataset from the world's largest online marketplace for fonts. We quantify font shapes by constructing embeddings from a deep convolutional neural network. Each embedding maps a font's shape onto a low-dimensional vector. In the resulting product space, designers are assumed to engage in Hotelling-type spatial competition. From the image embeddings, we construct two alternative measures that capture the degree of design differentiation. We then study the causal effects of a merger on the merging firm's creative decisions using the constructed measures in a synthetic control method. We find that the merger causes the merging firm to increase the visual variety of font design. Notably, such effects are not captured when using traditional measures for product offerings (e.g., specifications and the number of products) constructed from structured data. JEL Numbers: L1, C8. Keywords: Convolutional neural network, embedding, high-dimensional product attributes, visual data, product differentiation, merger.

Introduction

Many differentiated products considered in economic analyses have important attributes that are unstructured. Examples include design elements in products such as automobiles, houses, furniture, and clothing or digital products such as mobile applications. Other obvious examples include creative features in books, music, movies, and fine arts. Unstructured attributes in these products are typically in visual or textual forms and thus are high-dimensional. More generally, products well beyond these categories are often presented to consumers in visual and textual forms: for example, product packages in supermarkets and online catalogs in e-commerce (e.g., Amazon, Airbnb, Yelp, Zillow). These attributes are one of the first pieces of information consumers receive along with more structured attributes such as price and product specifications. As a result, unstructured attributes are important decision factors for consumers and thus are key decision variables for producers.

Economists are aware of the role of unstructured attributes. Product attributes are an important component of economic models, such as discrete-choice models (mcfadden1973conditional, berry1995automobile) and hedonic models (rosen1974hedonic, bajari2005demand). These models treat product attributes as both low-dimensional observable variables and a (typically scalar) unobservable variable. In these models, the scalar unobservable variable normally captures the high-dimensional, unstructured attributes, including design and other original features of the products.

Although this tradition has its own merits, certain economic questions are better answered by treating unstructured attributes as observables. For example, one may ask how the style of products evolves over time---that is, how the fashion changes---in accordance with market conditions. One may further ask how product differentiation in this creative dimension affects the product\textquoteright s market power or is affected by market shocks such as vertical integration. These questions entail the issues of measurability and dimensionality of unstructured attributes. Due to these challenges, there has been little understanding about the production of creative attributes---namely, product differentiation decisions---in the realm of quantitative economics.\footnote{galenson2000age,galenson2001creating study artists and their career choices and paths in fine arts. Their main approach to quantitative analyses is to use price as a proxy to measure the value of artworks.}

In this paper, we propose a framework to quantify the design-oriented attributes of a product and construct a low-dimensional space of products and measures for the degree of product differentiation using the Euclidean distance endowed in the space. We then illustrate how some of the economic questions above can be answered using this framework, such as market-driven product differentiation decisions of product designers in multi-product firms.

For the purpose of this paper, we consider a particular design product: fonts. We use a dataset obtained from the world's largest online marketplace for Roman alphabet fonts. There are a few reasons to study the market for fonts. First, font is one of the simplest visually differentiated products. The shapes (i.e., two-dimensional monochrome visual information) of a fixed number of characters mostly describe the product.\footnote{These shapes are called typefaces.} This visual simplicity greatly facilitates our analysis. Second, the visual information is simple to understand but important in predicting the functionality and value of the product. Third, fonts are ubiquitous products; thus, the market for fonts is large with frequent productions and transactions. The online marketplace we consider is the world's largest and has over 28,000 fonts (produced by font design firms called foundries) and 2,400,000 transactions over the past six years. Fourth, interesting policies are involved in this market, such as vertical integration. Finally and most importantly, font is a stylized product that saliently captures a key aspect many products in the market have in common: design attributes.

The main challenge in quantitatively analyzing the market for fonts is that the main product attributes, their shapes, are high-dimensional. To address this challenge, we represent font shapes as low-dimensional neural network embeddings and construct a corresponding space of fonts. In particular, we adapt a state-of-the-art method in convolutional neural networks (wang2014triplet, schroff2015facenet), where the network directly learns to map font images to a compact Euclidean space---that is, the embedding space---in which perceived visual similarity is preserved. The algorithm is sophisticated enough to recognize the style of font shapes, which is crucial for our purpose.

Another challenge is to ensure that the resulting embeddings represent economic agents' perceptions of the product's visual information. To this end, we employ two strategies. First, we use the images of entire pangrams (instead of individual alphabet letters) as inputs in the neural network.\footnote{A pangram is a sentence that contains all the alphabet letters.} Because pangrams effectively capture important design elements that cannot be seen in individual letters (e.g., spacing, deep-height, up-height, and ligature), they are the most relevant decision variables for font designers and consumers. Second, we demonstrate that the obtained image embeddings contains a substantial amount of information that is mutually shared with tags, which are word phrases that describe fonts (e.g., “curly,” “flowing,” “geometric,” “organic”). Tags are assigned to each font by font designers and consumers and thus, we believe, represent the economic agents' perceptions of the product. We construct word embeddings from these tags using a simple neural network and calculate the mutual information between the word and image embeddings.

Why do we consider a convolutional neural network? Font shapes involve a non-linear interaction between many neighboring pixels. Considering individual pixels separately or recognizing interactions in a restrictive model provides little information about the overall shape of a font. By considering how neighboring pixels interact in a very flexible model, the deep neural network outperforms other machine learning methods such as LASSO, random forest, and boosting that use pixels or other hand-designed features (edges, corners, etc.); see Goodfellow-et-al-2016. In particular, the deep convolutional neural network is designed to effectively capture the spatial correlation between nearby pixels. Although neural networks are generally known to be less interpretable (friedman2001elements) than other learning methods due to model flexibility,\footnote{One approach to overcome this is to consider “hand-crafted features.” However, feature selection can generally be arbitrary and there can be arbitrarily many possible features, which hinders the interpretation.} we show how an interpretable embedding space can be learned through visual similarity. Most importantly, instead of attempting to interpret each embedding value, our approach is to give meanings to the distance metric of the embedding space and subsequently construct the product differentiation measures based on it.

Given the embedding space, a font designer's decision of a typeface design is equivalent to choosing a location in the space. This location choice is a strategic decision that depends on the choices of other designers in the space. In this sense, the abstract space of fonts we construct can be viewed as a location-analog model (hotelling1929) with Lancasterian characteristics (lancaster1966new,lancaster1971consumer). Based on this space, we construct two alternative differentiation measures using the image embeddings, namely a distance to Averia (i.e., the average font) and a gravity measure, which succinctly represent the location choice of a designer relative to others' choices.

To illustrate the usefulness of our approach, we conduct a causal analysis of how a merger affected the merged firm's design decisions before and after the merger. In June 2014, a major font foundry was acquired by the company that owns the online marketplace. Before the merger, the merging firm was selling fonts as a third party and competing with the foundries owned by the marketplace. The main motivation of the analysis is that designers are not only artists but also economic agents who are affected by market conditions. We use the constructed design differentiation measures as the main dependent variables. To estimate the effect of merger on differentiation, we use the synthetic control method (abadie2003economic, abadie2010synthetic). We employ this method because, although there is only one treated (i.e., merged) foundry, we can construct a suitable weighted average of a comparison group from untreated foundries. We provide arguments why strategic spillovers may be weak in our setting by showing that each control foundry produces substantially different products than the treated foundry, even though the synthetic control behaves similarly as the treated before the merger. The latter is achieved due to the rich information contained in the embeddings, which serve as the main predictors in constructing the synthetic control.

Our main finding is that, relative to the synthetic control, the merged foundry produced fonts with greater visual variety after the merger and that this effect was statistically significant. One of the explanations is that the degree of product differentiation may have increased after the merger to avoid cannibalization. Notably, we find that such effects are not captured when traditional measures for product offerings are used from structured data (i.e., the number of products and specifications; berry2001). This illustrates the importance of more sophisticated product offerings measures as employed in the current paper.

Contributions and Related Literature

Machine Learning and Social Science Research

To our knowledge, this is one of the first few papers using neural network embeddings for visual data in the economic analysis of markets and industries. Earlier work that analyzes visual data uses methods that are partly or fully human-aided. glaeser2018big use visual data from Google Street View to predict the economic prosperity of neighborhoods, but they assign scores to street images based on human surveys on the visual perception of street quality and safety. gross2016creativity investigates how competition influences creative production in a commercial logo design competition. Whereas gross2016creativity uses hand-crafted features to create a perceptual hash code for comparing images, we learn image embeddings based on recent advances in deep learning that have been shown to work extremely well for high-dimensional image data (krizhevsky2012imagenet, simonyan2014very, he2016deep). Deep learning methods are used in more recent studies. zhang2017much show how the quality and specific attributes of property images on Airbnb can affect the demand. The quality and attributes are human-labeled, and a convolutional neural network is used to train the images based on the labels. Our approach does not require subjective human labeling. As independent and contemporaneous work, bajari2018hedonic estimate a hedonic function for apparel consumption using a deep neural network based on visual and textual data. We also utilize textual data but in such a way that gives economic meanings to the image embeddings we obtain. The most significant difference between our approach and the two aforementioned studies is that we consider images as the main response variables of market shocks and construct differentiation measures based on neural network embeddings. Using these measures also helps gain interpretability and robustness in the quantities resulting from the network training. In addition, we use a different neural network algorithm based on image triplets that is suitable to our setting. Another recent work by magnolfi2022triplet considers triplet embeddings to characterize the product space for demand estimation. Although we share a similar motivation, their embeddings are computed after human making comparisons of product triplets while ours are elicited directly from neural network.

This paper is also among the first social science studies that make use of embeddings as part of empirical analyses. As another form of unstructured data, text data have recently gained much attention in economic analyses; see gentzkow2019 for a thorough review of the machine learning applications. Kozlowski_2019 use word embeddings to understand cultural norms. gentzkow2019measuring analyze political polarization using congressional speeches as text data. hobergetal2016 use text data from firms' 10-K product descriptions across industries to classify competing products and construct a product location space as we do in our paper. Unlike their paper, however, we use image embeddings and employ neural networks as a classification method. We also use word embeddings as a way of validating the image embeddings. Moreover, we focus on a particular industry as opposed to multiple industries and utilize detailed structured and unstructured data about product offerings.

Merger and Product Differentiation

Market structures and endogenous product differentiation have been important themes in the field of economics. In a theoretical paper, mazzeo2018 find that the effects of a merger on product differentiation can be ambiguous, implying that the question is more of empirical research, as in empirical industrial organization.\footnote{mankiw1986free theoretically consider a more general question of oligopolistic competition and product differentiation and again suggest that the direction of entry bias can be unclear.} sweeting2013 studies the dynamic aspects of product differentiation in the radio industry and finds some evidence suggesting that increased concentration increases variety. fan2013 finds in newspaper markets that mergers between local competitors has effects on vertical differentiation, leading to a decrease in the news quality. In related work, fan2020competition consider multi-product firms in smartphone markets and show how mergers may lead to a decline in the number and variety of products, and wollmann2018trucks's analysis of a commercial truck industry implies the opposite direction of the merger effect on product variety. Apart from this important line of work, the literature using a structural approach to study the effects of merger on non-price and potentially unstructured attributes is rather scarce due to difficulties in modeling endogenous product offerings with high-dimensional attributes. This is in contrast to merger effects on price, which have received relatively more attention in the literature, e.g., building on the structural framework of nevo2001measuring.

The structural approach has been complemented in the literature by studies of merger and concentration from a viewpoint of treatment effects and program evaluation. berry2001 and sweeting2010effects document the effects of mergers on product variety in local radio markets.\footnote{See also atalay2020post for a related merger analysis in consumer goods markets. They find that mergers lead to dropping of products that are dissimilar to existing ones, where the measure of dissimilarity is calculated based on structured attributes.} hastings2004vertical and ashenfelter2008effect study the effects of mergers on prices using the difference-in-differences and instrumental variable methods, respectively. Although the treatment effect approach to merger analyses has limitations (nevo2010taking), it is still suitable to highlight the rich information contained in the visual dimension of product attributes previously neglected in the literature on mergers.\footnote{See angrist2010credibility and nevo2010taking for discussions on how the two approaches can complement each other and what their pros and cons are.} We introduce unstructured data of images and related machine learning methods to derive new insights into creative product differentiation and its relationship to mergers. On the other hand, all the empirical studies listed in this and the previous paragraphs use structured data to construct response variables including price and non-price attributes (e.g., product variety measures).\footnote{For example, fan2013 uses data on the number of opinion section staff members, the number of reporters, local news ratio, variety, frequency of publication, and edition. sweeting2013 uses Neilson data on broadcasts. berry2001 use the number of stations and programming formats as product offerings.} As mentioned, we cannot find in our analysis the merger effects on the traditional product offerings measures. We believe this paper is a good starting point for introducing embeddings into economic analyses.

Machine Vision

Recognizing letters (e.g., distinguishing handwritten “G” from “Q”) is one of the most well-studied areas of machine vision as is done with the MNIST database (lecun2010mnist). Our paper, however, is one of the first that applies machine vision techniques to recognizing the style of font images (e.g., distinguishing typeface “G” from “G”), which is a more challenging vision problem. tenenbaum2000separating propose bilinear models that separate style and content with fonts as one of the examples. o2014exploratory develop a method for searching fonts using relative attributes based on the work on attributes and whittle search by parikh2011relative and kovashka2012whittlesearch. campbell2014learning develop a procedure for learning a font manifold by parametrizing font shapes and reducing the dimension of the resulting model.

The method of training the font embeddings builds on the works of schroff2015facenet and wang2014triplet, who develop a face recognition algorithm that directly learns an embedding for images via neural network training. schroff2015facenet show that their approach performs substantially better than the earlier approaches of training a classification network for face recognition, such as those in taigman2014deepface and sun2015deeply. The former approach is suitable for our purpose, as the procedure produces embeddings as the intermediate output of the classification algorithm. Although we are not directly interested in the classification of font identity, embeddings serve as our object of primary interest.

Fonts can be viewed as fashion products. Our quantitative analysis of the trend in font style is related to work by, e.g., al2017fashion, mall2019geostyle, and yu2019thinking, who apply advanced machine vision techniques they develop to recognize the style of clothes and shoes in the fashion industry and understand the trend. The analysis of the visual attributes of design products has also been considered, e.g., in burnap2016estimating and dosovitskiy2016learning using deep generative models with applications to furniture and automobile designs. However, none of the studies conduct causal analyses to answer economic questions.

Organization of the Paper

In the next section, we provide the background about the font industry and the online marketplace for fonts considered in this paper. Section (ref) describes the data obtained from this market. In Section (ref), we construct the embedding and the product space using the neural network. In this section, we demonstrate that the embeddings are meaningful by calculating the mutually shared information between the font embeddings and tags. Section (ref) presents the causal analysis of a merger using the embeddings. Section (ref) concludes. Sections (ref)--(ref) in the Appendix respectively present the internal evaluation of the trained network, the descriptive analyses of style trends, and supplemental results from the merger analysis.

Online Marketplace for Fonts

We consider the world's largest online marketplace \href{https://www.myfonts.com}{MyFonts.com} that sells around 30,000 different fonts. This market is a superset of all major global online stores for fonts. MyFonts.com and the other stores are all owned by Monotype Inc. A font is a delivery mechanism for typefaces. Therefore, fonts are sold as a piece of software, for which consumers purchase a license. Licenses are protected by the End User License Agreement (EULA). Mostly two types of licenses are sold to consumers. A web font license allows fonts to be displayed on a website, and a desktop license is for printed material. The marketplace also serves as a platform for third-party fonts. As a result, it sells fonts designed by foundries owned by Monotype as well as fonts from third-parties foundries.\footnote{A foundry is a group of designers who create fonts.}

Figure (ref) shows the home page of MyFonts.com.

figure[figure omitted — 196 chars of source]

An example of a font family page on MyFonts.com is captured in Figure (ref).

figure[figure omitted — 207 chars of source]

In the font industry, a family represents the identity of a font, which name is the name of the font. A font family is consist of several different styles, such as regular, light, bold, and italic. Figure (ref) shows different styles of Gilroy family. In this market, typical consumers are independent designers who use fonts as intermediate goods. They produce printed material (e.g., posters, pamphlets, cards), for which a desktop license is purchased, or webpages and digital ads, for which a web license and digital ads license are purchased, respectively. In the data collection period of 2012--2018, around 2,400,000 purchases were made.

Data

Overview

Our sample comprises data from 2002 to 2017. The dataset includes, in total, 28,659 fonts and 2,446,604 orders. The main information contained in the dataset is product attributes and transactions for each consumer. The unstructured high-dimensional attributes include images of typefaces and tags (i.e., descriptive words assigned by producers or consumers). The structured attributes include price, category types, license types, the number of languages supported, the number of glyphs supported, the foundry and designer information, and the date of introduction in the market. There are roughly six category types: sans serif, serif, slab serif, display, handwritten, and script. Transaction data include information on individual orders made by consumers and consumer characteristics such as the country and city of origin. We will revisit this dataset in Section (ref) for the merger analysis. For now, we focus on the unstructured data.

Visual Attributes

Fonts are displayed on the webpage using pangrams that contain all the alphabet letters in one sentence.\footnote{In font markets, many different pangrams are used: “The quick brown fox jumps over a lazy dog.” “Six quite crazy kings vowed to abolish my pitiful jousts.” “Quincy Jones vowed to fix the bleak jazz program.” “Mozart's jawing quickly vexed a fat bishop.” Here we chose to use one of the shortest pangrams to minimize the size of the image.} Pangrams effectively capture important design elements that cannot be seen in individual letters, such as spacing, deep-height, up-height, and ligature. We use pangrams as direct inputs in the neural network in order to mimic a consumer's actual perception of the products. Figure (ref) shows the examples of a pangram that roughly correspond to the product categories (i.e., sans serif, serif, slab serif, display, handwritten, and script).

figure[figure omitted — 743 chars of source]

The format of pangram images is a bitmap with $200\times1000$ pixels, where each pixel is a greyscale with a value between 0 and 255. We use random crops of $100\times100$ pixels as inputs in the network.

Construction of Embeddings

Neural Network Embedding

We employ a method in which the network directly learns a mapping from pangram images to a compact Euclidean space. This mapping is called an embedding. We map each pangram to a 128-dimensional embedding, denoted as $f(x)\in\mathbb{R}^{d}$ for pangram image $x$ with $d=128$.\footnote{It is important to have sufficient dimensions so that the neural network can allow for variation in the embedding space based on the actual images. The embedding with a larger number of dimensions would perform better, but it requires more training data to achieve the same level of accuracy while avoiding the risk of overfitting. We additionally normalize each embedding so that it lies on a 128-dimensional hypersphere, i.e., $\left\Vert f\right\Vert _{2}=1$ with the Euclidean norm $\left\Vert \cdot\right\Vert _{2}$.} The Euclidean distance in the resulting embedding space corresponds to the measure of similarity of font shape. For training the network embedding, we adapt a modern algorithm developed by schroff2015facenet for face recognition. Their algorithm determines the identity of a person based on face images in two steps. In the first step, a deep convolutional network is trained to learn an embedding space of faces. The rationale is that similar faces should occur closer in the embedding space than dissimilar faces. In the second step, the identities of the images are classified by choosing a threshold in the space below which the embeddings have the same identity. They show that this approach performs substantially better than the earlier approaches of training a classification network (taigman2014deepface, sun2015deeply). Our algorithm builds on schroff2015facenet, which approach is suitable for our purpose. First, pangram images can be classified based on a structure similar to that for face images. In the dataset of face images, multiple images are associated with the same identity. Analogously, multiple styles are associated with the same font family as detailed below. Second, the approach produces embeddings as the intermediate output of the algorithm. Although we are not directly interested in the classification of font identity, embeddings serve as our object of primary interest.

Triplet Loss and Network Training

The neural network is trained to produce embeddings and classify fonts that are within the same family in the resulting embedding space. In practice, we accomplish this by constructing triplets of images. Triplet $i$ comprises anchor $x_{i}^{a}$, positive $x_{i}^{p}$, and negative $x_{i}^{n}$. An anchor is a pangram image of a given font family (e.g., Helvetica), positives are pangram images of the same family but different styles (e.g., Helvetica Regular, Helvetica Light, Helvetica Bold, Helvetica Italic), and negatives are pangram images of different families (e.g., Time New Roman).\footnote{To be precise, pangram images here refer to crops of the images. Therefore, we also use different crops within the same pangram as positives.} The way triplets are sampled is analogous to that in the face recognition problem, where positives are images of the same person as the anchor and negatives are images of different persons. For triplet $(x_{i}^{a},x_{i}^{p},x_{i}^{n})$ in the entire set of font images, we enforce the following inequality during the network training:

align[align omitted — 158 chars of source]

where $f(x)\in\mathbb{R}^{128}$ is the embedding of image $x$, $\left\Vert \cdot\right\Vert _{2}$ is the Euclidean norm, and $\alpha$ is an enforced margin. That is, we ensure that the distance between an anchor and positive is smaller than that between the anchor and a negative.\footnote{The choice of triplets is very important for the fast convergence of the algorithm. As such, we make sure we choose sufficiently many triplets with hard positives and negatives, namely triplets that violate (ref).} The margin $\alpha$ allows the images for one font family to stay closer in the embedding space, while still discriminating the images of other families.

Then, a triplet-based loss function that is minimized in the network training is

equation[equation omitted — 172 chars of source]

We optimize this objective using stochastic gradient descent (SGD; bottou2010large). SGD is an iterative method for optimizing an objective function---in our case, for creating a reasonable embedding space. Because our dataset size is in gigabytes, it would be computationally challenging to compute the gradient of the entire dataset and optimize it using a more traditional optimization algorithm (e.g., Nelder-Mead or the conjugate gradient method). SGD can be regarded as a stochastic approximation of gradient descent optimization. It replaces the actual gradient by an estimate of the gradient.

We use approximately 20,000 images of fonts to train the neural network. The training iteratively improves the parameters of the network using small batches of images to estimate the gradient and then update the parameters accordingly. As the gradient is evaluated at more batches, the parameters in the network are adjusted. The training of the network is completed when the loss function reaches below a certain threshold. Section (ref) in the Appendix contains the details.

As internal evaluation, we show in Section (ref) how the neural network performs in the original classification task of identifying the font family. Although we are interested in the embeddings and not the classification, it is important to evaluate the embeddings based on their ability to classify. If the embeddings perform poorly, then they are not likely to generate a reliable embedding space. As mentioned, the classification of font families is conducted by thresholding the distances between embeddings. Overall, the neural network embeddings perform well in differentiating between fonts in different families.

Constructed Product Space

The trained neural network produces embeddings with 128 dimensions, which define a 128-dimensional space of font products. Figure (ref) visualizes this space by projecting it onto a two-dimensional space using t-distributed stochastic neighbor embedding (t-SNE).\footnote{t-SNE is a useful tool to visualize high-dimensional data in a two or three dimensional space. In our setting, we visualize 128-dimensional objects in a two-dimensional space.} For expositional purposes, the thumbnail is created with the word “Quick” that is cropped from the full pangram. Each thumbnail corresponds to the embedding of the regular style of each font family as a representative style. Therefore, for example, boldface in this figure is a design aspect of a particular font family, not a bold style within a family. A visual inspection reveals that different styles of fonts are well-clustered together even in this two-dimensional projection. More specifically, we can see that the western region is populated with fonts that have narrower characters and the eastern region with fonts that are thicker. In addition, the shape of fonts in the northeast is more geometric while the shape in the southwest is more curly. It is worth noting that these patterns are observed even in this space with restricted dimensions. The space we base our economic analysis on is the original 128-dimensional space.

figure[figure omitted — 467 chars of source]

To further understand how well the fonts are clustered in the 128-dimensional product space, we present in Figure (ref) the examples of regular fonts that are close to each other in the space. In Figure (ref), the images of six nearest neighbors in terms of the Euclidean distance are listed in each row of the table.

figure[figure omitted — 230 chars of source]

Even though we only list regular fonts, it is clear that fonts with different thicknesses are clustered together in the space. Also, the middle and the last rows show that geometric and curly fonts tend to cluster together, respectively.

Unlike individual embeddings, the distance metric endowed in this space has the clear interpretation of visual similarity and is more robust to model specifications. These aspects allow us to use the distance metric as the main building block in the economic analysis in Section (ref).

Relevance of Embeddings

Although the neural network embedding performs well in terms of classifying fonts of similar shapes, good predictive performance does not necessarily imply the relevance of the embeddings for economic analyses. To address this concern, we verify that the visual attributes captured in the resulting embeddings are relevant to economic agents' perceptions by measuring (a generalized notion of) their correlation with “perceived” attributes. For the latter, we use information from tags, which are short descriptive words assigned to each font family by font designers and consumers. Examples of tags are “curly,” “flowing,” “geometric,” “organic,” “decorative,” and “contrast.” These descriptive words of fonts are also high-dimensional, as the tags include nearly 30,000 different words. Therefore, we consolidate tags that are synonyms to create meta-tags that encompass many related words (e.g., “handwritten” and “cursive” would be in the same meta-tag). To this end, we create clusters of tags using a standard word embedding “Word2vec” (two-layer neural network) by mikolov2013efficient.

To measure the relevance between the image and word embeddings, we create clusters of the image and word embeddings, respectively, using $K$-means clustering.\footnote{To perform $K$-means clustering of font shape, we use the obtained 128-dimensional image embeddings directly clustered into 60 clusters using the elbow method. To perform $K$-means clustering of tags, we use 100-dimensional word embeddings that are first reduced to 10-dimensional embeddings and then clustered into 6 clusters using the elbow method. } Then, we measure how well the clusters of word embeddings match the clusters of image embeddings by using mutual information. Specifically, we use the normalized mutual information (NMI).

align*[align* omitted — 56 chars of source]

In this formula, $F$ is the distribution of clusters based on font image embeddings, $W$ is the distribution of clusters based on word embeddings, $H(\cdot)$ is information entropy, and $I(F,W)=H(F)-H(F|W)$ is the mutual information between $F$ and $W$. The NMI can be interpreted as how informative $W$ is in determining $F$. Its value ranges between $0$ and $1$ (the value $0$ implies that $W$ contains no information regarding $F$).

table[table omitted — 416 chars of source]

Table (ref) reports the values of NMI for two different pairs of distributions. The first row corresponds to the NMI for the image and word clusters described above. We obtain $NMI(F,W)=0.473$, which is quite promising. To compare this value with the baseline case, the second row of the table shows the NMI when the pre-existing product categories are used as structured attributes instead of the image embeddings. In general, when images were treated as unobserved product attributes, then the main structured attribute available in many design products is product categories. Recall that sans serif, serif, slab serif, display, handwritten, and script are such product categories defined by the font industry.\footnote{Product categories are generally available to analysts, but it is specific to this online marketplace that tags are observed. This motivates the setup of having tags as the ground truth in this section.} In this case, we obtain $NMI(F,W)=0.261$, which is roughly only half of the value in the first case. These results suggest that the learned image embeddings arguably capture economic agents' perceptions better than the structured attributes and can be relevant for economic analyses.

Effects of Merger on Design Differentiation

Background

On June 15, 2014, one of the major font foundries called FontFont was acquired by Monotype. At the time, Monotype sold fonts created by the foundries it owned as well as by third-party foundries. Before the merger, FontFont had sold its fonts through MyFonts.com as a third party. We study the causal effect of this merger on the change in the product differentiation decisions of the merging firm (i.e., FontFont foundry). As we show below, a favorable aspect of this market for merger analysis is that price seems to play little role in competition so that we can focus on product differentiation as the main response variable to merger. The major channel for product differentiation is the design of font shapes. This creative decision of a foundry may be affected by the merger through many different channels. By merging, the firms could increase efficiency by reducing transaction costs. If costs are reduced, the merged firm might find it profitable to create more experimental products. The merged firm may also be concerned with cannibalization---that is, competition among their own products---which would increase the diversity of product design. On the other hand, if Monotype has a preemptive motive with the merger, it will crowd its products to prevent entries. In the presence of these opposing factors, whether the merger increases the design diversity is an empirical question. We answer this question using the embeddings we created as the main ingredient of our statistical model.

Design Differentiation Measures

In this merger analysis, the outcome of interest is the degree of design differentiation. We construct two measures for design differentiation based on the constructed embeddings. The first measure is the distance to Averia. We calculate the Euclidean distance between an individual font and a benchmark font:\footnote{For each font embedding in our economic analyses, we use a representative embedding in a given family, namely the embedding of a regular style.} for image $x_{i}$ of font $i$,

align*[align* omitted — 74 chars of source]

where $f(\cdot)\in\mathbb{R}^{128}$ is the embedding and $f_{averia}$ is the embedding of a benchmark font called “Averia,” which is calculated by averaging the values of the embeddings of all the existing fonts in the market. The distance measure $D_{i}^{A}$ is intended to capture the degree of product differentiation of font $x_{i}$.

In addition to $D_{i}^{A}$, we consider the following gravity measure: for image $x_{i}$ of font $i$,

align*[align* omitted — 97 chars of source]

where the sum is for all other fonts $j$'s in the market. This measure effectively captures how font $i$ is located relative to other fonts $j$'s in the product space; it takes a large value when $i$ is individually far from all $j$'s and a small value when it is close to any of them. Compared to the distance to Averia, the gravity measure acknowledges the aspect of spatial competition among designers. For illustration, consider the interval $[0,1]$ as the product space. Suppose two existing fonts are located at $\{0,1\}$ in terms of their shapes, and the third font chooses to enter a location between one of $\{0,1/2,1\}$. The shape of the third font $i$ would be most differentiated in terms of $D_{i}^{G}$ if it is located at $\{1/2\}$. On the other hand, $i$ would be the most differentiated product in terms of $D_{i}^{A}$ if it is located at one of $\{0,1\}$ (even though these points are already populated).

Finally, we aggregate each measure for all fonts created by foundry $k$ in period $t$. In particular, for $j\in\{A,G\}$, we construct

align*[align* omitted — 91 chars of source]

where $I_{kt}$ is the set of all fonts created by foundry $k$ in period $t$. In the subsequent merger analysis, we consider both $\bar{D}_{kt}^{A}$ and $\bar{D}_{kt}^{G}$ as the outcome variables, henceforth referred to as the mean deviation and gravity measures, respectively. Despite the distinct feature, both $\bar{D}_{kt}^{A}$ and $\bar{D}_{kt}^{G}$ are meant to succinctly capture each foundry's creative decision of product differentiation in terms of font design every period. We show that the subsequent empirical analysis is robust to the choice of the measure between the two.

A Theoretical Example for Merger and Differentiation

We first consider a simple Hotelling-style model to illustrate that, in the market for fonts, the degree of product differentiation may be larger with merger than without while price may remain the same and consequently welfare may be greater with merger. The prediction from this theoretical model serves as the motivation of the subsequent empirical analysis. Consider two representative foundries each of which produces one font by locating its shape in a unit interval $[0,1]$. Foundry 1 chooses location $a$ and foundry 2 chooses location $1-b$. Based on tastes for fonts, consumers are located at $x$ distributed uniformly over $[0,1]$. Each consumer experiences aesthetic “transportation cost” $t$, which is incurred by “traveling” from her specific taste to an available font on the market. Consumers then pay $p_{1}$ for foundry 1's font and $p_{2}$ for foundry 2's font. The utility of consumer located at $x$ from buying font 1 is \[ V_{1}(x)=u-p_{1}-t|a-x|, \] where $u$ is a utility parameter, and the utility from buying font 2 is \[ V_{2}(x)=u-p_{2}-t|1-b-x|. \] Finally, consumers' utility if no product is purchased is normalized to be $V_{0}=0$. This model is slightly more general than what is considered in berry2001.

When travel cost $t$ is high relative to $u$, the firms in the Hotelling model may not compete with each other (economides1989). In this case, the model predicts (as shown below) that some consumers may not buy the product (i.e., $V_{0}\geq V_{1}(x)$ and $V_{0}\geq V_{2}(x)$ for some $x$) and each firm becomes a local monopoly, only selling to consumers with positive net surplus. We derive the equilibrium location and price under local monopoly. We start by finding foundry 1's profit function, which depends on the share of consumers who buy its product. To find this share, consider the consumer who is indifferent between buying and not buying font 1: $V_{1}(x)=0=V_{0}$. Let $x_{1}$ be the westmost location of a consumer who will buy font 1. Solving for $x_{1}<a$, we have $x_{1}=a-(u-p_{1})/t$. Let $x_{2}$ be the eastmost location of a consumer who will buy font 1. Solving for $x_{2}>a$, we have $x_{2}=a+(u-p_{1})/t$. Assume that the marginal cost is zero, which is plausible in this market. Then foundry 1's operating profit function of choosing $a$ and $p_{1}$ would be \[ \pi_{1}(a,p_{1})=p_{1}(a-x_{1})+p_{1}(x_{2}-a)=2p_{1}(u-p_{1})/t. \] Maximizing this profit yields the optimal price of $p_{1}=u/2$. Note that location $a$ and the price $p_{2}$ chosen by foundry 2 are not part of foundry 1's profit function (and analogously for foundry 2's profit), which reflects the fact that each firm is a local monopoly. As a result, there are multiple possible equilibrium locations for $a$ (and for $b$), subject to the fact that the share of consumers who buy font 1 cannot overlap with the share who want to buy font 2. We can analogously derive the optimal price for foundry 2, which yields $p_{2}=u/2$, the same as the optimal $p_{1}$.

Consider a concrete example of the model by assume $t=1$ and $u=1/3$. Further assume there is an fixed entry cost $F=1/20$. Then, the optimal prices are $p_{1}=p_{2}=1/6$ and $(1/3,2/3)$ is one of the equilibrium location. Under this solution, the consumer at the center of the space will not buy anything because $V_{1}(1/2)=V_{2}(1/2)=-1/3<V_{0}=0$. Therefore, the location and price are the equilibrium under local monopoly. However, this is an inefficient equilibrium as only $2/3$ of consumers buy the fonts; consumers $[1/6,1/2]$ buying font 1 and $[1/2,5/6]$ buying font 2. Still, neither firm has an incentive to deviate and introduce an new product that yields enough profit justifying the fixed cost.

Now consider a market where both foundries are merged under the same parameter values for $(t,u,F)$. The merged firm may introduce three products in the location $(1/6,1/2,5/6)$. To see why this is a local monopoly equilibrium, note that none of the products compete with each other; consumers from $[0,1/3]$ buy font 1, consumers from $[1/3,2/3]$ buy font 2, and consumers from $[2/3,1]$ buy font 3. Moreover, this is a unique and efficient equilibrium as the products serve all the consumers in the market. In terms of product differentiation, note that the maximum differentiation (i.e., the distance between furthest endpoints) is $2/3$, larger than $1/3$ without the merger. Similarly, in terms of our gravity measure, differentiation is approximately $-2.0$, larger than $-2.19$ without the merger. To see why the price remains the same, the merged firm's joint profit function involving all 3 products is $\pi_{1}(p_{1})+\pi_{2}(p_{2})+\pi_{3}(p_{3})$ where $\pi_{1}$, the profit for font 1, does not depend on the prices of fonts 2 and 3 and so on. Then, the optimal price would be $p_{1}=p_{2}=p_{3}=1/6$, the same as the optimal price without the merger. In this case, the profit is $1/6$ with the three products, compared to $1/9$ which is the aggregate profit of the two firms without the merger. Because more consumers are served with the merger while price remains the same, the welfare would be also greater with the merger than without.

For all other two-product equilibria without the merger that are close to $(1/3,2/3)$,\footnote{Specifically, they are equilibria where (approximately) $a\in[0.3157,1/3]$ and $b\in[2/3,1-0.3157]$.} a similar pattern is predicted: increased differentiation with the merger while price stays the same. When the locations move further away from $(1/3,2/3)$ toward $(1/6,5/6)$, then one of the firms may introduce another product, which results in a three-product equilibrium. An example would be $(1/6,1/2,5/6)$ and $p_{1}=p_{2}=p_{3}=1/6$, which are identical to the unique equilibrium location and price with the merger. For all other three-product equilibria without the merger,\footnote{They are equilibria where (approximately) $a\in[1/6,0.3157]$, $b\in[1-0.3157,5/6]$, and the third product locating at a point in $[0.3882,1/2]$ by one of the firms and the exact locations are dependent to one another.} differentiation increases with the merger and price may stay the same (for the single-product foundry) or decrease (for the two-product foundry). In all these other cases, since the price is no larger with the merger while the market is fully served, the welfare would be no smaller than without the merger.

To summarize the prediction under local monopoly, for a large range of equilibria, we expect increased differentiation and invariant price with merger than without. For all equilibria, there is a welfare gain with merger. The model's prediction is starkly different when there is no local monopoly. This occurs when $t<u$. In contrast to the prediction under local monopoly, differentiation falls and price increases with merger than without. Therefore, in this case, the welfare consequence is ambiguous.

Given the theoretical findings, the effects of merger on product differentiation and price (and consequently on welfare) in the market for the current design products remain an empirical question. Although this illustration is based on a simple stylized model, a similar intuition would continue to hold when there are more than two foundries and each foundry produces multiple fonts. Below, as our main focus of investigation, we empirically show that the responses of product differentiation and price in the real-world merger case of FontFont are in fact consistent with the prediction of the model under local monopoly.

Data and Empirical Strategy

For our empirical analysis, we construct the panel based on the dataset described in Section (ref). The cross-sectional units of this panel are foundries and the time dimension is a bi-annual time series between 2002 and 2017. Table (ref) shows the summary statistics for the variables (except the raw embeddings) in the panel. All the variables are constructed to be foundry-level. The mean deviation and gravity measures are the main response variables in the merger analysis. Glyph count is an important product specification and is considered to be related to quality.\footnote{Glyph count is the number of characters including special characters in each font family. We calculate the average glyph count for all fonts produced by each foundry each period. } The release frequency (i.e., maximum period between two releases), sales, number of orders, average price, and age of foundries introduced are the other control variables we use.

table[table omitted — 810 chars of source]

Section (ref) in the Appendix contains a simple descriptive analysis of the supply- and demand-side trends for the shapes of fonts captured in the embeddings. The supply-side trend plots the differentiation measures constructed in the previous subsection for fonts newly introduced in the market every period. For the demand-side trend, we construct and plot analogous measures for fonts purchased every period. We find that, on average, newer fonts tend to be more differentiated than older fonts. How such differentiation decisions are affected by the change in market structure is the subject of the merger analysis. We also find that the demand-side trend is markedly stable, especially around the time of merger.\footnote{In general, the stable demand-side trend is consistent with the industry norm. According to the industry experts we interviewed, font markets tend not to experience seasonality or short-term trends in consumer preferences, unlike in other design industries such as clothing. } This supports the argument that the change in the producer behavior we find below in the merger analysis is not demand-driven.

To estimate the causal effect of merger on product differentiation, we find a comparable group that would behave similarly as the treated foundry (i.e., FontFont) if it were not for the merger. There are challenges in using this approach. First, it is difficult to find a single untreated foundry that resembles the treated foundry in terms of the product differentiation measure (i.e., the mean deviation or gravity measure). A naive average within a control group would be a poor candidate. Second, candidate foundries for the control group whose products are too similar to the treated foundry's products may be in direct competition with the treated, which can create strategic spillovers. To overcome these challenges, we use the synthetic control method (abadie2003economic, abadie2010synthetic) and the proposed differentiation measures. First, the synthetic control method compares the treated unit with a “synthesized control unit” obtained from a weighted average of the control group. Here, the weights are estimated by minimizing the distance between the observed characteristics including the outcome (as presented in Table (ref)) of the treated unit and the weighted average of the characteristics of the control group. For the control group that serves as the basis for constructing a synthetic control, we use foundries whose merger status has not changed during the study period. Second, although we achieve a certain level of similarity in the differentiation measure between the resulting synthetic control and the treated foundry, this does not necessarily imply that they are competitors producing similar products. This is because the value of the measure of each control unit can be significantly different from that of the treated unit, even if the weighted average is close to the latter. In fact, this appears to be the case in our data as shown in Figures (ref) and (ref) below. Therefore, we assume that strategic spillovers between the treated unit and the control units are not substantial.\footnote{See abadie2021using for discussions on similar approaches to minimize strategic spillovers between treated and control units.} Nonetheless, the advantage of our setting is that the embeddings contain rich information about the design attributes of the fonts created by the foundries. Therefore, we can reliably create a comparable synthetic unit from the control group without their products being necessarily visually similar to those of the treated foundry. In total, our data include one treated foundry (FontFont) and 50 control foundries.

The Effects of Merger

figure[figure omitted — 665 chars of source]

Figure (ref) captures the main results of our causal analysis of merger using the gravity measure.\footnote{To reduce the scale, we take the logarithm of the positive part of $\bar{D}_{kt}^{G}$ and then put back the minus sign.} The results of the analysis using the mean deviation measure are presented in the Appendix. The solid line in the left panel presents the trend of the gravity measure of the fonts newly designed by FontFont in a given period. Around the period where FontFont was merged, which is indicated by a vertical line, the shape of FontFont's fonts substantially differs from that of other fonts in the market. This before and after comparison alone cannot yield the causal effect of the merger, as other market conditions may have changed around this period. Comparing this trend with the trend of the synthetic FontFont, depicted by a dashed line, removes possible confounding factors. First, the two trends before the merger appear to be close to each other by construction. After the merger, however, FontFont tends to produce more experimental fonts (i.e., fonts that are far from others) relative to the trend of the synthetic control. We also confirm that even if we backdate the period of acquisition (i.e., does not use the information of the timing of merger), the relative increase still occurs around the same period as the vertical line.

To understand the virtue of our synthetic control method, we contrast the left panel with the right panel in Figure (ref). The latter depicts the trends of the treated unit and the naive average of the control group. Inspecting the pre-treatment period in the right panel, the naive average fails to mimic the trend of the treated unit.

table[table omitted — 439 chars of source]

Finally, Table (ref) reports the treatment effects averaged over each year after the merger. We also report their $p$-values using the permutation test by chernozhukov2017exact. The advantage of this inferential method in our context is that it does not assume random assignment of the policy intervention unlike earlier methods that rely on the assumption to ensure the properties of randomization tests (e.g., abadie2010synthetic).\footnote{See arkhangelsky2019synthetic for related discussions.} Consistent with Figure (ref) (left), the short-run effects in the first and second years are statistically significant. To understand the magnitude of the effect, note that the standard deviation (SD) of the gravity measure in the sample is 0.16 as shown in Table (ref). Therefore, the treatment effects are on average half of 1 SD. It is also informative to understand the increase in the measure (roughly from -9.8 to -9.6) relative to the distribution of the gravity measure (the left panel of Figure (ref) in the Appendix). After two years of strong positive effects, the effect of merger dissipates in the third year. The placebo test results, presented in Table (ref) in the Appendix, show that the treatment effects are not statistically significant before the merger.

figure[figure omitted — 597 chars of source]

We produce analogous results using the mean deviation measure instead of the gravity measure and find that the results are qualitatively similar. The effects are positive and statistically significant and on average larger than 1 SD; see Section (ref) in the Appendix. This suggests that our findings are robust to the choice of the measure of product differentiation as long as it is constructed based on the image embeddings. This robustness disappears if we use more traditional measures for product offerings, such as the number of products and specifications. Figure (ref) shows that the merger has no effects on both the Glyph counts (which is the key specification for fonts) and the number of new fonts. Tables (ref) and (ref) in the Appendix confirm that the effects are not statistically significant.

Based on this analysis, we conclude that the merger caused FontFont to explore a new territory of the product space, at least temporarily.\footnote{Given our data frequency, the foundry's immediate increase in design differentiation after the merger is feasible because it typically takes one or two months for a foundry to design a font. } That is, by being part of the parent organization, Monotype, FontFont increased the visual variety in font design. Before the merger, foundries owned by Monotype produced fonts and sold them on MyFonts.com. After the merger, Monotype may have incentives to diversify the product scope owing to the increased size and efficiency of the firm. It may also be the case that Monotype tries to avoid cannibalization by spreading apart its products, thus reducing competition amongst their own foundries. Such tendency disappears after two years, perhaps because the firm has either successfully foreclosed the market or found the strategy unprofitable.\footnote{According to the interview with the employees, we discovered that the company has gone through a structural change two years after the merger that is consistent with the second explanation. The details of the change cannot be publicly revealed.}

The explanation of our empirical findings echos the prediction under local monopoly in the simple theoretical model above. Indeed, the market for fonts may have the aspect of local monopoly in which consumers exhibit high travel cost. This seems plausible because most of the consumers in this market are professional designers with sophisticated tastes for font shapes. The theoretical model also predicts that the price response to merger would be minimal under local monopoly for a large range of possible equilibria. Figure (ref) corroborates this theoretical prediction. The figure plots the average prices of newly introduced fonts by FontFont and the top ten control foundries.\footnote{The range of periods differs from our main analysis due to data limitations.} The average prices of most of the foundries stay relatively stable, especially around the time of merger.

figure[figure omitted — 452 chars of source]

Conclusion

Certain policy questions are better answered by treating high-dimensional, unstructured attributes as observables and attempting to solve the resulting dimensionality problem. We propose to quantify the design-oriented attributes of a product by constructing a low-dimensional product space of these attributes. We use a modern convolutional neural network to accomplish this task. We find that these attributes are correlated with font designers' and consumers' perceptions, as reflected in the mutual information between tags used to describe the fonts and the neural network embeddings. We then turn to a causal analysis to understand the effects of a merger on product differentiation. We find that the merger increases the product variety in this market. The analytical framework of this paper can be applied to various products where unstructured attributes are present.

This paper motivates interesting directions for future research, some of which may require a structural approach. For structural economic analyses, we envision a two-step approach of using the embeddings: construct a neural network embedding to reduce the dimensions of the product image and then include the embeddings in structural economic models. Although we could use the neural network to directly predict economic variables, such as demand and price, the neural network is not a causal model. Therefore, its counterfactual predictions are not as credible as those of a more traditional structural model. An alternative structural approach would be to incorporate dimension reduction as part of the structural estimation (chernozhukov2018double, foster2019orthogonal).

Using these structural approaches, we may answer economic questions related to this market. Related to the economic analyses in this paper, we may view product differentiation as spatial competition. There is a close parallel between spatial differentiation and the neural network embedding product space we constructed. Analogous to siem2006spatial, who studies the location choice of video rental stores, we can investigate the relationship between design choices and the local market power of the designers. We can also study how third-party and in-house foundries differ in their product differentiation decisions. One relevant policy question is the effect of the commission fee of third parties. In fact, Monotype changed the commission policy during the data collection period, which may serve as a key policy variation.

Alternatively, we may view product differentiation as intellectual property. Agents in this industry are subject to license agreements, which aim to protect the originality of font shapes. Heuristically, this policy states that “one cannot produce fonts which shapes are substantially similar to existing fonts.” Given the product space we characterized, one can interpret this policy as imposing a ball centered around each font, thus preventing other productions within: producing another font inside the ball is considered a violation. Then, one can ask what the welfare maximizing level of the policy (i.e., the optimal radius of the ball) would be and whether the current level in the market is optimal or suboptimal.

appendix

Neural Network Training

Details of Network Training

The training iteratively improves the parameters of the network using batches to estimate the gradient and then update the parameters accordingly (wilson2003general). Each batch contains 270 cropped images of fonts, or equivalently, 90 triplets. We cropped each image based on the number of characters in the image.\footnote{For example, the first half of the pangram sentence has 20 characters; therefore, to crop 5 characters, we would take 40 percent of the pixels in the first half of the image. We tried crops with 3, 4, 6, and 7 different characters. We also tried different cropping schemes such as using the white space between characters.} As the gradient is evaluated at more batches, the parameters in the network are adjusted. The number of trainable parameters are $90,000=3\times(3\times100\times100)$ (layers $\times$ input size). The training of the network is completed when the loss function reaches below 0.7. The learning rate is initially 0.05, and then lower to 0.01 to finalized the model. Here are the remaining hyper-parameter values: the batch size is 90, epoch size 500, weight decay $1\times10^{-4}$, and the margin $\alpha$ is 0.2. The training time took approximately 24 hours with 4 GPUs (Nvidia 1080-TI).

Evaluation of Trained Network

To evaluate the neural network embeddings, we create test sets of triplets that the neural network has never seen during the training. Using the test sets, the task is to identify whether the image in a triplet is a positive or a negative. It is classified as a positive if its Euclidean distance from the anchor in the trained embedding space is less than a pre-defined threshold (via cross validation) and a negative otherwise. We describe the results of tests on two different test sets. The first test set, “easy,” randomly samples pairs of a positive crop within the same family as an anchor and a negative crop from different families, for which the performance is evaluated. The second test set, “hard,” samples pairs of a positive crop within the same style of a family as an anchor and a negative crop from different styles within the same family.

The analysis involves true positives (TP), true negatives (TN), false positives (FP), and false negatives (FN). Table (ref) presents the accuracy and validation rates for the neural network in each of these test sets with different crop sizes. The accuracy is a measure of how well the neural network embeddings perform in the classification task. Overall, it gets about 90 percent of the triplets correct.

We also analyze the trade-off between Type-I and Type-II errors in classification. The validation rate is the true acceptance rate and is calculated as TP/(TP+FN). The false acceptance rate (FAR) is given by FP/(TN+FP). In addition, we plot precision and recall curves. Precision is related to the Type-I errors and is calculated as TP/(TP+FP). Recall is related to the Type-II errors and is given by the formula TP/(TP+FN). The precision recall curve in Figure (ref) shows the trade-off between these types of errors. There exists a steeper trade-off between precision and recall in the hard test set.

table[table omitted — 639 chars of source]
figure[figure omitted — 341 chars of source]

Overall, all the statistics we report here show that our neural network performs well. The low validation rate in the hard test set suggests that there is room for improvement in differentiating between different styles.

Trend Analyses

To further illustrate the usefulness of the embeddings, we analyze the supply- and demand-side trends in font style. This also provides an additional background for the causal analysis of the merger in Section (ref). On the supply side there are constant entries of new products in the marketplace. On the demand side, on average, more than 1,000 fonts are (stably) sold per day. Of course, demand and supply are endogenous, so this analysis is only a descriptive analysis. We, however, believe it reveals some interesting patterns in this market.

Supply-Side Trend

figure[figure omitted — 268 chars of source]

We analyze the supply-side trend in the Euclidean distance from Averia. We are particularly interested in how the style of fonts newly entering the market differs from the style of incumbent fonts. In Figure (ref), each dot represents the daily mean distance from Averia for a range of periods. The red horizontal line is the mean distance of all fonts introduced before 2001, which we view as incumbents. The black line depicts the quarterly moving averages of the daily mean for fonts newly introduced since 2001 each day, which we view as entrants. Overall, we find that entrants have font shapes that are more experimental or innovative than those of incumbents, possibly to avoid competition and establish market power distant from the incumbents in the product space.

Demand-Side Trend

figure[figure omitted — 605 chars of source]

We now analyze the trend in the Euclidean distance between Averia and fonts purchased by consumers between 2014 and 2017.\footnote{The time span is shorter than that for the supply-side analysis due to the missing license type data for earlier periods.} To remove the supply-driven factor, we condition on price and focus on the price ranges of USD 25--35 and USD 35--45, where promotions are rare. These price ranges are the most common in the market. Figure (ref) plots the trend conditional on the price being USD 25--35 and USD 35--45. We also plot the trend for the two major license types: desktop license and web font license. Again, each dot represents the daily mean distance weighted by the number of purchases. We also plot monthly moving averages.

Interestingly, in both figures, we find that fonts sold under desktop license are more experimental (or less conservative) than fonts sold under web font license. Because both licenses are offered for most fonts, this difference cannot be attributed to the supply-side decision but is the result of consumer decisions.

Summary of Findings

Based on our analyses, we find the following stylized facts. (i) Entrants are more innovative than incumbents. (ii) Consumer preferences are stable during 2014, the year that witnessed the merger of our interest.\footnote{This feature is uniformly found in all price ranges, but we do not report the results for succinctness.} Based on this finding, we assume that the demand side has a negligible influence on the change in product differentiation decisions by the merging firm around the time of merger. (iii) Consumers have different preferences over shapes depending on the license type they purchase. Consumers prefer more conservative shapes for web font licenses and more experimental shapes for desktop licenses. Usually, web font licenses are used on webpages, where legibility is important, whereas desktop licenses are used in printed material (e.g., posters, cards), where designers (as consumers) have more control over the design environment.