EconBase
← Back to paper

Content vs. Form: What Drives the Writing Score Gap Across Socioeconomic Backgrounds? A Generated Panel Approach

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

107,041 characters · 18 sections · 23 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Content vs.\ Form: What Drives the Writing Score Gap Across Socioeconomic Backgrounds? A Generated Panel Approach

abstractStudents from different socioeconomic backgrounds exhibit persistent gaps in test scores, gaps that can translate into unequal educational and labor-market outcomes later in life. In many assessments, performance reflects not only what students know, but also how effectively they can communicate that knowledge. This distinction is especially salient in writing assessments, where scores jointly reward the substance of students’ ideas and the way those ideas are expressed. As a result, observed score gaps may conflate differences in underlying content with differences in expressive skill. A central question, therefore, is how much of the socioeconomic-status (SES) gap in scores is driven by differences in what students say versus how they say it. We study this question using a large corpus of persuasive essays written by U.S. middle- and high-school students. We introduce a new measurement strategy that separates content from style by leveraging large language models to generate multiple stylistic variants of each essay. These rewrites preserve the underlying arguments while systematically altering surface expression, creating a “generated panel” that introduces controlled within-essay variation in style. This approach allows us to decompose SES gaps in writing scores into contributions from content and style. We find an SES gap of 0.67 points on a 1–6 scale. Approximately 69% of the gap is attributable to differences in essay content quality, Style differences account for 26% of the gap, and differences in evaluation standards across SES groups account for the remaining 5%. These patterns seems stable across demographic subgroups and writing tasks. More broadly, our approach shows how large language models can be used to generate controlled variation in observational data, enabling researchers to isolate and quantify the contributions of otherwise entangled factors.

\epigraph{Style is the dress of thoughts; and a well-dressed thought, like a well-dressed man, appears to great advantage.} {Philip Stanhope, Earl of Chesterfield, Letters to His Son, Letter CXXVIII, January 21. 1751}

Introduction

For many of us, generating an idea and communicating it are distinct processes. Even when an idea feels clear internally, expressing it in writing can be difficult, particularly in settings where the writing is judged. We may have a good idea and a coherent answer, but we might struggle to express it in a manner that effectively communicates our reasoning and addresses all relevant points.

The ability to form ideas and clearly articulate them in writing are skills that develop over time, shaped by educational background, access to resources, and other socioeconomic circumstances. Differences in these factors can create systematic disparities in how students from different social groups tackle questions and express their thoughts on paper. Given that substantial grade gaps exist across students from different socioeconomic backgrounds hanushek2019achievement,morgan2024explaining,reardon2011widening, this naturally raises a question: do these gaps primarily reflect differences in the content of what students write, or do they stem from differences in how these ideas are expressed?

From the perspective of educational fairness, a central question emerges: to what extent do observed score differences reflect differences in the substance of students' ideas—their reasoning, argumentation, and use of evidence—versus differences in their ability to express those ideas in forms that evaluators reward? If score gaps are driven primarily by differences in expression that stem from unequal access to particular linguistic resources or familiarity with academic conventions, then interventions focused on writing style and mechanics might substantially reduce disparities. Conversely, if gaps reflect deeper differences in analytical skills and content knowledge, then addressing them will require more fundamental changes in curriculum and instruction. Distinguishing between these possibilities is essential for designing equitable and effective educational policy.

To answer these questions, we need to separate differences in "what is said" from differences in "how it is said"—to decompose the observed score gap into components reflecting content and style. But doing so poses substantial measurement challenges. First, content and style are not directly observable as separate quantities.irst, content and style are not directly observable as separate quantities. Ideas are expressed through words, which makes it hard to measure the role of ideas separately from the words that express them. Second, even if we had separate measures of content and style quality, the two are deeply entangled in real writing. Students who develop stronger analytical skills typically also develop more sophisticated expression. Those with access to better educational resources acquire both richer vocabularies and more complex reasoning strategies. Conversely, students who face linguistic barriers or whose home discourse patterns differ from academic conventions may struggle to express sophisticated ideas in forms that evaluators recognize as "clear" or "well-organized." This entanglement means that content and style do not vary independently, making it difficult to assess their separate contributions to observed score differences.

Traditional empirical approaches struggle with these challenges. One might attempt to control for observable stylistic features—sentence length, vocabulary sophistication, grammatical correctness—and attribute the remaining variation to content. But such controls rest on strong assumptions about which features of writing reflect "how" versus "what," and risk either absorbing content-related variation into the style measures (if the controls are too comprehensive) or leaving substantial style-related variation in the residual (if the controls are too limited). Alternatively, one might ask human coders to rate content and style separately. But human judgments are inherently subjective, difficult to standardize across many essays, and expensive to scale. Moreover, raters often cannot cleanly separate substance from presentation: when ideas are expressed poorly, raters may perceive them as weaker or less developed than they actually are, confounding the two dimensions.

We propose a novel measurement strategy that leverages recent advances in large language models. Rather than attempting to separate content and style through statistical controls or subjective coding, we generate multiple stylistic variants of each essay using an LLM. Specifically, for each original essay, we produce several rewrites designed to alter surface expression—word choice, sentence structure, grammatical polish—while preserving the underlying arguments, evidence, and logical structure. These rewrites create a "generated panel" data in which the same substantive content appears in different stylistic realizations.

This panel structure allows us to decompose the observed score gap into three components: differences in content (what students argue), differences in style (how they express those arguments), and differences in scoring functions (how essays are evaluated across groups). By examining how scores vary across rewrites of the same essay, we can isolate the component of the score that responds to stylistic variation while holding content fixed. Conversely, by averaging over rewrites to eliminate within-essay stylistic noise, we can recover a measure of each essay's content that is comparable across students, even when their original writing styles differ systematically.

To operationalize this approach, we first estimate group-specific scoring functions from the human-assigned scores in our data. We train separate predictive models for high- and low-SES students using rich features that capture both semantic content (through text embeddings) and stylistic properties (through large set of linguistic variables measuring syntactic complexity, lexical diversity, and cohesion). These learned scoring functions represent how essays are evaluated under current practices for each group, allowing us to predict scores for any essay—including the LLM-generated rewrites, which were not seen by human graders. Combined with the rewrite panel, these predicted scores enable us to separately identify the contributions of content, style, and scoring rules to the observed gap. As the rewrites preserve content and induce a common shift in the style component—this design permits a clean separation of these three components.

We apply this method to a large dataset of persuasive essays written by U.S. middle- and high-school students, linked to measures of socioeconomic status. On average, high-SES students score 0.67 points higher than low-SES students on a 1–6 scale. Our decomposition reveals that approximately 69% of this gap is associated with differences in content, 26% with differences in style, and the remaining 5% with differences in scoring functions across groups.

This pattern is stable across demographic subgroups and most writing prompts. Among both white and non-white students, and among both males and females, content differences account for roughly two-thirds to three-quarters of the score gap. The main systematic variation appears across grade levels: in higher grades, the share of the gap attributed to style increases, rising from around 25% in grade 6 to over 40% in grade 11. This shift may reflect increasing differentiation in students’ stylistic choices in later grades, either because students specialize in different styles of writing or because some students are more successful at adapting to the highly rewarded ‘school’ style.

These findings have several implications for how we understand SES disparities in writing assessment. First, they suggest that observed score gaps are driven primarily by differences in the substance of student writing—the quality of arguments, the coherence of reasoning, the relevance and use of evidence—rather than by superficial features of presentation. To the extent that these substantive differences reflect real inequalities in critical thinking, analytical training, and access to rich curricular content, addressing the gap will require more than teaching students to write grammatically correct sentences or format their essays according to conventions. It will require attention to the deeper educational experiences that shape students' capacity to construct and organize compelling arguments.

At the same time, the non-trivial contribution of style—roughly one-quarter of the gap—underscores that mastery of academic discourse conventions remains an important and consequential dimension of writing performance. Students who control the lexical, syntactic, and organizational features valued in school writing receive higher scores, even when their underlying ideas are comparable. To the extent that these conventions are themselves socially patterned—acquired more readily by students whose home language use aligns with school norms, or whose access to instruction includes explicit attention to register and style—disparities in expression can amplify differences in opportunity and reinforce existing hierarchies.

Our paper also speaks to an emerging policy question: what would happen to score gaps if all students had access to LLM-based writing tools that could "polish" their essays before submission? Our results suggest that this would not close the gap: rewriting low-SES essays with an LLM prompt raises their average predicted scores, but a substantial disparity remains because the content component dominates. Most of the score difference reflects what students argue and how they structure and develop their reasoning. This matters for equity debates about AI writing assistance. If stylistic conventions unfairly penalize students’ ideas, polishing tools could reduce one source of disadvantage; but if they mainly shift style while leaving content differences unchanged, their impact on overall gaps will be limited, and unequal access could introduce new advantages. More fundamentally, treating technological standardization of expression as an equity fix risks obscuring the underlying inequalities in content development and reinforcing a narrow, standardized conception of what counts as “good” writing.

Beyond the specific findings on writing assessment, this paper makes two broader contributions. Methodologically, it demonstrates how large language models can function as research instruments, not just as end-user tools for text generation, but as devices for creating controlled variation in observational data. The generated panel design we introduce is applicable to any setting in which substance and presentation are confounded and researchers need to assess their separate contributions. Potential applications include studying how the presentation of job market papers affects hiring decisions, how framing shapes persuasive communication, or how linguistic style influences the reception of creative and professional work.

\paragraph{Related Literature} It's well-documented that essay-scoring differences across social groups are significant and persistent nationsreport2012. A large education and applied linguistics literature studies why these gaps arise, emphasizing that performance in school writing reflects not only ideas and reasoning but also mastery of academic language and discourse conventions—features that are unevenly distributed across socioeconomic backgrounds and schooling contexts. This perspective motivates our focus on separating “what is said” from “how it is expressed,” and connects to work that links linguistic features such as syntactic complexity, cohesion, lexical sophistication, and organization to human-assigned writing scores wang2023multi. At the same time, the automated essay scoring (AES) literature documents that scoring models can place substantial weight on surface and stylistic properties, raising concerns about construct validity and about whether models reward form in ways that can amplify pre-existing inequalities (e.g., if academic register and polish are themselves socially patterned). Closely related, research on fairness in automated scoring evaluates whether scoring algorithms treat groups differently and shows that conclusions can depend on the fairness criterion used litman2021fairness.

Methodologically, our paper builds on two recent strands. First, it draws on the growing “text-as-data” tradition in economics and political science, which treats text as a high-dimensional object that can be represented, summarized, and manipulated to study substantive questions. Our contribution fits this agenda by using text representations to construct interpretable “content” and “style” components and by using controlled textual edits to form counterfactuals. Second, it leverages recent progress in large language models (LLMs) as tools for generating rewrite panels: for a fixed underlying essay, we create multiple stylistic variants intended to preserve meaning while shifting expression along targeted dimensions. This design provides a new way to probe which aspects of writing are rewarded by scorers, and to do so at scale while holding the underlying message approximately fixed. In this sense, LLM-based rewriting functions like a structured perturbation device that complements purely observational correlations between linguistic features and scores.

Our approach to measuring disparities also aligns with recent trends in economics and computer science that define disparities conditional on a variable of interest dwork2012fairness, kusner2017counterfactual. For instance, arnold2022measuring examines disparities in judicial treatment by comparing bail decisions among individuals with similar predicted risk, and bohren2022systemic studies group gaps in employment probabilities among workers of equal productivity. Such approaches highlight the normative aspect of measuring disparities across social groups by specifying which attributes should be held fixed. We adapt this logic to writing by treating the essay’s underlying content as the “attribute to be held fixed,” and then quantifying how much of the remaining score gap is attributable to non content factors. At the same time, our goal is not only to document differential impacts but also to identify the mechanisms that generate them. This objective connects to the decomposition tradition in economics, which explains observed gaps by constructing counterfactual distributions—fixing one component and modifying the relevant conditional distribution to isolate its contribution fortin2011decomposition.

Finally, our design resonates with the explainable AI literature, which seeks to understand algorithmic predictions by modifying inputs and examining resulting changes in outputs. In text analysis, prior work has proposed targeted edits (e.g., flipping sentiment-bearing adjectives) to diagnose model behavior feder2021causalm, vig2020investigating. Our use of rewriting is conceptually similar but more conservative in its claims: we treat the edits as a way to construct counterfactual score distributions and to decompose gaps, rather than as a causal estimate absent additional identifying assumptions (e.g., random assignment of essays to graders). Related input-modification approaches have also been used outside text. ludwig2023machine, for instance, studies why defendant photos predict bail decisions by generating minimally perturbed images that induce large changes in predicted release probabilities, and then interpreting the differences. Analogously, our rewrite-based counterfactuals illuminate which dimensions of writing are most strongly rewarded by scoring systems, and how those rewards map into socioeconomic score gaps.

\paragraph{Roadmap.} The remainder of the paper proceeds as follows. Section 2 develops the empirical framework and identification strategy. Section 3 describes the data and variable construction. Section 4 presents descriptive statistics and assesses the performance of our scoring models and rewrite procedure. Section 5 reports the decomposition results and robusntess. Section 6 concludes.

Empirical Framework

We observe essays indexed by $i=1,\dots,N$. Let $G_i\in\{H,L\}$ denote group membership of essay $i$'s author (e.g SES). For each scoring concept or ranker $m\in\{H,L\}$, a deterministic scoring function $S^{(m)}$ maps a rendered text $x_{i}$ to a numeric score: \[ s^{(m)}_{i} \;=\; S^{(m)}(x_{i}) \in \mathbb{R}, \] where $x_{i}$ collects the full set of observed essay and author characteristics, including the text embedding and a rich set of style variables. In the framework below we treat $S^{(m)}$ as given. In the empirical implementation, $S^{(m)}$ is not observed directly but is estimated from the observed human scores. The predicted values from these models serve as estimates of $S^{(m)}(x_{i})$.

Write $C_i$ for the (latent) content of essay $i$. Let $R_{i}$ denote the phrasing/style realization of essay $i$. We make the following assumption on the scoring function.

assumption[Separable content and style] For any scorer $m$ the essay scoring is separable in content and style: \begin{equation} s^{(m)}_{i} = \underbrace{\theta_m(C_i)}_{Content index under $m$} + \underbrace{\rho_m(R_{i})}_{Style component under $m$}. \end{equation}

This assumption on the ranker implies that scoring varies additively with the content and with the style, where $\theta_m$ and $\rho_m$ are mappings from content and style, respectively, to the score under $m$. Notice that we allow for different rankers to weight content and style differently. For example, rankers for one group of students may be more responsive to style and less responsive to the actual content of the answer, while another may be more sensitive to the content. This ranking function can also allow for difference in the overall ranking by simply giving lower scores to the content and style, capturing cases of discrimination or differences in who is assigned to rank low SES students vs. High SES students.

Measuring What Drives The Gap in Scores

We are interested to measure how much of the gap in scores between two groups is driven by difference in content and how much the gap is driven by difference in writing style. To do so we use Kitagawa--Oaxaca--Blinder decomposition approach (Kitagawa1955, Oaxaca1973, Blinder1973). Suppose the policy grades H originals with $S^{(H)}$ and L originals with $S^{(L)}$. The observed gap is \[ \Delta_{\mathrm{obs}} := \mathbb{E}_H[s^{(H)}_{i}] - \mathbb{E}_L[s^{(L)}_{i}]. \] We suggest decomposing the gap as follows. Let the group means \[ A^{(m)}_G:=\mathbb{E}_G[\theta_m(C_i)], \quad U^{(m)}_G:=\mathbb{E}_G[\rho_m(R_{i})], \] The following identity holds:

equation[equation omitted — 497 chars of source]

The first bracket captures how much of the score gap is driven by differences in content quality across the two groups, when essays are ranked by the high-SES scorer. These differences reflect what students choose to argue about and how they structure their reasoning. For example, some students may present a well-developed causal argument with clear evidence, while others provide only loosely connected claims. Whether an argument is convincing, rational, and relevant—all of these contribute to this component.

The second bracket measures how much of the gap arises from differences in writing style, again evaluated through the lens of the high-SES scorer. This component reflects control of language, correctness of word choice, coherence, and other stylistic decisions students make in expressing their ideas. For instance, two students may present equally strong arguments, but one might write in a more polished and grammatically precise style. The gap here reflects how the scoring rule—estimated from high-SES writing—responds to stylistic features more common among low-SES writers.

Finally, the third component captures how much of the score difference is attributable to the rankers themselves, holding the underlying texts fixed. Even when reading the exact same set of essays, rankers functions may differ systematically in how they assign scores. For example, a low-SES student might receive a lower score simply because one scorer is more severe or holds implicit biases, or because high-SES and low-SES students happen to be assigned different rankers with different grading tendencies.

In our discussion, we mostly focus on their relative sizes—comparing the share of the writing-score gap explained by each component, rather than on the absolute sizes.

Identification

Equation (ref) is an accounting identity, but its components are not directly observed. The objects \[ A_G^{(m)}=\mathbb E_G[\theta_m(C_i)] \qquad\text{and}\qquad U_G^{(m)}=\mathbb E_G[\rho_m(R_{i})] \] involve the latent content index $\theta_m(C_i)$ and the latent style component $\rho_m(R_{i})$. Even under separability (ref), these two primitives are not separately pinned down from a single observed score, because the decomposition is only defined up to an additive constant (shifting $\theta_m$ by a constant and shifting $\rho_m$ by the opposite constant leaves $s_{i}^{(m)}$ unchanged). Identification therefore requires an anchoring restriction that uses the within-essay panel of rewrites to create a common reference style environment.

Let $k\in\{0,1,\dots,K\}$ index versions of essay $i$, where $k=0$ denotes the original and $k\in\mathcal K:=\{1,\dots,K\}$ indexes a fixed set of $K$ LLM-generated rewrites. Each rewrite is intended to preserve the semantic content of essay $i$ while altering its stylistic realization. Let $C_{ik}$ denote the content of version k of essay $i$, $R_{ik}$ its realized style, and $x_{ik}$ its rendered text. Let $s_{ik}^{(m)} = S^{(m)}(x_{ik})$ denote the score assigned to version $k$ of essay $i$ by scorer $m$. We make the following assumptions on the rewrites.

assumption[Rewrite content fidelity] For all $k\in\mathcal K$, the rewrite preserves content: \[ C_{ik}=C_{i0} := C_i. \]

Assumption (ref) formalizes the design goal that rewrites change the realization $R_{ik}$ while leaving latent content $C_i$ fixed, so that within-essay variation across $k\in\mathcal K$ can be attributed to style rather than content.

Our second assumption characterizes how the LLM rewrites enter the style component.

assumption[Rewrite style homogeneity] For each ranker $m$ and rewrite $k\in\mathcal K$, the effect of rewrite $k$ on the score is an additive constant $\lambda_{k,m}$: \[ \rho_m(R_{ik}) = \lambda_{k,m} + u_{ikm}, \] where the residual satisfies \[ \mathbb E[u_{ikm}\mid C_i,G_i] = \mathbb E[u_{ikm}\mid C_i]=0. \]

Assumption (ref) says that, conditional on the essay's content $C_i$, the effect of asking the LLM for rewrite type $k$ on the style component under ranker $m$ is a constant shift $\lambda_{k,m}$, up to idiosyncratic noise $u_{ikm}$ with mean zero. The second part rules out systematic differences in the rewrite noise across groups, conditional on content. This implies that for two essays with identical content—one from each group—the conditional mean of $u_{ikm}$ is the same for both essays, and equals zero.

Assumption 3 has a testable implication. It implies that rewrite type $k$ acts like a group-invariant vertical shift in scores (up to mean-zero noise conditional on content). In Section 4.3 we evaluate this implication directly using the rewrite panel: for each pair of rewrite types \(k,k^{\prime}\) we compute within-group mean score differences and then a difference-in-differences across SES, which should be approximately zero under additivity. The resulting “Difference-in-difference matrix” should be close to zero relative to the magnitude of rewrite-induced score changes.

Given Assumptions (ref)–(ref), the content gap can be identified as follows. Define the rewrite-averaged score \[ \bar{s}^{(m)}_i := \frac{1}{K} \sum_{k\in\mathcal K} s^{(m)}_{ik}. \] Using (ref) and Assumptions (ref)–(ref) we obtain

align*[align* omitted — 245 chars of source]

since $\mathbb E[u_{ikm}\mid C_i,G_i]=0$ by Assumption (ref). The constant $\bar\lambda_m$ does not depend on group membership. Therefore the difference in rewrite-averaged scores between groups exactly recovers the content gap under ranker $m$: \[ A_H^{(m)} - A_L^{(m)} = \mathbb E_H[\theta_m(C_i)] - \mathbb E_L[\theta_m(C_i)] = \mathbb E_H[\bar{s}^{(m)}_i] - \mathbb E_L[\bar{s}^{(m)}_i]. \] The rewrites thus serve to equalize the distribution of writing style across essays, allowing us to isolate differences in the way content is scored.

To identify the style component, define the essay-level deviation between the original score and the rewrite-averaged score under ranker $m$: \[ d^{(m)}_i := s^{(m)}_{i0} - \bar{s}^{(m)}_i. \] Using (ref) and Assumptions (ref)–(ref) again,

align*[align* omitted — 292 chars of source]

Taking expectations and using $\mathbb E[u_{ikm}\mid C_i,G_i]=0$ yields \[ \mathbb E_G[d^{(m)}_i] = \mathbb E_G[\rho_m(R_{i0})] - \bar\lambda_m. \] The constant $-\bar\lambda_m$ cancels when we take differences across groups, so that \[ U_H^{(m)} - U_L^{(m)} = \mathbb E_H[\rho_m(R_{i0})] - \mathbb E_L[\rho_m(R_{i0})] = \mathbb E_H[d^{(m)}_i] - \mathbb E_L[d^{(m)}_i]. \]

In words, subtracting the rewrite-averaged score from the original score nets out differences in content and the average shift induced by the rewrite procedure. The remaining group difference reflects how the scoring rule responds to the styles actually used by the two groups relative to the common benchmark style implemented by the LLM rewrites.

Finally, the ranker-tilting term in (ref) is directly observed as \[ \mathbb E_L\big[s_{i0}^{(H)} - s_{i0}^{(L)}\big], \] which compares how the two rankers score the same set of original low-SES essays. \paragraph{Discussion.} The key objects in this identification argument are Assumptions (ref) and (ref). Taken together with separability (ref), they say the following. First, rewrites change only the realization $R_{ik}$ but not the latent content $C_i$, so all within-essay variation across $k$ can be attributed to style. Second, conditional on $C_i$, the effect of requesting rewrite type $k$ under ranker $m$ can be summarized by a constant shift $\lambda_{k,m}$ in the style component, plus idiosyncratic noise with mean zero and a distribution that does not depend on group membership. Under these conditions, the rewrite-averaged score $\bar{s}^{(m)}_i$ is the sum of the content index $\theta_m(C_i)$ and a constant $\bar\lambda_m$ that is common across groups, so differences in $\bar{s}^{(m)}_i$ identify the content gap $A_H^{(m)}-A_L^{(m)}$. Similarly, the deviation $d^{(m)}_i = s^{(m)}_{i0}-\bar{s}^{(m)}_i$ nets out both content and the average rewrite shift, so differences in $d^{(m)}_i$ identify the style gap $U_H^{(m)}-U_L^{(m)}$. In this sense, the role of the rewrites is purely anchoring: they define a reference style environment under which content and style can be separated in a way that is comparable across groups.

These assumptions and the identification argument impose discipline on how the rewrite operator should be constructed. First, to make content fidelity plausible, prompts to the LLM should explicitly emphasize preserving meaning, arguments, and logical structure, and only allow modifications to surface features. Second, Assumption (ref) is fragile to interactions between the rewrite mechanism and the essay itself. If the LLM’s rewrites respond strongly and systematically to the content of the essay, the induced differences in style will be correlated with $\theta_m(C_i)$ and will be absorbed into the content component. To mitigate this, prompts should aim either to produce rewrites in (approximately) a fixed writing style with limited sensitivity to content, while still generating enough variation to pin down the content component, or to generate rewrites that move essays symmetrically along a stylistic axis that is correlated with scores. In the latter case some rewrites tend to increase the score and some tend to decrease it for any given essay, so that \[ \frac{1}{K}\sum_{k\in\mathcal K} \rho_m(R_{ik}) \approx 0 \qquad\text{for all $i$,} \] and small residual deviations from zero will not materially affect the decomposition. In both designs the goal is that the rewrite operator induces a common reference distribution over styles, rather than introducing systematic group- or content-specific shifts that would contaminate the identified content and style components.

In the appendix we show that without assumptions (ref) and (ref) one can still construct a meaningful decomposition if one is willing to fix a particular neutralizing rewrite operator $T$. There we introduce an operator $T$ that maps each original essay to a neutralized version $T(\text{original text})$, and define, for a given scoring function $S$, \[ \mu_G^{S,\mathrm{orig}} := \mathbb{E}\big[S(\text{original text})\mid G\big], \qquad \mu_G^{S,\mathrm{neu}} := \mathbb{E}\big[S(T(\text{original text}))\mid G\big]. \] The observed gap can be decomposed as \[ \Delta_{\mathrm{obs}} = \underbrace{\big(\mu_H^{S^{(H)},\mathrm{neu}} - \mu_L^{S^{(H)},\mathrm{neu}}\big)}_{\text{Content}^{(H)}} + \underbrace{\Big[(\mu_H^{S^{(H)},\mathrm{orig}} - \mu_H^{S^{(H)},\mathrm{neu}}) - (\mu_L^{S^{(H)},\mathrm{orig}} - \mu_L^{S^{(H)},\mathrm{neu}})\Big]}_{\text{Style}^{(H)}} + \underbrace{\big(\mu_L^{S^{(H)},\mathrm{orig}} - \mu_L^{S^{(L)},\mathrm{orig}}\big)}_{\text{Scoring-function tilt (vs.\ $S^{(H)}$)}}. \] Here the “content” term measures the gap that would remain if all essays were rewritten into the same neutral style before scoring with $S^{(H)}$, the “style” term captures the differential premium of the original high- and low-SES writing styles relative to that neutral baseline under $S^{(H)}$, and the tilt term compares how $S^{(H)}$ and $S^{(L)}$ score the same low-SES originals.

This baseline decomposition is valid for any scoring functions $S^{(H)}$ and $S^{(L)}$ and any rewrite operator $T$, without invoking the separability assumptions (ref) and (ref). It requires fewer structural assumptions, at the expense of committing to a particular baseline distribution over styles. It is therefore particularly useful for policy questions such as: “If all students were allowed (or required) to pass their essays through an LLM-based neutralizer before grading, how much of the score gap would remain?” When the separable model (ref) and Assumptions (ref)--(ref) hold, and when the neutral rewrite $T$ can be interpreted as imposing the same reference style environment as the rewrite panel, the neutral content term $\mu_H^{S^{(H)},\mathrm{neu}} - \mu_L^{S^{(H)},\mathrm{neu}}$ coincides with the structurally identified content gap $A_H^{(H)} - A_L^{(H)}$ derived above. In this sense, the structural decomposition and the neutral rewrite-based decomposition are two views of the same underlying idea: using rewrites to define a common style benchmark and separating differences in content from differences in style and scoring functions.

remark[When does the decomposition have a causal interpretation?] Up to this point, our decomposition is descriptive: it partitions the observed score gap into components indexed by content, style, and scoring functions under the observed joint distribution of essays. Each component suggests a natural “lever” (changing content, changing style, changing the scoring rule), but reading the terms as causal effects requires strong conditional independence assumptions about how the remaining components would behave under such interventions. To fix ideas, consider a hypothetical policy that “equates content” by shifting the content distribution of low-SES students from $F_L^C$ to $F_H^C$. Let \[ m_L^{\rho}(c) := \mathbb E[\rho(R_{i0}) \mid C_i=c, G_i=L] \] denote the mean of our style functional at content level $c$ for low-SES students. Under the status quo, \[ \mathbb E_L[\rho(R_{i0})] \;=\; \int m_L^{\rho}(c)\, dF_L^C(c). \] Under a content-only policy that changes the marginal distribution of $C_i$ but leaves the conditional law of style given content unchanged for low-SES students, the mean style component would instead be \[ \int m_L^{\rho}(c)\, dF_H^C(c). \] By contrast, our content component is constructed holding the mean style term $\mathbb E_L[\rho(R_{i0})]$ fixed when varying content. Therefore, the content component coincides with the causal effect of a feasible content-only policy only under an additional restriction ensuring that the style term is (approximately) insensitive to the induced change in the content mix---for example, that $m_L^\rho(c)$ is (approximately) constant in $c$ over the relevant support (or, more generally, that $\int m_L^\rho(c)\, dF_H^C(c)\approx \int m_L^\rho(c)\, dF_L^C(c)$). In realistic settings, content and style are typically intertwined, so equalizing content is more naturally interpreted as a descriptive reweighting of observed essays than as the effect of a feasible content intervention.\footnote{A further requirement is overlap: the policy is only well-defined on regions where $F_H^C$ assigns mass to content values attainable by low-SES students.} An analogous issue arises for policies that “equate style”: our style component replaces $R_{i0}$ with an LLM-induced reference style while holding $C_i$ and the scoring functions fixed, and it is causal only if this intervention changes style alone—neither shifting expressed content nor changing how scorers respond once style is modified; since writing tools can affect what is articulated and graders may adapt to standardization, we read the style term as a mechanical association in the current environment. The same caveat applies to the scoring-function component: the “scoring-function tilt” compares how $S^{(H)}$ and $S^{(L)}$ grade the same texts and is causal only if scoring rules are stable objects and unifying them would not induce equilibrium responses (e.g., grader selection or institutional feedbacks).

Estimation

In the empirical analysis we implement the decomposition in two layers. We first estimate the scoring functions $S^{(H)}$ and $S^{(L)}$ from the human data, and then recover empirical counterparts of the content and style components using a fixed-effects decomposition of the resulting predicted scores. For each scoring concept $m \in \{H, L\}$, we train a gradient-boosted tree model (XGBoost\footnote{The XGBoost model uses 396 trees with a maximum depth of 4, a learning rate of 0.05, subsampling rates of 0.8, and L1/L2 regularization set to 0.055 and 0.026, respectively. These parameters were selected by performing a simple randomized hyperparameter search to tune the XGBRegressor, focusing on key parameters that control model capacity and regularization. The search varied the number of trees, tree depth, and L1/L2 regularization strengths, with regularization terms sampled from log-uniform distributions to cover multiple orders of magnitude. Model performance was evaluated using RMSE as the optimization metric. Hyperparameters were selected via cross-validation over the full dataset, using shuffled $K$-fold cross-validation with three folds.}) to predict the human-assigned score using the feature vector $x_{i0}$ of the original essay. This feature vector includes the text embedding, the full set of style variables, student characteristics, and an indicator for the essay prompt. Exact definition of the variable is in section (ref). This delivers a fitted scoring rule $\hat S^{(m)}(\cdot)$ for each $m$. With a slight abuse of notation, we write $S^{(m)}$ for these estimated functions in what follows and treat them as fixed when we construct the decomposition. For each essay $i$ and each version $k$ (the original and all rewrites), we compute the predicted score \[ \hat s^{(m)}_{ik} := S^{(m)}(x_{ik}), \] which serves as the empirical analogue of $s^{(m)}_{ik}$ in the framework. To avoid overfitting, predicted scores are obtained via sample splitting, similar to cross-validating: the scoring model is trained on 80% of the data and used to predict scores on the remaining 20%. This procedure is repeated so that every observation is scored using a model trained on data that excludes it.

Given these predicted scores, we recover estimates of the latent content and style components via a fixed-effects regression that mirrors the additive structure in (ref). For each $m$ we estimate \[ \hat s^{(m)}_{ik} = \alpha^{(m)}_i + \sum_{k\in \{1,...,K\}} \gamma^{(m)}_k \mathbf{1}\{k\} + \varepsilon^{(m)}_{ik}, \] where $\alpha^{(m)}_i$ is an essay fixed effect and the dummies $\mathbf{1}\{k\}$ indicate the rewrite type, with the original $k=0$ absorbed into the fixed effect. Under Assumptions (ref), (ref), and (ref), the fixed effect $\alpha^{(m)}_i$ provides an estimate of the content index $\theta_m(C_i)$ up to an additive constant, while the fitted style component for version $k$ of essay $i$ is \[ \hat\rho_m(R_{ik}) := \hat s^{(m)}_{ik} - \hat\alpha^{(m)}_i = \hat\gamma^{(m)}_k + \hat\varepsilon^{(m)}_{ik}. \] We then construct sample analogues of the group-specific content and style indices by averaging these estimated components over essays. For each $G\in\{H,L\}$ and each $m$ we define \[ \widehat{A}^{(m)}_G := \frac{1}{N_G}\sum_{i:G_i=G} \hat\alpha^{(m)}_i, \qquad \widehat{U}^{(m)}_G := \frac{1}{N_G}\sum_{i:G_i=G} \hat\rho_m(R_{i0}), \] where $\hat\rho_m(R_{i0})$ is the estimated style component for the original version of essay $i$. By construction, $\widehat{A}^{(m)}_G$ is the empirical analogue of $\mathbb E_G[\theta_m(C_i)]$ and $\widehat{U}^{(m)}_G$ is the analogue of $\mathbb E_G[\rho_m(R_{i0})]$, up to constants that cancel when we take differences across groups. The estimated content and style components of the gap under $S^{(H)}$ are therefore given by \[ \widehat{\text{Content}}^{(H)} = \widehat{A}^{(H)}_H - \widehat{A}^{(H)}_L, \qquad \widehat{\text{Style}}^{(H)} = \widehat{U}^{(H)}_H - \widehat{U}^{(H)}_L, \] while the ranker-tilting term is estimated directly from the scoring functions as \[ \widehat{\text{Tilt}} = \frac{1}{N_L}\sum_{i:G_i=L}\big(\hat s^{(H)}_{i0} - \hat s^{(L)}_{i0}\big), \] which compares how the two learned scoring rules grade the same set of original low-SES essays.

Inference on these components is obtained by bootstrapping the entire procedure at the essay level. In each bootstrap replication we resample essays with replacement, re-estimate the fixed-effects model, and recompute the empirical decomposition. The dispersion of the resulting bootstrap distribution for each component provides standard errors and confidence intervals that account jointly for uncertainty in the first-stage learning of the scoring functions and in the second-stage decomposition based on the fixed effects.

Data

In this section, we describe our main data sources and how we construct the data. We begin with our primary dataset of student essays, PERSUADE 2.0. We then explain how we construct the content variables. We then describe how we construct the style variables, which capture stylistic choices in writing. Finally, we outline how we produce the essay rewrites.

The PERSUADE 2.0 dataset

PERSUADE 2.0 (Persuasive Essays for Rating, Selecting, and Understanding Argumentative and Discourse Elements) is a large-scale, open-source corpus designed to advance research into argumentative writing quality, discourse elements, and their effectiveness crossley2024large. The corpus was created to address the need for systematic analysis of student argumentation and to support the development of unbiased computational algorithms for assessing writing quality. PERSUADE 2.0 builds upon the PERSUADE 1.0 corpus by adding holistic essay quality scores and effectiveness ratings for individual discourse elements. The dataset contains essays written in response to 15 different prompts across two distinct writing tasks: independent writing (8 prompts) and source-based writing (7 prompts, each with a single source text).

The dataset comprises 25,996 argumentative essays written by students in grades 6-12 across the United States. The corpus reflects the diversity of the U.S. student population, with writers identifying as White (45%), Hispanic (25%), Black (19%), and other racial/ethnic groups. Essays are linked to detailed demographic information including gender and race/ethnicity. Notably, 20,759 essays (approximately 80% of the corpus) include data on student eligibility for federal assistance programs, which we use as an indicator of socioeconomic status (SES). All essays in the dataset have been assigned holistic quality scores by trained expert raters.

The holistic scores were assigned using a standardized 1-6 point SAT essay scoring rubric, where 6 indicates "clear and consistent mastery of writing." Expert raters were assigned to rate the essays underwent prompt-specific training and employed a double-blind rating process, with 100% adjudication by a third rater when necessary. The rubric for source-based essays was slightly modified to include evaluation of how effectively students incorporated evidence from source texts. Inter-rater reliability before adjudication demonstrated strong agreement, indicating consistent and reliable quality judgments across the corpus. crossley2024large

“Content” Variables

To represent the semantic content of the PERSUADE essays, we embed each entire essay using contextual representations from a pre-trained BERT encoder devlin2019bert. We use bert-base-uncased, which is pre-trained on large-scale general-domain English text and is therefore well-suited to the naturalistic language in these essays. Because BERT accepts sequences of up to 512 WordPiece tokens, essays exceeding this length are truncated to the first 512 tokens. For each essay, we take the final-layer embedding of the special \([CLS]\) token, which is commonly used as a pooled summary representation of the input sequence. This procedure maps each essay to a fixed-dimensional vector in a shared high-dimensional space. Prior work shows that contextual BERT representations encode higher-level abstractions—such as topic, intent, and relational meaning—beyond surface lexical overlap devlin-etal-2019-bert,reimers-gurevych-2019-sentence,ethayarajh-2019-contextual,tenney-etal-2019-bert,jawahar-etal-2019-bert,clark-etal-2019-bert,rogers-etal-2020-primer,karpukhin-etal-2020-dense. We therefore interpret distances between \([CLS]\) vectors as reflecting conceptual relatedness between essays rather than mere word overlap.

“Style” Variables

To quantify stylistic dimensions of the essays, we employ computational linguistic tools from \href{https://www.linguisticanalysistools.org/}{the Suite of Automatic Linguistic Analysis Tools} potter2025assessing. These tools enable systematic measurement across three key dimensions of writing style. In total, our stylistic representation comprises 355 linguistically variables, spanning syntactic, lexical, and discourse-level dimensions.

\paragraph{Syntactic complexity and sophistication} We assess the structural properties of student writing using TAASSC (Tool for the Automatic Analysis of Syntactic Sophistication and Complexity, version 1.3.8; kyle2016measuring). This tool captures two complementary aspects of syntax: traditional complexity metrics such as clause subordination and phrase elaboration, alongside usage-based sophistication measures that reflect developmental difficulty. The sophistication metrics draw on frequency and contingency patterns from large-scale corpora, notably the Corpus of Contemporary American English, to identify constructions that are statistically less common and therefore more challenging to acquire kyle2016measuring, potter2025assessing. Under this framework, rarer syntactic patterns signal greater linguistic maturity. From TAASSC, we extract 149 measures, including indices such as mean clause length, frequency-weighted measures of multi-clause constructions, verbal dependency counts, and proportions of finite and non-finite structures.

\paragraph{Lexical variation} The diversity of vocabulary within each essay is measured using TAALED (Tool for the Automatic Analysis of Lexical Diversity; kyle2021assessing, zenker2021investigating). This tool generates indices that capture the degree of lexical variety employed by writers. Prior validation work demonstrates strong correspondence between TAALED metrics and expert assessments of vocabulary richness, establishing these measures as reliable proxies for stylistic development kyle2021assessing, potter2025assessing. TAALED contributes 38 lexical diversity and sophistication variables, including moving-average type–token ratios, measures of how common or rare words are in large reference corpora, and word-length–based indices that reflect vocabulary development.

\paragraph{Textual cohesion} We evaluate the connectivity of ideas using TAACO (Tool for the Automatic Analysis of Cohesion, version 2.1.3; crossley2016taaco, crossley2019taaco2). TAACO quantifies how explicitly writers signal relationships among concepts at multiple scales: adjacent word and sentence connections, paragraph-level coherence, and document-wide integration. The indices encompass various cohesive devices including repetition of lexical items, semantic similarity, and explicit connectives that guide reader comprehension. TAACO provides 168 cohesion-related variables, which include local and global cohesion metrics such as lexical overlap between adjacent sentences, semantic similarity across paragraphs, and the use of explicit connectives.

Collectively, these three tools, TAASSC, TAALED, and TAACO, capture stylistic features that characterize how writers communicate rather than, providing dimensions of expression, organization, and connectivity.

Rewrites

As discussed in Section (ref), separating content from style requires generating stylistic variation while holding content fixed. For this purpose, we construct two sets of rewrites.

table[table omitted — 4,183 chars of source]

\paragraph{SAT Score-Based Rewrites.} For each original essay, we generate six rewrites by applying prompts that target specific SAT score tiers and operationalize each tier using the SAT rubric dimensions that primarily reflect language and style (e.g., “the essay exhibits skillful use of language” versus “the essay displays fundamental errors in vocabulary”). The prompts instruct the model to adjust only these stylistic components, such as word choice, sentence fluency, grammar/mechanics, and cohesion, to match the target tier, while keeping the essay’s substantive content and meaning fixed. The resulting rewrites are intended to span a quality gradient along a single “writing quality’’ axis, yielding six versions per essay aligned with distinct quality levels. Table (ref) illustrates how the same student-written paragraph is rewritten at different levels. As the SAT level increases, the rewrites become more academic and formal in tone.

\paragraph{Neutral Rewrites/GPT Baseline.} We also generate rewrites using a single neutral prompt that serves as a control without explicit stylistic instructions. Rather than varying the prompt, we generate six outputs per essay by sampling repeatedly from the model conditional distribution of rewrites, on the same input. These texts are therefore stylistically more uniform, reflecting the model's default (“neutral”) writing style. Table (ref) illustrates how the model rewrites the same student-written paragraph when prompted to rewrite without any specific stylistic instructions.

All rewrites are generated using GPT-4o (via Azure), with temperature = 1 and max_tokens = 2048. The exact prompts are reported in Appendix (ref).To satisfy Assumption (ref), rewrites must preserve the underlying content of the original essay. We therefore subject each rewritten essay to an additional verification step using GPT-4o, which compares the rewritten text to the original and evaluates whether the substantive content has been altered (see Appendix (ref) for the verification prompt). Rewrites that are flagged as altering content are discarded from the analysis. Across the six SAT based rewrite samples, the vast majority of generated texts are retained after verification. Specifically, the number of rejected rewrites (i.e., flagged as altering content) is 54 (0.27%) for rewrtie SAT of level 1, 38 (0.19%) for level 2, 37 (0.19%) for level 3, 676 (3.38%) for level 4, 107 (0.54%) for level 5, and 791 (3.96%) for level 6. For the no-description/neutral rewrite condition, 329 rewrites (1.59%) are rejected. Overall, rejection rates remain low across all rewrites, indicating that essay content in the large maintained in the majority of cases.

Using an LLM for content verification is appropriate in this setting because the task of assessing semantic equivalence between two texts is precisely the type of judgment for which large language models have been shown to perform reliably, and because the verifier is used only as a binary filter rather than as a source of continuous measurement. This design limits the scope for systematic bias from the verification step and ensures that identification relies on content-preserving rewrites by construction.

Descriptive Statistics

In this section, we first describe the data and the score gaps between high- and low-SES students in our sample. We then present our estimated scoring function, and finally we discuss the rewrites.

Descriptive Statistics and The Score Gap

Table (ref) reports descriptive statistics for our sample. The sample is roughly balanced by socioeconomic status, with 46.5% of students classified as low SES. Women make up 51.2% of the sample, with a slightly higher share among low-SES students. Racial composition differs sharply by SES: 24.7% of low-SES students are White, compared with 63.0% among high-SES students. Finally, although the PERSUADE data cover grades 6–12, our SES-linked sample is concentrated in grades 6, 8, 10, and 11, where coverage is richest.

Next, we move the descire the score gaps. Figure (ref) shows the score gap we seek to explain by displaying the score distributions for high and low SES students. The dashed lines denote the mean score within each group. On average, high-SES students score about 0.67 points higher than low-SES students on the 1–6 scale.

Beyond the difference in means, the distributions exhibit distinct shapes. Scores for low-SES students are more concentrated at the lower end of the scale, particularly at scores of 2 and 3. In contrast, scores for high-SES students arec more dispersed, with the modal score at 4. High-SES students are also substantially more likely to receive top scores of 5 or 6, whereas such scores are relatively rare among low-SES students (24.5% vs. 8.5%). Overall, the low-SES score distribution is shifted leftward relative to that of high-SES students.

table[table omitted — 3,728 chars of source]

The SES score gap appears across demographic subgroups and grades. The gap is larger among non-White students (0.746) than White students (0.569), and similar for males (0.657) and females (0.682). By grade, the gap is smallest in grade 6 (0.238) and remains around 0.48-0.58 in grades 8–11.

Figure (ref) in the Appendix shows that this gap is relatively stable across writing prompts. The estimated differences range from 0.73 points for the “Mandatory Extracurricular Activities” prompt to 0.41 for “Community Service” and 0.23 for “A Cowboy Who Rode the Wave.” The consistency of the gap across assignments and topical domains suggests that the observed score differences are not driven by differences in specific area, but are likely to exists across wide array of domains.

figure[figure omitted — 611 chars of source]

The Scorer Functions

Figure (ref) assesses the performance of the ranking model separately for low- and high-SES students. The figure plots the expected holistic score conditional on the model’s predicted score (x-axis). Intuitively, if the ranker is unbiased and successfully captures the underlying scoring rule, the average realized score should coincide with the prediction, placing the conditional expectation along the 45-degree line. The figure shows that this relationship holds closely for both low and high SES scorers. Deviations from the 45-degree line are small and statistically non significant, indicating limited systematic bias and suggesting that the model captures the underlying ranking function well for both groups.

Figure (ref) reports the $R^2$ values for the two group-specific predictors. Both predictive models explain a large share of the variation in holistic scores (0.735 for low-SES students and 0.764 for high-SES students), indicating strong predictive performance in both groups and suggesting that our set of variables captures most of the systematic variation in scores.

figure[figure omitted — 1,064 chars of source]

Figure (ref) compares the rankings assigned by the high- and low-SES models for the same essays. The figure plots binned averages of the high-SES predicted scores against the corresponding low-SES predicted scores. This relationship directly reflects the “tilting” of the score distribution across groups. Large deviations from the 45-degree line would indicate substantial differences in how the two models rank essays, and thus a larger contribution of ranking differences to the overall score gap. In contrast, points lying close to the 45-degree line imply broadly similar rankings across groups. The figure shows that across most of the support, high-SES predicted scores lie slightly above the 45-degree line. This indicates that, on average, the high-SES model assigns marginally higher ranks to the same essays than the low-SES model. These differences are small in magnitude and increase modestly at higher score levels, suggesting limited but systematic tilting rather than large discrepancies in ranking.

Finally, Figure (ref)\footnote{The corresponding distribution under the low-scorer model is shown in Figure (ref) in the Appendix.} plots the distribution of predicted essay scores for high- and low-SES students using the same (high-SES) scoring function and the full set of explanatory variables. Even under this common scoring rule, the predicted distribution for low-SES students is systematically shifted to the left, placing substantially more mass on lower scores relative to high-SES students. This comparison isolates differences in inputs rather than differences in how inputs are rewarded. The persistence of a sizable distributional gap under a fixed scoring function therefore indicates that disparities in the scoring rule alone cannot account for the observed SES gap. Instead, the figure reinforces the conclusion that differences in underlying content and style features play a central role.

figure[figure omitted — 1,188 chars of source]

Content and Style

We now turn to examine differences in the content and style variables. Figure (ref) displays the density of a two-dimensional projection of the text embeddings, obtained using the t-SNE algorithm vanDerMaaten2008tsne, pooling across all writing prompts. The figure shows that high- and low-SES students exhibit distinct regions of concentration in several prompts. To the extent that the embeddings capture semantic differences across topics, this pattern indicates that the writing of high- and low-SES students occupies different regions of the content space, suggesting systematic differences in the ideas they choose to express. We also perform an exercise in which we predict the SES of essay authors using the embeddings. We find strong separation, achieving an AUC of 0.705, again implying that content differs between high- and low-SES students.

Figure (ref) in the Appendix explores these distributions across the 12 different writing prompts. We again find that the observed content differences are not driven by any single prompt, but instead appear consistently across prompts. This pattern suggests that high- and low-SES students systematically differ in how they respond to these questions, rather than reacting differently to a particular prompt.

figure[figure omitted — 628 chars of source]

Turning to style, table (ref), shows a sample of the style variables in our sample\footnote{We include an explanation of each style variable that appears in table (ref), in Appendix (ref)}. The table shows that high-SES students write essays that look systematically more “academic” and elaborated than low-SES students along multiple stylistic dimensions. A major difference between high and low SES is sheer volume: high-SES essays are about 83 words longer on average (451 vs. 368, roughly 22% longer), which goes hand-in-hand with higher-scoring writing and may mechanically allow more development of ideas. Beyond length, high-SES essays use slightly more lexically sophisticated and information-dense language: average word length is higher (4.390 vs. 4.295), lexical density is modestly higher (0.443 vs. 0.433), and lexical diversity is higher as measured by MTLD\footnote{MTLD (Measure of Textual Lexical Diversity) is a length-robust measure of vocabulary variety that tracks how quickly word repetition accumulates as a text unfolds. Higher MTLD indicates more diverse vocabulary because longer segments are needed before the text’s type–token ratio falls below a fixed threshold. } (56.2 vs. 54.2), with a smaller but consistent increase in lemma-based MATTR (0.741 vs. 0.737)\footnote{MATTR (Moving-Average Type–Token Ratio) measures lexical diversity by computing the type–token ratio within a fixed-length sliding window and averaging it across the text. Higher MATTR indicates more varied vocabulary, and using a fixed window makes it less sensitive to overall text length than the raw type–token ratio.}. High-SES essays also show somewhat greater syntactic complexity: clauses are longer (MLC 12.316 vs. 11.911), T-units are longer (29.157 vs. 28.721), verbs carry more syntactic attachments (mean verbal dependencies 6.207 vs. 5.992), and non-finite constructions appear more often (infinitives 0.154 vs. 0.148; non-finite clauses 0.424 vs. 0.400), consistent with more embedding and syntactic compression. Looking at nominalization, we find a pronounced gap between high- and low-SES essays. High-SES essays contain substantially more nominalizations (13.4 vs. 8.9; a difference of 4.49, roughly a 50% increase), consistent with a more abstract, noun-heavy style characteristic of academic prose. In contrast, discourse cohesion and explicit causal signaling differ little: adjacent-sentence content-word overlap is nearly identical (0.163 vs. 0.160), and causal connectives are if anything slightly less frequent in high-SES writing (0.012 vs. 0.013). Finally, high-SES essays rely a bit less on pronouns relative to nouns (pronoun-to-noun ratio 0.299 vs. 0.312), indicating more explicit noun-based reference, which often reads as clearer and more formal.

These descriptive patterns, indicating higher levels of writing among high-SES students, also show up in a single composite style score. Figure (ref) plots the distribution of predicted-score probabilities from a model trained only on the style variables. High-SES essays are spread more evenly across the score range, whereas low-SES essays are disproportionately concentrated at low predicted holistic scores. This separation is qualitatively similar to what we see using the full feature set (all variables and embeddings), though the style-only model produces a more pronounced pile-up at the bottom among low-SES essays.

figure[figure omitted — 1,121 chars of source]

Comparing figure (ref) and (ref) demonstrates our main challenge in isolating the role of content and the style in shaping outcomes for students. Figure (ref) shows the predicted value distribution based on the embedding only. As we can see the implied distribution are very similar to to the style distribution, implying that the two may carry similar infomration. Figure (ref)explore this more and shows the correlation between the scores assigned using predictors who were trained only on a subset of the variables (all variables, embedding, all only style varaibles, TAACO, TAASSC, TAALED). The figure shows very high correlation in predicted values, indicating that the different set of variables carry similar information on on the holistic score.

Figure (ref) makes the redundancy between “content” (embeddings) and our style measures especially clear. Using embeddings alone already predicts a large share of score variation: $R^2=0.64$ for low-SES students and 0.66 for high-SES students. Style variables on their own perform slightly better ($R^2=0.71$ and $0.75$, respectively). But combining content and style adds only a modest incremental gain (to $0.73$ for low SES and $0.764$ for high SES) despite doubling the information set.

This small marginal improvement implies that content and style are strongly correlated in the data: much of what style “explains” is already encoded in the semantic/content representation, and vice versa. That pattern is intuitive. More nuanced ideas often require more precise lexical and syntactic control, so stronger content tends to co-occur with more “academic” style. Likewise, some ideas are inherently harder to express tersely—complex arguments naturally demand longer sentences and more elaborate structure—so stylistic sophistication can be partly a byproduct of the content being conveyed. More so, although we treat embeddings as capturing semantic meaning, contextual representations also encode surface and syntactic regularities that reflect stylistic choices jawahar2019does, reinforcing that content and style are difficult to cleanly separate in practice.

To quantify redundancy between style features and content (embedding) features, we measure how much of the variance explained by content is also explained by style. Let $R_S^2$ denote the $R^2$ from regressing the holistic score on style variables only, $R_C^2$ the $R^2$ from regressing the score on content variables only, and $R_{SC}^2$ the $R^2$ from regressing the score on both sets of variables. Define the incremental contribution of content beyond style as $\Delta_C \equiv R_{SC}^2 - R_S^2$. Then $R_C^2 - \Delta_C$ is the portion of the content-only explained variance that is not unique to content: it is variance that content explains in isolation, but that becomes redundant once style is included because style already accounts for it.

We summarize this redundancy via \[R^2_{\mathrm{expl}} =\frac{R_C^2-\Delta_C}{R_C^2} =\frac{R_S^2+R_C^2-R_{SC}^2}{R_C^2}.\] The numerator, $R_C^2-\Delta_C = R_S^2+R_C^2-R_{SC}^2$, is the standard commonality (shared explained variance) of style and content with respect to holistic scores. Dividing by $R_C^2$ expresses this overlap as a fraction of the explanatory power of the content model. Thus, $R^2_{\mathrm{expl}}$ measures the share of the content model’s signal that is recoverable from style.

In Figure (ref), we find $R^2_{\mathrm{expl}}\approx 0.98$, for both high and low SES models, implying that about 98% of the variance in scores explained by content features is also explained by style features, leaving only about 2% as uniquely incremental content signal beyond style. Overall, this indicates that style and content features contain highly overlapping information about the score.

figure[figure omitted — 786 chars of source]

Rewrites

We next turn to the rewritten-essay analysis. Figure (ref) plots the mean predicted holistic score for each rewrite style, separately for low- and high-SES students. The figure validates our “move along a writing axis” design: as the target SAT level increases from 1 to 6, average predicted scores rise monotonically for both groups, indicating that the rewrites successfully shift writing quality in the intended direction. At the lower end, SAT-1 and SAT-2 rewrites score below the original essays for both groups, confirming that the model is not mechanically “improving” text but can also degrade it when instructed, with our scorer responding accordingly.

The figure also helps calibrate the standard “GPT-polish” baseline. On average, the generic GPT rewrite raises predicted scores to roughly the same level as our SAT-4 rewrites, suggesting that this commonly used approach corresponds to a sizable upward move along the style axis.

Importantly, none of the rewrite regimes comes close to eliminating the SES gap. Even when both groups are pushed toward a common reference style, a substantial score difference remains. This persistence is consistent with the view that part of the observed gap reflects differences in content---ideas, argument structure, and substance---rather than expression alone.

figure[figure omitted — 598 chars of source]

These patterns speak directly to Assumption (ref). As we vary the target style from SAT level 1 to 6, scores for both groups move in parallel: the transformation shifts the level of predicted scores up or down, but leaves the High--Low gap approximately unchanged. This “parallel shift” is what we would expect if the rewrite operates primarily through a style channel that enters additively and similarly for both groups, while content differences remain as a stable wedge. Appendix Figure (ref) provides complementary distributional evidence. It plots kernel density estimates of the predicted-score distribution for each rewrite level. The densities shift upward or downward with the target SAT level but remain similar in shape, consistent with the rewrites acting approximately as a constant (additive) style shift.

figure[figure omitted — 1,798 chars of source]

To probe Assumption (ref) more sharply, we run a difference-in-differences diagnostic that isolates the incremental effect of a rewrite level net of within-group content. The first difference compares two rewrite levels within the same SES group, so any stable group-specific content differences cancel. The second difference then compares these within-group rewrite effects across High- and Low-SES students. If rewrites operate as a common, additive style shift, our Assumption (ref), this difference-in-differences should be zero.

Formally, for group $G\in\{H,L\}$ define the within-group rewrite contrast \[ \Delta^{kk'}_G \;:=\; \mathbb{E}\!\left[s_{ik}-s_{ik'}\mid G\right]. \] Under Assumption (ref), rewriting from $k'$ to $k$ changes the score only through the style component, so \[ \Delta^{kk'}_G =\mathbb{E}\!\left[\rho_m(R_{ik})-\rho_m(R_{ik'})\mid G\right] =\lambda_{m,k}-\lambda_{m,k'}, \] which does not depend on group membership. Consequently, the difference-in-differences contrast \[ \widehat{\Delta}^{kk'}_H-\widehat{\Delta}^{kk'}_L \] should be close to zero for all $(k,k')$.

Figure (ref) reports these contrasts across all pairs of rewrite levels. The lower triangle (blue) shows the magnitude of the within--High-SES rewrite premium $|\Delta^{kk'}_H|$, which is large: moving from lower to higher SAT targets shifts predicted scores by several tenths to well over a point. The upper triangle (red) shows the corresponding difference-in-differences, $\widehat{\Delta}^{kk'}_H-\widehat{\Delta}^{kk'}_L$. Overall, these red entries cluster tightly around zero across the matrix, with only a modest tendency for comparisons involving the lowest rewrite level ($k=1$) to be slightly larger in magnitude. Moreover, they are an order of magnitude smaller than the underlying rewrite premia in blue. Taken together, the figure provides strong, direct evidence for Assumption (ref): rewrites substantially move score levels, but they do so in a nearly parallel way across groups, leaving the High--Low gap essentially unchanged.

Decomposition Results

In this section we discuss our decomposition results using the SAT rewrites. Figure (ref) summarizes our central decomposition of the high–low SES writing score gap into three parts: (i) content differences, (ii) style differences, and (iii) the scoring-function differences (“tilt”) component, capturing other unexplained contributors to the gap.\footnote{Full results, with bootstrap standard errors can be found in table (ref) in the Appendix.}. This structure is designed to separate “what is being said” from “how it is said,” and allow us to learn about how each of these components contribute to the realized writing score gap.

figure[figure omitted — 1,239 chars of source]

The first takeaway is that most of the score gap is attributed to content, with style playing a sizable but secondary role, and the tilt component typically smaller. In the pooled sample (“All”), the content share accounts for 68.8% of the total SES gap, the style share accounts for 26.1% of the total gap, leaving 5.2% for the scorer function tilt. These results suggests that the SES gap is driven primarily by substantive differences in what students say (i.e. arguments, reasoning, and the relevance and organization of evidence), rather than by surface how they say it. At the same time, the style share is nontrivial: even holding fixed the underlying content benchmark induced by the rewrite panel, differences in writing style, phrasing, grammatical control, and coherence explain an meaningful portion of little more than quarter of the gap. Interpreted through a policy lens, this implies that interventions aimed purely at improving expression (or mechanically standardizing it) could potentially reduce, but not eliminate, the SES gap.

Figure (ref) displays the estimated distributions of the content and style components, separately for high and low SES students. As both the content component (for each essay) and the style component are identified only up to an additive constant, the x-axis levels in these histograms are arbitrary up to a common shift. We therefore focus on the shapes and relative shifts of the high-vs. low-SES distributions rather than on absolute levels. Panel (a) shows that the content component distributions are clearly separated: the mean content index is 2.644 for high-SES students versus 2.188 for low-SES students. The difference between these two numbers capture the content gap (0.456). The distribution are clearly different, where low SES has high mass of content on lower scores, compare to the high SES stuendes. The support of the content component is wide, and both distributions are realitvly spread out, showing a substantial heterogeneity in content quality within each SES group while maintaining a systematic shift in means. The fact that the variation in the content is large, support the notion that content is the primary driver of the score gap. Panel (b) presents the style component distributions, which reveal a smaller but still meaningful separation. High-SES students have a mean style premium of 0.981 relative to the rewrite-induced benchmark, while low-SES students average 0.807 (implying gap of 0.173). Both distributions exhibit tighter dispersion than the content distributions, where deviations from the mean value, are modest in magnitude, with shifts of more than one point being relatively rare. The relatively narrow spread of the style component implies that, the difference in style can have is relatively small, compare to variation in the content.

figure[figure omitted — 1,044 chars of source]

Figure (ref) examines within-student correlation between content and style. Among high-SES students, the Pearson correlation is 0.384, implying that students who score higher on content also tend to receive higher style premia; among low-SES students, the correlation is weaker at 0.187. Overall, the positive correlations indicate that content and style co-move: better ideas tend to be expressed more effectively, making the two dimensions difficult to separate empirically.

The weaker correlation for low-SES students suggests a looser mapping from ideas to expression. One channel is heterogeneity in skill development: some students may construct strong arguments but lack familiarity with academic conventions, while others may master surface mechanics (grammar, templates) without comparable gains in substantive reasoning—either pattern attenuates the content–style link. A second channel is linguistic distance and “code-switching” demands: students whose home language patterns differ from standardized academic English may need to translate and suppress native forms, so sophisticated ideas can coexist with lower measured style, further decoupling the two components.

figure[figure omitted — 728 chars of source]

Next, we turn to how the decomposition compares across subgroups. In Figure (ref), we see how the roles of content, style, and tilt change across gender, race, and grade. The shares are broadly similar by race and gender (the content share ranges from 67% to 71%), suggesting that the relative importance of content versus style is not driven by a single demographic subgroup. Where we do see systematic movement is across grade levels: the bars indicate that the content share tends to decline and the style share tends to rise in higher grades (from a content share of 68.5% in grade 6 to 51.2% in grade 11 and style component share going from 29.2% to 42.2%). This pattern is consistent with greater stylistic differentiation as students mature: students may increasingly develop distinct voices and specialize in particular forms of writing that are rewarded differently by the scoring rubric; some students may become more adept at adopting the styles most strongly rewarded in school; and, independently, evaluators may place greater weight on stylistic conventions at higher grade levels.

Finally, Figure (ref) in the Appendix examines how the decomposition varies across prompts. The prompt-level results suggest that the content–style split is not driven by any single prompt; instead, it fluctuates within a fairly narrow range across tasks. This matters because the raw SES score gap itself is relatively stable across prompts. In most prompts, the gap is primarily explained by content differences, though several prompts exhibit noticeably larger style shares. For example, Distance learning (45%), Summer projects (41.4%), Community service (33.2%), Driverless cars (35%), and Seeking multiple opinions (31.4%) all attribute more than 30% of the gap to style differences.

With the exception of Driverless cars,\footnote{In Driverless cars, students are given an article but are also asked to express their own opinion.} these prompts largely involve “independent” writing: students are asked to produce an essay in response to a question without an accompanying text that anchors what content to include. In such settings, style may play a larger role because reviewers have less of a shared benchmark for what a “complete” answer should contain, making the presentation of ideas more salient.

It's important to emphasize how we read these results. The decomposition is descriptive: it partitions the observed gap under the current joint distribution of essays and scoring practices, and the components naturally map to “levers” (content, expression, scoring rules), but causal interpretations require additional assumptions, as discussed in remark (ref). With that caveat, the pattern in Figure (ref) points to a clear conclusion: SES disparities in writing scores appear to be driven primarily by differences in substantive content, with a meaningful additional contribution from style, and a smaller role for systematic differences in scoring rules.

Robustness

In this section, we assess the robustness of our results. Figure (ref) compares our main SAT decomposition with four alternative specifications. First, we restrict attention to essays whose original human SAT score lies in $\{2,3,4,5\}$. For each such essay with original score $s_i$, we retain only rewrite versions with scores $k \in \{s_i-1, s_i, s_i+1\}$ (e.g., $s_i=3$ implies keeping SAT-2, SAT-3, and SAT-4). We then re-estimate the same fixed-effects regression and SAT decomposition as in Section (ref) on this restricted rewrite panel. If upward and downward rewrites are symmetric, then $\sum_k \lambda_k \approx 0$, allowing us to net out potential rewrite effects. In this subsample, the total gap is smaller, 0.54 points. The relative importance of content and style is largely unchanged compare to our baseline SAT results, at 69.3% and 24.6%, respectively.

Next, we repeat the decomposition while restricting attention to rewrites involving SAT scores 1 and 6, and 2 and 5. Appendix Figure (ref) shows that the average difference between rewrites and baseline essays for scores 1 vs.\ 6 and 2 vs.\ 5 is roughly symmetric, again implying that the rewrite effects net out \(\sum_k \lambda_k \approx 0\). This restriction has little effect on the decomposition: content and style account for 66.5% and 28.1%, respectively.

In Figure (ref) (Section (ref)), we saw in the difference-in-differences design that SAT-1 rewrites may be affected by content, as they exhibit slight difference in the effects on high- and low-SES essays. In the “Drop SAT 1” specification, removing the SAT-1 rewrites has little impact on the decomposition: the content and style shares are 71% and 23%, respectively.

Finally, the last bar considers a decomposition that replaces our targeted SAT rewrites with a “standard” GPT rewrite, as described in section (ref), where we simply ask the model to rewrite the essay without specifying a target style level. In this case, the content share is slightly higher (76%), while the style share is lower (18.5%). In Appendix (ref), we expand our discussion on this alternative rewrite.

Taken together, Figure (ref) shows that the qualitative conclusion is stable: across specifications, content remains the dominant contributor to the SES gap, with a meaningful but secondary role for style.

figure[figure omitted — 1,117 chars of source]

Conclusions

The ability to communicate - ideas, intentions, and concepts - shapes life chances. In schools, workplaces, and professional settings, people are judged not only on what they know or their ideas, but on how well they can express it in the forms institutions recognize and reward. This paper examines a central question in the study of inequality and evaluation: to what extent do observed disparities reflect differences in underlying ideas and reasoning versus differences in individuals’ ability to articulate those ideas in forms institutions reward? In many settings—education, labor markets, and professional evaluation—success depends not only on what people want but also on how effectively they express it. When articulation is unevenly distributed, differences in expressive skill may amplify or distort differences in substantive ability. Distinguishing content from style is therefore essential for understanding how inequality arises, and for identifying where policy interventions can be most effective or warranted.

We find that the writing-score gap between high- and low-SES students is driven primarily by differences in content, which account for roughly 69% of the gap. Differences in style explain about 26%, with the remaining share attributable to differences in scoring functions. This decomposition is stable across demographic subgroups and most prompts.These results underscore that the ability to develop and communicate ideas clearly is both consequential and unevenly distributed, and that systematic differences in this skill can generate large disparities across many domains. From a policy perspective, the findings suggest that interventions focused solely on surface-level mechanics or standardized expression—whether pedagogical or technological—may substantially narrow, but not eliminate, SES disparities in writing assessment. The dominant role of content points to deeper inequalities in the development of argumentation, reasoning, and organization. At the same time, the sizable contribution of style indicates that academic conventions remain consequential and are themselves shaped by unequal access and exposure.

Beyond writing assessment, the paper contributes a general methodological framework for separating tightly confounded components in observational data. Content and style naturally co-move: stronger ideas are often expressed more effectively, and expressive choices influence how substance is perceived. Traditional approaches—statistical controls or human coding—struggle with this entanglement, relying on strong assumptions, limited feature sets, or subjective judgments.

We show how large language models can serve as research instruments to overcome this limitation. By generating multiple stylistic rewrites of each essay that preserve underlying arguments and evidence, we construct a "generated panel" that induces within-essay variation in presentation while holding content fixed. Exposing each essay to a common distribution of styles allows us to separately identify how scorers respond to substance versus form under mild structural assumptions.

This approach has broad applicability. Many social science questions require disentangling idea quality from presentation—whether in job market papers, persuasive communication, or other creative and professional outputs. LLM-generated variation offers a scalable alternative to costly experiments or restrictive modeling choices.

In sum, SES disparities in writing scores reflect primarily differences in argumentative substance rather than stylistic polish, though both matter. More broadly, the paper illustrates how LLMs can be deployed not just as end-user tools but as instruments for scientific measurement, enabling identification strategies that were previously infeasible.