Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
70,002 characters · 22 sections · 30 citation commands
Digital Divide: Evidence from the 2020 Canadian Internet Use Survey
What we know about the Canadian digital divide. There is broad agreement that digital participation in Canada is unequally distributed along familiar socioeconomic lines. Haight-et-al(2014), using the 2010 CIUS, show that age, income, and education are the primary predictors of internet use and that the digital divide extends well beyond infrastructure access to encompass the skills and willingness needed for meaningful engagement. Wavrock_2022 document significant differences in the breadth and depth of digital engagement across age, education, and income groups in the 2018 CIUS, while administrative sources show that the urban--rural gap in broadband access remains substantial despite federal investments under the Universal Broadband Fund CRTC2023, ISED2024. The resulting picture is one of layered disadvantage: low-income households, older adults, and individuals with limited educational attainment are less likely to access the internet and, conditional on access, less likely to use it for financially or socially consequential activities VDVD2019, TheDais2024.
\noindentWhat distinguishes Canada. Canada's digital inequality challenge has features that set it apart from other developed economies. The central issue is not absent infrastructure: by 2020, 92.2% of CIUS respondents reported using the internet. Instead, inequality reflects differences in the cost, reliability, and quality of connections across geographic areas; the absence of a coordinated national digital literacy strategy (Canada remains one of the few OECD countries without one); and the concentration of exclusion among specific populations --- Indigenous communities, persons with disabilities, low-income seniors, and recent immigrants TheDais2024, ESDC2024, Kwok-Korpela-2024 --- who are also among those most reliant on digital services to access health information, government benefits, and financial services as physical alternatives continue to decline.
\paragraph*{Four gaps this paper addresses.} Despite a growing descriptive literature, four questions remain unanswered or only partially answered.
\paragraph*{Contributions.} The four research questions map directly onto four empirical contributions using the 2020 CIUS, a nationally representative cross-sectional survey of 17,409 Canadians aged 15 and older across the ten provinces.
\paragraph*{Main findings.} The svy LLasso identifies age, education, and employment as the most consistent predictors of digital financial service adoption, with education the only covariate that remains significant at every rung of the digital ladder. Income-related inequality is significant across all digital financial services and is especially pronounced for virtual-wallet adoption; for online banking, employment and education together account for nearly half of the pro-rich concentration, indicating a broad socioeconomic gradient rather than a purely income-based divide.
The age--education decomposition shows that income, disability, health, and preferences together explain a substantial but incomplete share of the observed gaps, with preferences the dominant mediator and large residual gaps persisting across all cells. The sequential logit shows that disadvantage emerges at different stages for different groups: education generates significant barriers throughout the ladder, disability produces its largest penalty at the digital-payments stage, and the most excluded profiles --- older, less-educated respondents with low income or non-employment --- face compounding barriers at multiple rungs.
The IRT analysis shows that digital literacy eliminates the education gradient at internet entry and substantially reduces it at the online banking rung, but a residual persists, pointing to behavioral and institutional frictions beyond measurable competence. The IRT analysis also reveals counterintuitive patterns: a Gen Z deficit in structured information-seeking and a security-behavior gap concentrated among recent immigrants and visible minorities rather than among seniors.
\paragraph*{Paper organization.} Section (ref) describes the CIUS 2020 data, and Section (ref) sets out the analytical frameworks. Section (ref) presents the svy LLasso results and concentration-index analysis (Q.1), while Section (ref) reports the age--education decomposition (Q.2). Section (ref) presents the sequential logit analysis (Q.3), and Section (ref) reports the IRT-based digital literacy measure and mediator analysis (Q.4). Section (ref) concludes. Additional technical details, supplementary results, and robustness checks are provided in the appendix.
We use the publicly available 2020 Canadian Internet Use Survey public-use microdata file (PUMF).\footnote{The dataset is available at \url{https://abacus.library.ubc.ca/dataset.xhtml?persistentId=hdl:11272.1/AB2/NUVBX2}.} The PUMF is close in scope to the Statistics Canada Research Data Centre version, although some variables are available only in more aggregated form. CIUS 2020 comprises 17,409 observations on Canadians aged 15 and older living in one of the ten provinces. First Nations persons living on reserve are excluded by the sampling frame. The survey uses a stratified sampling design at the province/census-metropolitan-area/census-agglomeration level and combines landline and cellular telephone frames; the overall response rate is 41.6%. All estimates use the person weight WTPG, which reflects the adjusted selection probability, non-response adjustment, and calibration to independent population totals.\footnote{ Precise variable definitions, item codes, and further information on sampling and weighting correspond to the publicly available CIUS 2020 codebook and User Guide.}
The three main outcomes used in Sections (ref)--(ref) are internet use, online banking, and digital payments. Internet use is defined on the full sample. Online banking is defined for internet users. The digital payments outcome indicates whether the respondent used a virtual wallet or credit card for online purchases; this outcome is defined for respondents who ordered goods or services online. The svy LLasso analysis also considers virtual wallet and credit card use as separate outcomes; the corresponding results are reported in Appendix (ref).\footnote{Email use is not reported as a standalone outcome because it is nearly universal among internet users; it enters the analysis as the second rung of the sequential logit in Section (ref).}
The Q.2 decomposition uses two unconditional outcomes: online banking and online health-information search. For these outcomes, non-internet users are coded as zero. Wavrock_2022 highlight these measures as especially policy-relevant and well suited to the survey's sequential routing structure.
The explanatory variables include household income quintile, educational attainment, employment status, age group, gender, Indigenous identity, visible minority status, immigration status, language, household composition, disability status, self-reported health, preference for non-use, and province. This covariate set is used across all empirical analyses, with the svy LLasso specification allowing for a flexible variable-selection step.
To address Q.1, we adopt the svy LLasso of JMT2026, which obtains the survey-weighted logistic Lasso estimator $\hat{\theta}$ by minimizing the penalized weighted negative log-likelihood:
where $\beta=(\beta_1,\dots,\beta_p)'$,
is the weighted log-likelihood of the logit model, and $w_i$ is the survey weight. The intercept $\alpha$ is left unpenalized, and the $\ell_1$ penalty shrinks weak predictors toward zero. The tuning parameter $\lambda$ is chosen by ten-fold cross-validation using the AUC criterion, implemented in the R package glmnet; person weights are rescaled to mean one within the estimation sample. In Appendix (ref), we consider two alternative tuning rules: a design-aware weighted cross-validation procedure Iparragirre2023 and the bootstrap-after-cross-validation rule of Chetverikov2025. Both yield very similar empirical patterns.
Following JMT2026, we apply the debiased Lasso correction to conduct valid post-selection inference:
where $H(\cdot)$ and $S(\cdot)$ denote the sample Hessian and score of the full weighted logistic log-likelihood in (ref). This one-step estimator removes regularization bias and is asymptotically normal, facilitating standard $t$-ratio inference on coefficients and AMEs. Throughout, we report $\tilde{\theta}^{\mathsf{DB}}$ together with the debiased estimates of the AMEs, denoted by $\widetilde{\mathrm{AME}}^{\mathsf{DB}}$; technical details are given in Appendix (ref).
To address Q.2, we estimate two nested survey-weighted logit models for each outcome $y_i$:
where $\theta_g$ denotes the coefficient for the age$\,\times\,$education cell $g$ relative to the omitted reference cell of adults aged 65 and older with high school or less education. The control vector $x_i$ includes household-size-adjusted income quintile, a disability indicator, indicators for self-reported health, and a preference indicator for non-use due to lack of interest or time.
Let
denote the conditional probability for cell $g$. We partition the covariate vector into four explanatory blocks,
corresponding to income, disability, health, and preferences. We define the characteristics-explained component of the age--education gap as
where $x_g$ denotes the observed covariate values for respondents in cell $g$, and $x_r$ denotes the corresponding values for respondents in the reference cell. Thus, $m_g(x_g)$ is the average fitted probability evaluated at the observed characteristics of group $g$, whereas $m_g(x_r)$ evaluates the same fitted probability after replacing those characteristics with the reference-group values.
To allocate $E_g$ across the four explanatory blocks, we implement the exact simulation-based Shapley decomposition of Shorrocks2013, which adapts the Shapley value Shapley1953 to regression decomposition. Let $K$ denote the set of blocks. For any subset $S \subseteq K$, let $m_g(x_r^{(-S)},x_g^{(S)})$ denote the fitted probability obtained by assigning the blocks in $S$ their group-$g$ values and all remaining blocks their reference-group values. Then the contribution of block $k$ to the explained component for cell $g$ is
where the bracketed term is the marginal contribution of block $k$ when it is added after the blocks in $S$ have already been switched from reference-group values to group-$g$ values. By construction, the decomposition is exact,
and invariant to the ordering of the blocks because it averages marginal contributions over all possible permutations. These properties make it especially attractive in the present nonlinear logit setting, where first-order linearization can perform poorly when fitted probabilities lie near the boundaries of the unit interval.
We estimate the decomposition on three nested samples --- urban non-disabled respondents, all urban respondents, and the full sample --- to distinguish connectivity constraints from disability-related barriers. The underlying logit models are estimated using svyglm in the R package survey, and the Shapley decomposition is then applied to the fitted values.
The CIUS routing structure described in Section (ref) implies a natural four-rung adoption ladder: internet use, email use, online banking, and digital payments. We exploit this structure by estimating a continuation-ratio logit in which each stage is modelled as a binary logit on the subsample that has cleared the previous gate.
Let $p_{g,1}$ denote the probability that group $g$ clears Stage 1, and let $c_{g,j}$ denote the conditional probability of clearing Stage $j$ given clearance of Stage $j-1$. The unconditional probability of reaching the digital-payments rung is $\pi_{g,4}=p_{g,1}c_{g,2}c_{g,3}c_{g,4}$. The gap relative to a benchmark group $\bar g$ admits the exact sequential decomposition
obtained by sequentially replacing the benchmark-stage probabilities with those of group $g$ from Stage 1 onward. The weighting terms reflect the sequential nature of the process: gaps at later stages matter only for the subset that clears the earlier rungs. The same covariate vector is used at each stage, and each stage is estimated using svyglm.
To measure digital competence, we estimate a weighted bifactor two-parameter logistic (2PL) item-response model Birnbaum1968 on 20 binary CIUS items covering information-seeking, software and file management, and security and privacy tasks. Let $y_{ij}\in\{0,1\}$ denote the response of individual $i$ to item $j$, where $y_{ij}=1$ if the respondent reports performing the corresponding digital task and $y_{ij}=0$ otherwise. The model includes one latent general factor, interpreted as overall digital literacy, and three latent domain-specific factors that capture residual covariance within the information-seeking, software/file-management, and security/privacy domains. The response probability is specified as
where $\Lambda(\cdot)$ is the logistic CDF, $\theta_i^{(G)}$ is the latent general factor for individual $i$, $\theta_i^{(D_j)}$ is the latent domain-specific factor for the domain to which item $j$ belongs, $a_j^{(G)}$ and $a_j^{(D_j)}$ are item discrimination parameters, and $b_j$ is an item difficulty parameter. The latent factors are assumed to be mutually orthogonal and normalized to unit variance for identification. The bifactor structure is identified without rotation because the general factor loads on all items and each domain-specific factor loads only on items within its domain Reise2012.
Our digital literacy score is the estimated general-factor score $\hat{\theta}_i^{(G)}$, which summarizes each respondent's overall digital competence net of domain-specific residual variation. Given the estimated item parameters $\hat{\psi}=\{\hat{a}_j^{(G)},\hat{a}_j^{(D_j)},\hat{b}_j\}_{j=1}^{20}$, this score is computed as the expected a posteriori (EAP) estimate (see Appendix (ref) for the formal definition), obtained using the EM or quasi-Monte Carlo EM algorithm implemented in the R package mirt Chalmers2012, with survey weights normalized within the estimation sample. For use in the concentration-index and sequential-logit analyses, we rescale the estimated general-factor score $\hat{\theta}_i^{(G)}$ to the unit interval and denote the resulting literacy index by $\hat{L}_i\in[0,1]$.
The standardized loadings reported in Appendix (ref) are obtained via the Schmid--Leiman orthogonalization Schmid1957, which decomposes the factor solution into a general factor and orthogonal domain-specific residual factors, ensuring that $\hat{a}_j^{(G)}$ reflects the unique contribution of the general factor to item $j$ net of domain-specific variance.
Item selection and dimensionality diagnostics are reported in Appendix (ref). Reliability is assessed using McDonald's hierarchical omega McDonald1999, which measures the proportion of composite-score variance attributable to the general factor net of domain-specific variance; values above 0.70 are generally taken as evidence that a single general factor dominates reliable variation.
To complement the regression analysis with a distributional perspective, we use concentration indices to measure whether digital outcomes are disproportionately concentrated among respondents higher in the income or digital-literacy distribution.
\paragraph*{Income-ranked concentration index.} Let $R_i$ denote respondent $i$'s fractional rank in the household-income distribution, and let $\mu_y=\mathrm{E}[y_i]$ denote the population mean of outcome $y_i$. The population concentration index is
A positive value indicates that the outcome is concentrated among higher-income respondents. In the data, we estimate $\mathcal{C}_y$ using survey weights. With $w_i>0$ denoting the person weight, let $ W:=\sum_{i=1}^n w_i$ be the total survey weight, and let
and $\hat r_i$ denote the weighted sample mean of $y_i$ and the weighted midpoint fractional rank in the income distribution, respectively. The survey-weighted sample analogue is
Because $y_i$ is directly observed, the influence function of $\widehat{\mathcal{C}}_y$ admits a closed-form linearization and inference is based on the resulting first-order variance estimator Kakwani1997.\footnote{In the implementation, we use the equivalent three-moment representation $\widehat{\mathcal{C}}_y = 2\bar{y}^{-1}W^{-1}\sum_i w_i y_i \hat{r}_i - 1 - 2(W^{-1}\sum_i w_i \hat{r}_i - \tfrac{1}{2})$, which reduces to the formula (ref) when $W^{-1}\sum_i w_i \hat{r}_i = \tfrac{1}{2}$ exactly. After excluding observations with missing outcome values, the weighted mean rank in the estimation sample need not equal $\tfrac{1}{2}$, so the three-moment form is used throughout.}
Since the outcomes of interest are binary, the standard concentration index is bounded by $[-(1-\mu_y),\,1-\mu_y]$ rather than $[-1,1]$, which complicates comparisons across outcomes with different means Wagstaff2005. We therefore also report two normalized versions. Following Wagstaff2005, the Wagstaff-normalized index is
and, following Erreygers2009, the Erreygers index is
The income-ranked decomposition in Section (ref) uses the standard index, which admits the standard linear decomposition of Wagstaff2003.
\paragraph*{Literacy-ranked concentration index.} To study how digital activities are distributed across the competence distribution, we define a literacy-ranked concentration index by replacing the income rank with the weighted rank of the estimated general-factor digital literacy score. Let $r_i^L$ denote respondent $i$'s weighted midpoint fractional rank in that literacy distribution. The corresponding population index is
with survey-weighted sample analogue
A positive value indicates that the outcome is concentrated among the more digitally literate.
Because the literacy rank is constructed from an estimated bifactor IRT model, $\widehat{\mathcal{C}}_y^L$ is a two-step estimator. Inference therefore uses Murphy--Topel-type standard errors MurphyTopel1985 that account for both second-stage sampling variation and first-stage IRT estimation uncertainty; details are in Appendix (ref).
Table (ref) presents the svy LLasso results for the online banking model. Age is the strongest predictor: relative to the reference group aged 45--54, respondents aged 25--34 and 35--44 are 7 and 6 percentage points more likely to use online banking, while those aged 55--64 and 65 and older are 3 and 9 percentage points less likely. Employment status is also important: employed respondents are about 9 percentage points more likely to use online banking ($\widetilde{\mathrm{AME}}^{\mathsf{DB}}=0.09$, $p<0.001$), consistent with payroll direct deposit and employer-linked financial access. Educational attainment generates a clear gradient: high school (HS) or less reduces the probability of online banking by 7 percentage points, while a university degree raises it by 5 percentage points relative to some post-secondary. Visible minority status is associated with a 5 percentage point lower probability ($p=0.003$). Among provinces, Manitoba is associated with a 5 percentage point lower probability ($p=0.010$) and Quebec with a 5 percentage point higher probability ($p=0.021$). Gender, rural residence, Aboriginal identity, immigration status, and disability are not statistically significant after conditioning on the full covariate set.
Language of use is also a notable predictor. Relative to respondents reporting neither English nor French, English speakers and English-French speakers are 17 and 13 percentage points more likely to use online banking, respectively. These are among the largest AMEs in the model, though their magnitude partly reflects the small and linguistically heterogeneous reference category rather than a uniform language penalty. Within the Canadian financial system, online banking interfaces have historically been designed around English and French, and some platforms impose official language requirements for digital enrollment. The language gradient is therefore consistent with a combination of interface accessibility barriers and the occupational and socioeconomic sorting that correlates with language of use. The French-speaker coefficient is positive but falls short of conventional significance ($p=0.178$), consistent with Quebec's above-average online banking rate.
Table (ref) presents the svy LLasso results for virtual wallet and credit card use. For virtual wallet adoption, age remains the dominant predictor: those aged 15--24, 25--34, and 35--44 are 11, 8, and 4 percentage points more likely to adopt relative to the reference group, while those aged 55--64 and 65 and older are 6 and 8 points less likely. Rural residence reduces adoption by 5 points. Unlike in the online-banking model, visible minority status has a positive association ($\widetilde{\mathrm{AME}}^{\mathsf{DB}}\approx 0.04$, $p<0.01$), pointing to a distinct adoption pattern rather than a general financial-inclusion gradient. University education and top-quintile income are also positively associated, with AMEs of about 2 and 8 percentage points.
For credit card use, the youngest group (15--24) is about 9 percentage points less likely to use a credit card online. Education generates the sharpest gradient: high school or less reduces the probability by 7 percentage points, while a university degree raises it by 7 points. Family households without children and Ontario residence are positively associated, while disability is associated with a 7 percentage point reduction and Quebec with a 6 point reduction. Visible minority status is weakly negative for credit card use, in contrast to its positive association with virtual wallet adoption. This sign reversal is consistent with substitution between payment instruments and differential access barriers across socioeconomic groups.
Table (ref) reports the standard concentration index together with the Wagstaff-normalized and Erreygers versions for online banking, virtual wallet use, credit card use, and the composite digital-payments outcome. Inference for the standard concentration index is based on a first-order linearization estimator.
Online banking is significantly concentrated among higher-income respondents: $\widehat{\mathcal{C}}_y=0.027$ (SE $=0.003$; 95% CI $[0.020,\,0.034]$; $p<0.001$). Panel B shows that the Wagstaff2003 decomposition associates 33.1% of this concentration with income, 26.6% with employment, 22.3% with education, and 14.4% with age. Employment and education together account for 48.9% --- nearly one half --- while income, although the single largest component, does not dominate the decomposition on its own. Inequality in digital banking access therefore reflects a broad socioeconomic gradient rather than a purely income-based divide.
The comparison across digital-finance outcomes in Panel A reveals a much sharper income gradient for newer payment technologies. Virtual-wallet use is far more concentrated among higher-income respondents ($\widehat{\mathcal{C}}_y=0.106$, SE $=0.023$, 95% CI $[0.061,\,0.152]$; $p<0.001$) than online banking ($0.027$), credit card use ($0.023$), or the broader digital-payments measure ($0.024$). The grouped decompositions in Panel B show that income accounts for 93.0% of the measured concentration in virtual-wallet use, compared with 51.6% for credit card use and 57.8% for the composite digital-payments measure. The role of income therefore becomes much more prominent as one moves from general digital banking toward newer payment instruments.
This section focuses on unconditional digital use. Unlike conditional measures, unconditional outcomes assign zero to those who do not reach a given stage, thereby capturing cumulative barriers along the full pathway to digital participation. Unadjusted rates and gaps relative to the reference cell are reported in Appendix (ref); for online banking, rates range from 0.381 for the reference cell to 0.977 for the 15--24 $\times$ University cell, an unadjusted gap of 0.596.
Table (ref) reports the exact Shapley decomposition of age--education gaps after conditioning on income, disability, health, and preferences; results for health-information search are included alongside online banking as a cross-outcome robustness check and show qualitatively similar patterns throughout. Adjustment reduces the gaps but does not eliminate them. In the full sample, the adjusted gap in online banking remains 0.379 for the 15--24 $\times$ University cell, 0.261 for the 25--44 $\times$ University cell, and 0.225 for the 45--64 $\times$ University cell.
Among the mediators, preferences account for the largest explained share in most cells. Income contributes a meaningful but smaller portion. Disability and health effects are generally modest; health contributions are sometimes slightly negative, indicating that differences in self-reported health do not reinforce the age--education gradient once other covariates are held fixed.
The total explained component is substantial but incomplete: it is 0.202 for the 15--24 $\times$ University cell and 0.247 for the 25--44 $\times$ University cell in online banking, leaving sizable residual gaps. These persistent adjusted gaps point to digital skills, confidence, interface familiarity, and other unobserved barriers. Subsample results for urban and urban non-disabled respondents are qualitatively similar and are reported in Appendix (ref). Section (ref) uses the IRT literacy score to assess how much of the residual is accounted for by measurable digital competence.
Figure (ref) reports average marginal effects (AMEs) from the four sequential logit models. The sequential logit is estimated as a continuation-ratio model with one binary logit at each rung of the adoption ladder. Model fit declines monotonically across stages: McFadden's pseudo-$R^2$ falls from 0.29 at internet entry to 0.12 for email use, 0.09 for online banking, and 0.06 for digital payments. Sociodemographic characteristics therefore explain a substantial share of variation in who gets online, but progressively less of the variation in higher-order activities. This pattern is consistent with early barriers reflecting structural constraints, while later stages depend more on trust, preferences, and transaction-specific behavior. Full goodness-of-fit statistics are reported in Appendix (ref).
The ladder is characterized by different barriers at different stages. Low educational attainment is the only disadvantage that remains significant throughout the entire sequence, from internet access to digital payment adoption, suggesting that education captures not only access differences but also the skills and confidence needed for progressively more demanding digital tasks.
Income-related gaps become sharper at later stages, with the role of financial capacity most pronounced for digital-payment adoption. Rural disadvantage shows the opposite pattern: it is concentrated at the early stages --- especially internet and email use --- and largely disappears once earlier access barriers are cleared. Gender is largely neutral across the ladder, with one exception: female respondents are 1.6 percentage points more likely to use email conditional on internet access ($p = 0.022$), consistent with occupational sorting toward communication-intensive roles.
Age-related gaps are largest at initial entry and narrow at later stages, suggesting that older adults face stronger barriers to getting online than to using digital transactions conditional on access. Disability, by contrast, is associated with negative effects across multiple stages, consistent with barriers that persist beyond entry. Figure (ref) in Appendix (ref) translates these stage-specific marginal effects into predicted adoption trajectories for six representative profiles, illustrating how multiple disadvantages compound across the ladder.
Figure (ref) traces the cumulative share of each demographic group reaching each rung.
\paragraph*{Income inequality becomes most visible at the digital payments rung.} Income quintile groups remain close through Stages 1--3 but diverge sharply at Stage 4: roughly 42% of respondents in Income Q1 reach digital payments, compared with about 70% in Income Q5. This gap reflects accumulated advantages rather than a single transition: Income Q5 raises the probability of internet use by 4.5 points, email by 2.4, online banking by 3.4, and digital payments by 4.3. Income Q1, by contrast, reduces internet adoption by 4.0 points while showing little additional conditional disadvantage at later stages.
\paragraph*{Education is the one disadvantage that persists at every rung.} High school or less generates a significant negative AME at all four stages: $-0.082$, $-0.044$, $-0.057$, and $-0.053$. A university degree produces significant positive increments at every rung: $0.034$, $0.056$, $0.034$, and $0.060$. No other covariate produces significant effects uniformly across all four stages. The persistence of education --- even after controlling for income, age, employment, disability, and health --- motivates the IRT mediator analysis in Section (ref).
\paragraph*{Disability generates compounding but non-monotonic disadvantage.} The AMEs are $-0.023$ at Stage 1, $-0.037$ at Stage 2, $-0.008$ at Stage 3 (not significant), and $-0.082$ at Stage 4. The pattern is cumulative but not monotonic. Online banking interfaces appear relatively more standardized for persons with disabilities, whereas retail payment checkout flows remain difficult to navigate. This points to the importance of accessibility enforcement beyond core banking platforms.
\paragraph*{Employment is specific to the banking rung.} Employment shows its largest AME at Stage 3 ($0.078$, $p<0.01$), while its effect at Stage 4 is essentially zero. The mechanisms connecting employment to banking --- payroll direct deposit, benefits management, employer-provided financial portals --- appear specific to the banking rung and do not carry through to retail payment adoption.
\paragraph*{Age effects and geographic heterogeneity.} Respondents aged 65 and older are 11.8 percentage points less likely to use the internet, 2.8 points less likely to use email, and 8.0 points less likely to use online banking, with no significant difference at digital payments. The 15--24 group shows a distinct profile: substantially more likely to use the internet ($0.069$) but less likely to use online banking ($-0.035$). Age therefore operates primarily through entry and banking barriers rather than a uniform resistance to all digital activities.
Rural residence reduces internet use by 3.0 points and email use by 2.3 points but has no meaningful effect at later stages, indicating that rural disadvantage in the PUMF data operates mainly through initial connectivity. Quebec displays a distinctive profile: negative effects at Stages 1--2, a positive banking effect ($0.057$), and a negative digital payments effect ($-0.073$), consistent with strong banking-platform penetration alongside lower adoption of newer payment methods.
To identify which combinations of characteristics jointly characterize the bottom tail of predicted digital payment adoption, we examine the modal characteristics among respondents in the bottom decile of the predicted adoption distribution, retaining those whose weighted modal share exceeds 0.80 (0.75 for virtual wallet). Appendix (ref) reports the corresponding profile summaries and predicted probabilities.
The virtual wallet exclusion profile is especially stark: respondents aged 65 and older, with high school or less education, in the lowest income quintile, and with non-visible minority status have a predicted adoption probability of only 0.010, against a weighted national average of 0.097 ($n=244$). For credit card and composite digital payment outcomes, the most excluded profile converges on older, non-employed, less-educated respondents, with predicted adoption probabilities of 0.315 and 0.316 against national averages of 0.595 ($n=671$) and 0.608 ($n=693$) respectively.
Across technologies, digital payment exclusion operates through overlapping socioeconomic disadvantages (especially age, education, and weak economic attachment) rather than any single barrier, while the most excluded subpopulations differ meaningfully across payment instruments.
This section uses the digital literacy score from Section (ref) to assess whether gaps in digital adoption reflect broad differences in digital competence or more specific domain-level deficits. The score is constructed from 20 binary CIUS items spanning information-seeking, software and file management, and security and privacy tasks.
The diagnostics support the use of the general-factor score as our main summary measure: the ratio of the first to second tetrachoric eigenvalue is 5.18 and McDonald's $\omega_h=0.76$, indicating a dominant general factor. The loading pattern is also substantively coherent: software and file-management items load most strongly on the general factor, while security and privacy items retain an additional domain-specific component. A scree plot, full loading estimates, and additional dimensionality diagnostics are reported in Appendix (ref).
Figure (ref) plots survey-weighted domain scores by age group. The general score exhibits a pronounced age gradient: the 25--34 group records the highest mean (0.275), while the 65-plus group records the lowest ($-0.509$), a gap of 0.784 standard deviations. The decline is modest through middle age but steepens markedly after age 55, driven largely by the software and file-management dimension, where scores fall below zero for the 55--64 group ($-0.104$) and decline further for the 65-plus group ($-0.207$).
The information-seeking dimension tells a different story. The youngest cohort (15--24) records the lowest information-seeking score ($-0.206$), while the 35--44 group records the highest (0.135). This suggests that younger Canadians may be highly engaged digitally without being equally proficient in structured information-search tasks --- locating services, comparing options, or navigating institutional websites --- at which middle-aged respondents perform better.
Security behavior exhibits a weaker age gradient. Most age groups cluster near the population average, but the 55--64 cohort scores significantly above zero on the security subscore (0.050, 95% CI [0.016, 0.084]), possibly reflecting workplace exposure to formal IT security practices during the organizational expansion of the late 1990s and early 2000s. Figure (ref) shows that the security dimension does not follow a simple socioeconomic gradient: visible minorities and landed immigrants score significantly lower, whereas rural residents score significantly higher. These patterns suggest that digital-inclusion policies should address privacy-risk awareness and protective practices, not only access and general skills.
Digital literacy itself is also unequally distributed across income. Applying the income-ranked concentration index of Section (ref) to the rescaled literacy score $\hat{L}_i$ (Panel A of Table (ref)) yields $\widehat{\mathcal{C}}_L = 0.029$ ($\mathrm{SE} = 0.003$, $p < 0.001$), confirming a statistically significant pro-rich gradient in digital competence and motivating the literacy-ranked analysis below.
The persistent education gradient in the sequential logit may reflect differences in digital capability, connectivity quality, or behavioral preferences. To assess the capability channel, we augment the sequential logit with the general digital literacy score and compare the AME of the high-school-or-less indicator before and after conditioning.
Table (ref) reports results for Stages 1 and 3. At Stage 1 (internet use), the baseline education gap is about 5.5 percentage points; after conditioning on literacy, the effect becomes statistically indistinguishable from zero, indicating that the initial education gradient largely reflects differences in digital capability. At Stage 3 (online banking), conditioning on literacy reduces the gap by 61%, but the effect remains significant at the 5% level, pointing to behavioral and institutional mechanisms --- trust, financial interface familiarity, and perceived risk --- that continue to shape adoption beyond basic skill. The literacy score itself raises the probability of online banking by 7.2 percentage points per standard deviation.
These are descriptive conditioning effects rather than causal mediation estimates, since digital literacy is jointly determined with education and other socioeconomic characteristics. Even so, the attenuation pattern is informative: capability accounts for most of the education gradient at entry, while behavioral and institutional frictions dominate at higher rungs.
Table (ref) reports the literacy-ranked concentration indices defined in Section (ref) for six digital activities. All six indices are positive, indicating that digital participation is systematically tilted toward the more digitally literate. Near-universal activities show the smallest gradients: email use ($\widehat{\mathcal{C}}_y^L = 0.026$) and smartphone use ($0.030$) are only mildly concentrated at the upper end of the literacy distribution. By contrast, routine transactional activities display moderate and tightly clustered gradients: online banking ($0.066$), credit-card use ($0.062$), and government online services ($0.067$) have very similar concentration indices, suggesting a common competence threshold for mainstream digital transactions.
Virtual-wallet adoption stands apart. Its literacy-ranked concentration index is $\widehat{\mathcal{C}}_y^L = 0.344$, with a Murphy--Topel standard error of $0.023$ and a 95% confidence interval of $[0.298,\;0.389]$, against a base rate of only 12.3%. Part of this larger value reflects the $1/\bar{y}$ scaling of the concentration index for relatively rare outcomes, but the contrast with credit-card use --- a similarly payment-oriented activity with a much higher base rate and an index of only $0.062$ --- still points to a substantially stronger literacy gradient for newer payment instruments. Taken together, these results place digital competence as a distributional link between economic resources and digital financial inclusion, complementing the income-ranked indices in Section (ref) and the conditioning evidence above.
This paper studies digital inequality in Canada using four complementary empirical approaches: survey-weighted logistic Lasso for digital financial adoption, an exact Shapley decomposition of age--education gaps, a sequential logit tracing where disadvantage emerges along the adoption path, and a bifactor IRT measure of digital literacy. The results show that digital inequality is multidimensional: income, education, disability, age, and digital capability matter in different ways and at different stages of participation.
Three conclusions stand out. First, education is the only determinant that remains significant throughout the full adoption ladder; this effect is only partly accounted for by income, health, disability, preferences, and measured digital literacy, pointing to behavioral and institutional frictions that persist beyond measurable competence. Second, income-based inequality is most pronounced for virtual-wallet adoption, where both the income and literacy gradients are especially steep, and for online banking the concentration reflects a broad socioeconomic gradient in which employment and education are nearly as important as income itself. Third, the disability penalty is stage-specific rather than uniform: it disappears at online banking but reappears with the largest penalty at the digital-payments stage, pointing to accessibility gaps in retail payment interfaces that banking platforms have largely addressed. The digital literacy analysis further reveals that general capability and security behavior follow different social patterns, with security deficits concentrated among recent immigrants and visible minorities rather than among seniors --- a finding with direct implications for digital-inclusion policy.
\paragraph*{Limitations.} Three limitations warrant acknowledgment. The CIUS 2020 covers only the ten provinces, excluding the territories and First Nations reserves, so the findings do not generalize to those populations. The digital literacy score is estimated on the routed subsample of internet users, which may introduce selection into capability measurement and limits the comparability of literacy-conditioned results with the full-population analyses. The decomposition and literacy-conditioning exercises are associational rather than causal; they identify plausible mechanisms but cannot rule out confounding. Future work using the CIUS 2022 wave could assess whether the stage-specific barriers documented here have persisted or shifted in the post-pandemic period.