Heterogeneous Peer Effects with Endogenous Network Formation
Source: arXiv:2606.24850 · Published 2026-06-23 · By Duong Trinh, Santiago Montoya-Blandón
TL;DR
This paper addresses the challenge of estimating heterogeneous peer effects in social and economic networks while correcting for endogeneity arising from the endogenous formation of network links. The authors introduce a novel Selection-corrected Heterogeneous Spatial Autoregressive (SCHSAR) model that jointly models network formation and outcome equations, allowing for data-driven latent clustering of units with distinct peer effect parameters. Unlike prior spatial autoregressive (SAR) and peer effect models that assume homogeneous interaction effects and exogenous networks, SCHSAR incorporates unobserved individual-specific heterogeneity affecting both network links and outcome disturbances, thus addressing selection bias. They develop a fully Bayesian MCMC estimation approach relying on data augmentation to overcome the high-dimensional latent variable inference challenges. Extensive simulations demonstrate improved inference accuracy compared to naive methods ignoring endogeneity or heterogeneity. Their empirical application on a large U.S. firm innovation network (n=1150) uncovers two latent firm types with significantly different peer effects and price elasticities on R&D investments, revealing that some firms amplify spillovers while others mainly respond to own cost shocks. This flexible framework produces credible estimates of direct and spillover effects, offering important implications for targeted innovation policy design.
Key findings
- SCHSAR model uncovers two latent firm types in U.S. innovation network: 34% peer-driven with network effect ~0.215 and own-price elasticity ~-2.2, and 66% self-driven with network effect ~0.127 and own-price elasticity ~-9.5.
- Allowing heterogeneous peer effects changes estimated magnitudes substantially compared to homogeneous SAR models, highlighting bias from assuming homogeneity.
- Ignoring endogenous network formation causes omitted variable bias due to latent individual traits correlated between network links and outcomes.
- Bayesian data augmentation MCMC achieves efficient posterior inference by sampling latent variables (network utilities, type indicators, unobserved heterogeneity) jointly.
- Simulation studies show SCHSAR provides valid frequentist coverage and improved parameter recovery over naive approaches.
- Network formation modeled as dyadic strategic link decisions with probit error and individual unobserved degree or homophily effects captures realistic formation mechanisms.
- Posterior predictive distributions enable probabilistic classification of individuals into latent types, facilitating deeper heterogeneity insights.
- The model quantifies firm-level spillin and spillout effects, revealing distinct roles of transmitters and absorbers of peer innovation.
Threat model
n/a; the paper is an econometric methodological contribution modeling social peer effects and endogenous network formation, not a security or adversarial threat study.
Methodology — deep read
The authors propose a two-stage joint econometric framework.
Threat Model & Assumptions: The adversary is absent as this is an econometric model. The key assumption is that individuals form network links endogenously based on observed and unobserved characteristics, which also affect their outcomes. Network formation and outcome equations share latent individual heterogeneity components inducing endogeneity. Peer effects are heterogeneous and latent.
Data: The method is demonstrated via a real dataset of 1,150 U.S. firms engaged in innovation collaborations, with data on firm characteristics, R&D investments, tax prices, and observed network links. The sample is cross-sectional with network adjacency matrix, firm covariates, and outcomes. Labels for latent types are unobserved.
Architecture / Algorithm: The core model consists of:
- Network formation modeled as a dyadic link formation probit model with latent utilities depending on observed dyad covariates and individual unobserved heterogeneity ai (degree heterogeneity or homophily).
- Outcome modeled as a heterogeneous spatial autoregressive (SAR) model with latent types g=1..G. Each latent type has distinct peer effect λg, direct effects βg, contextual effects δg, error variance σu,g, and latent heterogeneity impact κg.
- Individual latent types zi are unobserved multinomial variables estimated together.
- The model links both stages with the same latents ai inducing correlation between network selection and outcomes.
Training Regime: Estimation uses Bayesian MCMC with data augmentation (Albert and Chib 1993 style). Latent utilities for network links w*, type allocations z, unobserved heterogeneity a are sampled jointly alongside parameters {γ, λg, βg, δg, κg, σ2u,g, πg}. Priors are placed on parameters and mixture probabilities πg. The MCMC iterates between updating latent variables and parameters conditional on others, allowing efficient sampling despite high dimensionality.
Evaluation Protocol: The method is validated via Monte Carlo simulations comparing SCHSAR to naive SAR ignoring endogeneity and heterogeneity, showing improved bias and coverage. The empirical application quantifies type-specific parameters with credible intervals, and examines firm-level spillin/spillout effects. Network formation model fit and posterior predictive checks support model adequacy.
Reproducibility: Code and data provision not explicitly stated and likely proprietary for firm data. The Bayesian algorithm is described sufficiently to be replicable given suitable data. The latent variable augmentation approach follows established Bayesian econometric routines.
Example: For a given i, the observed outcome Yi depends on weighted outcomes of neighbors via W, with weight λg depending on latent type g. The network W is formed by dyadic probit decisions with utilities driven by observed dyadic covariates and latent social capital ai that also enters Yi equation through κg. MCMC samples latent w*ij (network utilities), ai (heterogeneity), and zi (type indicators) alternating with model parameters, enabling joint inference that accounts for selection and heterogeneity simultaneously.
Technical innovations
- Introduction of a finite mixture structure within a spatial autoregressive framework enabling data-driven latent clustering for heterogeneous peer effects.
- Joint econometric modeling of network formation and outcome determination correcting for endogeneity arising from shared latent traits.
- Fully Bayesian data augmentation MCMC algorithm for simultaneous estimation of latent types, network formation latent utilities, and outcome parameters overcoming high-dimensional integration challenges.
- Extension of spatial autoregressive models to allow endogenous network weights rather than assuming exogenous fixed network structure.
- Incorporation of unobserved individual-specific degree heterogeneity or homophily effects into network formation driving endogenous link formation.
Datasets
- U.S. innovation collaboration network — 1,150 firms — proprietary dataset from empirical application in Section 5
Baselines vs proposed
- Naive homogeneous SAR ignoring endogeneity: biased peer effect estimates with no credible uncertainty quantification vs SCHSAR: valid heterogeneity estimates and corrected bias (exact metric deltas not stated).
- Naive SAR ignoring network endogeneity: omitted-variable bias in λ and other parameters indicated by simulation vs SCHSAR: consistent estimates with frequentist coverage demonstrated in simulations.
Figures from the paper
Figures are reproduced from the source paper for academic discussion. Original copyright: the paper authors. See arXiv:2606.24850.

Fig 1: Distribution of R&D intensity among firms exhibits visible multimodality.

Fig 2: (a) Direct effects of a 1% reduction in a firm’s own R&D tax price. (b) Indirect spillin

Fig 3: (a) Total spillin effects on each firm due to a 1% reduction in the R&D tax price for all
Limitations
- The methodology requires large sample and rich network data to reliably identify latent clusters and correct for complex endogeneity.
- The assumption of normally distributed latent unobserved heterogeneity or categorical latent types may be restrictive in some applications.
- Empirical application is cross-sectional, lacking dynamic or panel data to assess temporal evolution of networks and peer effects.
- No explicit adversarial or causal robustness tests showing how well model copes with strategic manipulation or model misspecification.
- Computational complexity and scalability for very large networks are not discussed in detail.
- Data used in empirical study is proprietary, limiting immediate reproducibility.
Open questions / follow-ons
- Extension of SCHSAR to dynamic or panel network data to model temporal evolution of peer effects and link formation.
- Incorporation of richer latent trait models beyond finite mixture or Gaussian assumptions to capture more complex heterogeneity.
- Development of scalable approximate Bayesian inference methods to enable application on ultra-large networks.
- Investigation of robustness to model misspecification and strategic manipulation in network formation.
Why it matters for bot defense
While not directly connected to bot-defense or CAPTCHA technologies, the SCHSAR framework provides valuable insights for practitioners interested in modeling interactions in user or device networks with endogenous link formation. Understanding heterogeneous peer effects with correction for selection bias is relevant when analyzing user behaviors influenced by connected peers in partially observed or strategically formed interaction networks. This approach may inspire advanced models for detecting coordinated bot attacks or adversarial behaviors where peer influences and network structure evolve endogenously. Its Bayesian data augmentation techniques for latent variable inference could inform probabilistic modeling of bot or human clusters amid noisy network observations. Hence, bot-defense engineers could consider analogous joint modeling frameworks to capture heterogeneity and endogeneity in interaction graphs underlying botnet or CAPTCHA puzzle-solving networks. However, direct application requires adaptation to security domains and higher-frequency dynamic data. The paper's emphasis on carefully addressing selection bias and latent heterogeneity is a critical modeling principle transferable to bot defense analytics.
Cite
@article{arxiv2606_24850,
title={ Heterogeneous Peer Effects with Endogenous Network Formation },
author={ Duong Trinh and Santiago Montoya-Blandón },
journal={arXiv preprint arXiv:2606.24850},
year={ 2026 },
url={https://arxiv.org/abs/2606.24850}
}