Source-linked AI summary
Point process modeling for directed interaction networks
Patrick O. Perry, Patrick J. Wolfe
TL;DR
Repeated directed interactions raise questions about which traits and behaviors predict interaction, including homophily, dynamic network effects, and multicast structure. The paper develops a Cox point-process framework with history-dependent covariates and partial-likelihood inference, proves consistency and asymptotic normality under regularity conditions, and extends the approach to multicast interactions. Applied to corporate e-mail, the framework quantifies static shared traits and dynamic network effects associated with recipient selection.
Problem
The paper asks which sender, receiver, and network characteristics predict directed interaction, including shared attributes, past interaction behavior, and multiple-receiver events.
Method
It models directed interactions as a multivariate point process using a Cox proportional intensity model with predictable, history-dependent covariates and partial-likelihood inference, including a multicast extension.
Results
The framework provides consistency and asymptotic normality results under suitable regularity conditions and quantifies static and dynamic effects in corporate e-mail recipient selection.
Takeaways & Limitations
The proportional intensity model is presented as useful for repeated directed interactions because it directly models data and supports precise time-dependent quantification of network effects.
Takeaways & Limitations
A parsimonious dyadic-only model has χ2 = 21094, but the analysis retains triadic effects to seek lower bias and estimates for all network effects.
Abstract
from arXiv · showhide
Network data often take the form of repeated interactions between senders and receivers tabulated over time. A primary question to ask of such data is which traits and behaviors are predictive of interaction. To answer this question, a model is introduced for treating directed interactions as a multivariate point process: a Cox multiplicative intensity model using covariates that depend on the history of the process. Consistency and asymptotic normality are proved for the resulting partial-likelihood-based estimators under suitable regularity conditions, and an efficient fitting procedure is described. Multicast interactions--those involving a single sender but multiple receivers--are treated explicitly. The resulting inferential framework is then employed to model message sending behavior in a corporate e-mail network. The analysis gives a precise quantification of which static shared traits and dynamic network effects are predictive of message recipient selection.
1. Introduction
The paper frames repeated directed interactions as data for studying which sender, receiver, and network characteristics predict interaction. It introduces a framework addressing homophily, dynamic network effects, and multiple-receiver interactions.
- Research questions: Directed interaction data record triples indicating when sender i interacts with receiver j.The central modeling goal is to identify characteristics and behaviors predictive of interaction.
- Research questions: The paper asks whether shared attributes predict heightened interaction among similar individuals.This is the homophily question.
- Research questions: It asks whether past interaction behaviors predict future directed interactions, including possible transitive patterns i →h and h →j preceding i →j.These questions concern dynamic network effects.
- Research questions: It asks how to model multicast interactions in which one sender communicates with multiple receivers.The paper treats this as a distinct multiplicity problem rather than automatically reducing it to pairwise interactions.
- Approach: The framework combines a Cox proportional intensity model, partial-likelihood inference, asymptotic theory, multicast extensions, and an empirical corporate e-mail analysis.The stated program includes computationally efficient procedures and analyses of homophily and network effects.
2. A point process model and partial likelihood inference
The paper models directed interactions as a multivariate point process whose intensity depends on sender baselines and predictable pairwise covariates, including history-dependent network effects. Partial likelihood removes the baseline nuisance component, while specialized optimization addresses time-varying covariates and large datasets.
- Point process model: Each sender–receiver interaction is represented through a counting process with a stochastic intensity for events over time.The intensity is interpreted as the probability of an interaction in a short interval, up to first order.
- Point process model: The Cox proportional intensity model modulates sender i’s baseline rate for receiver j by exp{β^T x_t(i,j)}.The covariate vector is predictable, locally bounded, and may depend on interaction history.
- Covariates: Predictable covariates support static group indicators as well as history-dependent measures of reciprocation, transitivity, and other network effects.They cannot depend on the future or the immediate present.
- Partial likelihood: Partial likelihood estimates β while treating the sender baseline rate as a nuisance parameter.It is formed from conditional likelihoods of observed receivers given interaction times and senders.
- Computation: The concave log partial likelihood supports Newton or gradient optimization, but time-varying covariates prevent standard sufficient-statistic routines from sufficing for most large datasets.The paper describes a customized method exploiting sparsity in x_t(i,j).
3. Consistency of maximum partial likelihood inference
Under regularity conditions, the paper establishes consistency and asymptotic Gaussian behavior for maximum partial-likelihood inference as the number of interactions grows. The results address the observation-time asymptotic regime most relevant to interaction data.
- Inference: The MPLE and the inverse Hessian of the log partial likelihood provide natural estimates of β0 and its covariance matrix.The paper states conditions under which these estimators are consistent.
- Inference: The relevant sampling regime increases observation time because interaction studies generally cannot control the sender and receiver sets.Existing consistency arguments do not fully cover the continuous, time-varying setting considered here.
- Main results: Assumptions A1–A4 imply that the MPLE is consistent and asymptotically Gaussian.The assumptions concern covariate integrability, covariance behavior, finite arrival times, and equicontinuity.
- Main results: Theorem 3.1 establishes weak Gaussian-process convergence for the scaled score and supports consistency under its stated conditions.The result uses assumptions on covariance behavior, with additional conditions for the estimator conclusions.
- Main results: Theorem 3.2 states that √n(β̂_n − β0) converges weakly to a mean-zero Gaussian variable with a specified covariance under the required conditions.The covariance regularity includes a locally Lipschitz function with smallest eigenvalue bounded away from zero.
4. Multicast interactions
The paper extends its directed-interaction point-process framework to multicast events with multiple receivers, while analyzing duplication-based approximate partial likelihood. It establishes consistency and asymptotic results, quantifies approximation bias, and studies when the approximation is accurate.
- Model extension: Multicast interactions involve one sender and possibly multiple receivers, which pairwise-only models do not describe.The paper extends the model and asymptotic theory beyond strictly pairwise directed interactions.
- Approximate inference: Duplication converts one multicast event into multiple pairwise interactions, producing an approximate partial likelihood related to the multicast model.The paper studies the approximation because exact multicast partial likelihood involves an intractable combinatorial sum.
- Model extension: The multicast model uses a baseline L-receiver intensity for each sender and models observed sender–receiver-set interactions directly.The receiver set J is represented as a set with cardinality |J|, rather than as separate pairwise events.
- Approximate inference: The approximation error between the exact and approximate estimators is O_P(G_n/n) under bounded covariates and receiver-set sizes.The receiver-set growth sequence G_n controls the error introduced by replacing the exact partial likelihood with its approximation.
- Bias correction: If receiver-set growth makes G_n smaller than O_P(√n), the approximate MPLE is √n-consistent and asymptotically Gaussian with the same covariance but possibly a different mean.A parametric bootstrap is used to estimate the resulting bias, while the observed Hessian estimates the limiting covariance.
- Simulation: In simulations, when |J| = √n, the approximate estimator’s squared error is roughly O_P(n^-1), matching the theorem’s prediction.The study used 100 replicates across receiver counts from 32 to 1000 and sample sizes from 32 to 100,000.
5. Fitting the model to a corporate e-mail network
The Enron analysis uses static actor traits and history-dependent dyadic and triadic covariates in a multicast proportional intensity model to assess predictors of message recipient selection. Bootstrap correction and goodness-of-fit analyses show substantial predictive value for network effects, especially prior sending, while addressing bias from multicast duplication.
- 5.1. Data and methods: The study analyzes 21,635 messages sent among 156 Enron employees between November 13, 1998 and June 21, 2002.The data come from a publicly available Enron e-mail corpus compiled by Zhou et al. (2007).
- 5.2. Covariates: The model combines actor traits with six history-dependent network indicators representing dyadic and triadic interaction effects.Static covariates encode gender, department, seniority, and their interactions; the network indicators include send, receive, 2-send, 2-receive, sibling, and cosibling.
- 5.3. Bootstrap bias correction: 500 bootstrap replicates show that treating multicast messages as separate single-recipient messages produces bias on the order of the standard error for most coefficients.Dyadic-effect estimates are particularly vulnerable because their covariates are sparse, and the bootstrap is used to correct the resulting bias.
- 5.3. Bootstrap bias correction: The simulated standard errors are very close to the theoretical standard errors, supporting the reasonableness of the asymptotic approximations in this application.This remains the case even though the covariate norm may be unbounded, contrary to an assumption of Theorem 3.1.
- 5.4. Goodness of fit: Static effects account for 15% of null deviance, network effects for 37%, and the full model for 52%.The Send terms produce the largest decrease in residual deviance, accounting for 33% of null deviance with 8 degrees of freedom.
- 5.4. Goodness of fit: The full model fits better than the static model: its Pearson-residual sum of squares is 17,281 versus 596,253, over 34 times lower.More than 95% of full-model absolute residuals are below 1.21, compared with a 3.5 quantile for the static model.
- 5.4. Goodness of fit: A dyadic-only model has a Pearson-residual sum of squares of 21,094, but the analysis retains triadic effects to minimize bias and estimate all network effects.The authors therefore favor the fuller model despite the availability of a more parsimonious alternative.
6. Evaluating the strength of homophily and network effects
The Enron analysis finds that homophily and dynamic network effects both predict message-recipient selection, with dyadic effects stronger and longer-lived than triadic effects.
- Assessing evidence for homophily in the Enron data: Homophily is evident for almost all main effects, although Gender, Seniority, and some Department combinations are not significant.The estimated coefficients for L(j), T(j), and J(j) are negative, while female-gender similarity has a relative rate of approximately 1.2.
- Assessing evidence for homophily in the Enron data: Legal senders exhibit negative homophily, with relative rates of 0.76 for Legal recipients, 0.92 for Trading recipients, and 1 for Other recipients.These values compare recipient departments while holding other recipient characteristics fixed.
- Evaluating the importance of network effects: All network-indicator coefficients are positive, and 1{send} is more than three times larger than the other coefficients.The next tier comprises 1{receive}, 1{sibling}, and 1{2-send}, with estimated coefficients ranging from 0.67 to 1.06; 1{2-receive} and 1{cosibling} are not significant.
- Evaluating the importance of network effects: Dyadic effects persist for over three weeks, with coefficients decaying roughly exponentially as time since interaction increases.The corresponding relative sending rate therefore exhibits super-exponential decay.
- Evaluating the importance of network effects: Reciprocation dominates for up to two hours after receipt, after which prior sending to another comparable individual becomes more influential.The comparison is between receive(k) and send(k) coefficients across recency windows.
- Evaluating the importance of network effects: Triadic effects are weaker and shorter-lived than dyadic effects: about 86% of their estimated coefficients lie within 3 standard errors of zero.Most significantly nonzero triadic coefficients lie between −0.05 and +0.05, while sibling effects show the clearest pattern.
- Evaluating the importance of network effects: Sibling effects switch direction by recency: they increase cross-sending when the shared sender acted within 30 minutes or 2–8 hours, but decrease it at 30 minutes–2 hours.This pattern is reported for the sibling covariate.
- Evaluating the importance of network effects: Most triadic effects are insignificant, consistent with colleagues generally lacking direct knowledge of one another’s e-mail activity.The analysis attributes remaining predictive power to correlation with exogenous factors and notes that triadic effects have small time horizons.
7. Conclusion
The conclusion presents the proportional-intensity framework as a direct, flexible approach for directed interaction data and summarizes theoretical extensions for time-asymptotic inference and multicast interactions.
- Conclusion: The approach models directed interaction data directly and is expected to generalize beyond the Enron e-mail analysis.The authors contrast it with contingency-table, actor-oriented, and exponential-random-graph alternatives.
- Conclusion: Partial likelihood estimates the coefficient vector while treating each sender-specific baseline intensity as a nuisance parameter.Baseline intensities would need separate estimation for prediction, for example with a Nelson–Aalen estimator.
- Conclusion: The theory is asymptotic in time rather than population size, and multicast duplication can bias estimates, though correction is possible in certain regimes.These are the two stated extensions of Cox partial-likelihood theory for interaction data.
- Conclusion: Time-varying covariates make the proportional-intensity model useful for identifying traits and behaviors predictive of repeated directed interactions.The model is characterized as simple, flexible, and well established.
A. Implementation
The implementation computes partial-likelihood derivatives efficiently by factorizing sender-specific terms and exploiting sparsity and repeated sender-group structure.
- Implementation: Sender-specific partial-likelihood terms can be computed in parallel and summed to obtain the full log partial likelihood and its derivatives.This factorization follows from the sender-wise structure of the partial likelihood.
- Implementation: Naive derivative computation can require O(n J p2) time, motivating sparse updates for large networks.The expensive terms arise from iterating over messages and receivers.
- Implementation: When senders share a small number of covariate-defined groups, initialization complexity drops from O(I J p2) to O(¯I J p2).Here ¯I is the number of distinct sender-group covariate patterns.
- Implementation: Dynamic covariates are updated incrementally because their changes are usually sparse across sender–receiver pairs.The implementation restricts attention to receiver pairs whose dynamic covariates can change.
- Implementation: Sparse gradient updates require total time O(n ¯J d p + I p).The bound follows from computing changing terms separately for each sender and then aggregating them.
- Implementation: Sparse Hessian updates reduce the changing-term accumulation cost to O(n ¯J d p2 + I p2).Only terms that change with time are accumulated.
- Implementation: The total cost per Newton step is O(¯I J p2 + n ¯J d p2 + I p2 + p3), which is nearly linear in I, J, and n.The authors conclude that the resulting algorithm scales naturally to large datasets.
B.1. Proof of Theorem 3.1
The proof establishes asymptotic behavior for the partial-likelihood score by representing it with martingales, rescaling time, and applying a martingale central limit theorem.
- Proof of Theorem 3.1: The score function at the true parameter has a representation in terms of compensated interaction processes.The proof introduces local martingales formed by subtracting compensators from the counting processes.
- Proof of Theorem 3.1: Uniform boundedness of the covariates makes the score locally square integrable and permits control of its predictable variation.The proof uses Assumption A1 and localizing stopping times.
- Proof of Theorem 3.1: After time rescaling, the discretized score process is analyzed on the unit interval using its values at jump times.The rescaled process is right-continuous with left limits.
- Proof of Theorem 3.1: The rescaled score satisfies a Lindeberg condition and converges in distribution to a Gaussian process with covariance function Σα(β0).This follows from Lemma B.2 and Rebolledo’s martingale central limit theorem.
- Proof of Theorem 3.1: The proof shows the relevant remainder terms converge to zero in probability by combining consistency, equicontinuity, Lenglart’s inequality, and Assumption A2.The convergence argument controls all three terms in the decomposition.
B.2. Supporting lemmas for Theorem 3.1
The supporting lemmas establish the martingale and moment conditions needed for asymptotic normality of the partial-likelihood estimator under the stated assumptions.
- Together, these lemmas verify the probabilistic conditions used to derive the estimator’s asymptotic distribution.The supporting arguments control conditional expectations, square integrability, and large-jump contributions.
- The conditional expectation property follows from the martingale structure under assumption A1, with optional sampling justified by uniform integrability.Square integrability is also established for the stopped process.
- The Lindeberg condition required by Rebolledo’s central limit theorem is satisfied under assumption A1.The proof bounds the relevant integral using the covariate norm and the compensator, then shows the bound converges to zero in probability.
- Under assumptions A1 and A3, the relevant normalized remainder term converges to zero.Lenglart’s inequality and the compensator moment condition provide the required bound.
B.3. Proof of Theorem 3.2
The proof uses Newton’s method and Kantorovich bounds to show that the partial-likelihood estimator is well defined, consistent, and asymptotically normal under the theorem’s regularity conditions.
- Kantorovich’s theorem guarantees well-defined Newton iterates that remain in a controlled neighborhood and converge to a zero of the nonlinear system.The theorem requires an invertible Jacobian and a Lipschitz condition on its variation.
- The negative Hessian is positive semi-definite, so the log partial likelihood is concave and a local maximum is the unique global maximum near β0.The smallest-eigenvalue condition supplies the required local strictness for sufficiently large n.
- The Newton construction yields β̂n → β0, while √n(β̂n − β0) and √nZn converge weakly to the same limit.The proof uses the first Newton iterate and bounds the estimator’s distance from it.
- The multicast approximation replaces sampling without replacement by independent weighted sampling and controls the resulting gradient and Hessian errors.The weights are exponential in covariates, and the approximation error is bounded through a coupling of the two sampling laws.
C.2. Supporting lemmas for Theorem 4.1
The supporting results for the multicast approximation bound the differences between with- and without-replacement quantities and establish convergence of the approximate Newton procedure.
- The corresponding Hessian difference is bounded using covariance matrices under the two receiver-sampling laws.The proof applies the same comparison strategy used for the gradient bound.
- Newton’s method applied to the approximate partial likelihood converges to β̃n after sufficiently many iterations.The procedure initializes at β̂n and uses Kantorovich’s theorem to bound the distance to β̃n.
- The approximation error is controlled at order OP(Gn/n) under the theorem’s assumptions and bounded inverse-Hessian condition.This rate supports the required closeness between the exact and approximate estimators.
1. A comparative analysis based on contingency tables
The contingency-table analysis estimates homophily from message counts but ignores dependence among interactions and can marginalize over covariates, producing potentially misleading predictive effects.
- Junior-Junior homophily corresponds to a positive βJ + βJJ, whereas Senior-Senior homophily corresponds to a negative βJ.The multinomial-logit formulation conditions only on the sender and ignores network effects.
- −1.4 and 1.6 are the estimated coefficients β̂J and β̂JJ, each with Wald standard errors of about 0.02.These coefficients are obtained from the 2 × 2 table of messages exchanged between the two seniority groups.
- Marginalizing over Gender and Department can introduce Simpson’s paradox, so sender and receiver covariates should be included in the analysis.The resulting coefficient estimates are displayed in Fig. 1.
- The contingency-table approach assumes messages are conditionally independent and identically distributed given the sender, but reciprocation and other network effects violate that assumption.Ignoring these dependencies can exaggerate the apparent homophily effect.
2. Comparative analyses using actor-oriented and exponential random graph models
The paper compares its proportional-intensity framework with actor-oriented and exponential random graph models fitted to reduced versions of the Enron e-mail data. Both alternatives recover several network effects qualitatively, but data reduction prevents separating some dynamic effects and limits direct comparison of group-level effects.
- Actor-oriented model: The actor-oriented comparison models network snapshots as a first-order Markov chain in which actors change ties according to stochastic utility.It is designed for persistent ties rather than instantaneous events.
- Actor-oriented model: The actor-oriented model was fitted to three regularly binned Enron e-mail snapshots using covariates analogous to the proportional-intensity model.The comparison includes outdegree/density, group-level edge effects, reciprocity, cyclic and transitive structures.
- Actor-oriented limitations: Binning interaction counts into snapshots makes the 2-send, sibling, and cosibling effects inseparable and restricts quantification of dynamic-effect time decay.These limitations follow from the snapshot representation and the actor-oriented model's first-order Markov structure.
- Actor-oriented results: The actor-oriented estimates qualitatively agree with the main model: outdegree/density endowment is negative, while reciprocity and transitive triplets are positive.The dyadic coefficients are larger than the triadic coefficients, while 2-receive has a negligible effect in the proportional-intensity model.
- Exponential random graph model: The exponential random graph comparison reduces the data to one directed graph by defining an edge when a sender-receiver pair has at least 10 sent messages.The model includes sender, group-level edge, mutuality, cyclic-triple, and transitive-triple covariates; second-order group-level interactions could not be fitted computationally.
- Exponential random graph results: The exponential random graph estimates show positive mutuality and transitive-triple effects but a negligible cyclic-triple effect.Mutuality and transitive triples agree qualitatively with the main model, whereas the cyclic-triple result differs from the actor-oriented comparison.