Source-linked AI summary
Portfolio Risk Bounds without Cross-Asset Return Covariances: Distributional Fields from Language-Model Representations
Marcus Gawronsky, Chun-Sung Huang
TL;DR
Short, high-dimensional panels make cross-asset covariance estimation difficult, so the paper asks what firm-level information distributions can certify about portfolio risk beforehand. It uses maintained links from information to exposures and returns to derive coherent Wasserstein risk bounds and minimize a tractable certificate; in a 52-firm panel, the Qwen3-Embedding-8B allocation ranks between the 0.69th and 1.33rd in-sample variance percentiles, versus 21.1st–28.6th for equal risk weights.
Problem
Short, high-dimensional return panels make reliable cross-asset covariance estimates difficult even though covariance is central to portfolio risk assessment.
Method
The paper compares firm-level article-embedding distributions with Wasserstein distance and, under maintained transmission and return restrictions, minimizes a coherent or weighted-pairwise certified variance bound.
Results
0.69%–1.33% of feasible portfolios had variance no greater than the Qwen3-Embedding-8B news-only allocation, versus 21.1%–28.6% for equal risk weights.
Takeaways & Limitations
The framework converts distribution-valued firm information into an implementable allocation rule without cross-asset return covariances; with zero slack, normalized allocation depends only on observed W2 geometry.
Takeaways & Limitations
The certificate is conditional on maintained transmission restrictions that the text data do not identify, and representation choices can alter measured distances and portfolio weights.
Abstract
from arXiv · showhide
Portfolio risk assessment ordinarily relies on reliable estimates of cross-asset return covariances, which are difficult to obtain in short, high-dimensional panels. We show that firm-level distribution-valued characteristics can instead provide one-sided certificates of portfolio risk. Under maintained links from characteristics to systematic exposures and from exposures to returns, multi-firm Wasserstein-2 dispersion yields a sharp upper bound on systematic portfolio variance and a corresponding bound for standardized returns. A weighted pairwise relaxation produces an objective that is convex under a checkable condition and requires marginal volatility scales but no cross-asset return covariances. With zero firm-specific slack, the common-map scale changes the certified variance reduction but not the normalized allocation, which depends only on observed information geometry. In a 52-firm panel from 2018-2022, an allocation constructed from Qwen3-Embedding-8B news representations lies between the 0.69th and 1.33rd in-sample variance percentiles across four prespecified capped portfolio populations; equal risk weighting lies between the 21.1st and 28.6th percentiles. The lower in-sample variance ranking relative to equal risk also appears across the reported frozen language-model representations. The framework therefore distribution-valued firm information into a coherent risk bound and an implementable allocation rule constructed without cross-asset return covariances.
1 Introduction
The paper replaces direct covariance estimation with distribution-valued firm information that certifies portfolio-risk bounds under maintained transmission restrictions. It derives a coherent multi-firm certificate, a tractable allocation rule, and descriptive in-sample evidence from news representations.
- Motivation: Short, high-dimensional return panels make covariance estimation unstable while minimum-variance allocation remains sensitive to noisy weak directions.An unrestricted covariance matrix has n(n + 1)/2 entries, whereas demeaned returns of length T have rank at most min(T −1, n).
- Approach: Firm article-embedding distributions are compared with quadratic optimal transport to obtain observable information geometry.Sufficient separation, under maintained information-to-exposure restrictions, implies that latent systematic risks cannot be perfectly aligned.
- Certificate: The observed-information certificate deducts C(q) from the perfect-positive-dependence variance benchmark without estimating a covariance matrix.Larger observed distances tighten the admissible upper bound, while distortion and firm-specific slack weaken it.
- Decision rule: Minimizing the upper bound yields an information-certified long-only portfolio without an expected-return input.A directly checkable condition on the observed distance matrix makes the standardized objective convex.
- Evidence: 0.69%–1.33% of feasible portfolios had variance no greater than the news-only allocation, versus 21.1%–28.6% for equal risk weights.These rankings are descriptive and in-sample, not forecasts of out-of-sample performance.
- Contributions: The paper contributes a coherent multi-firm variance bound, a weighted pairwise decision relaxation, and a 52-firm in-sample evaluation.With zero firm-specific slack, the common carrier scale changes certified variance reduction but not normalized allocation.
2 Related literature
The paper connects portfolio theory, factor-risk models, distribution-valued characteristics, and textual finance while identifying a missing portfolio-level bridge from information geometry to coherent risk bounds.
- Portfolio risk: Regularized covariance methods impose structure but still estimate joint return risk before solving the portfolio problem.This motivates a route that uses observable firm information instead of cross-asset return covariances.
- Factor and characteristic models: Factor and characteristic-based models organize the latent exposures that covariance summarizes, but do not by themselves provide the paper’s portfolio certificate.The paper instead links observable distributional characteristics to latent systematic exposure laws.
- Distributional characteristics: Distribution-valued characteristics use quadratic transport to compare laws and support pairwise covariance envelopes under maintained restrictions.Retaining empirical embedding distributions preserves within-firm heterogeneity that pooled vectors would suppress.
- Positioning: Prior barycentric work organizes cross-sectional exposure adjustment, whereas the present contribution aggregates distributional separation into portfolio risk.The earlier pairwise and cross-sectional results do not resolve portfolio aggregation.
- Portfolio aggregation: A single coherent joint exposure law is required because pairwise optimal couplings may be mutually incompatible.The sharp multi-firm object preserves this compatibility, while weighted pairwise distances provide a computational relaxation.
3 Model and information-certified variance bound
The model maps observed firm information laws to latent exposure laws through maintained restrictions, then uses coherent Wasserstein dispersion to bound systematic portfolio variance and standardized returns.
- Model objects: Each firm has an observed characteristic law and a distinct latent systematic-exposure law connected by a maintained information-to-exposure transmission.The observed law is formed from article embeddings, while the latent exposure law lives in factor-risk coordinates.
- Transmission restrictions: An antilipschitz common carrier and firm-specific slack translate observed information distance into a lower bound on exposure separation.The restrictions prevent distinct information states from collapsing into identical exposures while permitting bounded mismatch.
- Coherence: One coherent joint law for all exposure marginals ensures that pairwise relations jointly define a valid portfolio risk configuration.The same joint law induces a positive-semidefinite covariance structure and rules out incompatible pairwise couplings.
- Variance bound: Theorem 1 bounds systematic portfolio variance by weighted marginal second moments minus an observable characteristic-distance credit.The credit C(q) grows with observed separation and is weakened by distortion or slack.
- Relaxation and sharpening: Weighted pairwise transport is conservative because pairwise optimal couplings may not coexist and pairwise slack deductions are separately truncated.The sharp multi-firm dispersion retains a common coupling and applies one aggregate slack allowance.
- Sharp certificate: The two-asset case exactly recovers the prior covariance-envelope correction, while the multi-firm construction extends it under coherent joint exposure laws.The sharp certificate is exact when an attaining coherent joint law exists and does not identify a return covariance matrix from text.
- Return bridge: The return bridge converts the latent systematic-risk bound into a standardized total-return statement but remains a maintained restriction rather than a consequence of information geometry.The decision certificate also includes a residual-covariance budget for sensitivity analysis.
4 Portfolio choice under the certificate
The paper minimizes a certified variance cap over feasible long-only capital weights rather than estimated variance. Existence and convexity follow under geometric conditions, while zero slack makes normalized allocations invariant to carrier scale.
- Decision rule: The allocation rule chooses feasible capital weights with the smallest certified upper bound at the benchmark δ = 0.The objective is a minimum-risk problem and excludes expected returns.
- Allocation trade-off: The objective trades off the perfect-dependence marginal-volatility benchmark against an information-based deduction for unaligned latent risks.Inverse volatility uses the first channel but not the information-geometry deduction.
- Existence: Every nonempty compact feasible long-only set admits a minimizer because the certified-variance objective is continuous.The pointwise variance certificate holds at that minimizer.
- Convexity: If the pairwise floor matrix is conditionally negative definite, q ↦ 1 − C(q) is convex on the simplex.The condition is checkable from the observed distance matrix, and minimizing the standardized certificate over a convex feasible set becomes a convex program.
- Scale invariance: With zero firm-specific slack, the carrier scale changes the certified variance deduction but not the normalized risk allocation.The normalized allocation depends only on observed W2 geometry, whereas capital implementation still requires marginal volatility scales.
5 Empirical design and data
The empirical design constructs a zero-slack news-only allocation from observed Wasserstein-2 information geometry, then evaluates it using standardized returns and full-sample covariance. It compares the allocation across capped reference populations, benchmarks, data screens, and representation sensitivities.
- Allocation construction: The zero-slack news-only allocation minimizes the certified standardized variance bound using observed information geometry before return covariance enters evaluation.With unit marginal scales, q = x; the common-map scale changes the certificate size but not the normalized allocation.
- Reference populations: 0.69%–1.33% of feasible portfolios had variance no greater than the news-only allocation across four capped reference populations.The comparison uses Monte Carlo estimates of descriptive lower-tail percentiles, not classical p-values.
- Reference populations: The four reference populations use clipped and renormalized symmetric Dirichlet draws under 12.5% or 15% single-name caps.The α values are 1, 0.8, and 0.5; α = 0.8 is calibrated to match the candidate’s effective number of names.
- Benchmarks: Equal normalized risk weights correspond to inverse-volatility capital weights and use marginal volatility without cross-asset covariance.The ex post sample GMV instead uses the full return covariance as a covariance-informed in-sample benchmark.
- Data: The return panel contains 52 coverage-screened firms with 1207 common observations from 2018-03-19 through 2022-12-30.WBA is dropped during return alignment, and missing dates are removed rather than imputed.
- Representation and evaluation: The canonical representation independently encodes each article with frozen 4,096-coordinate Qwen3-Embedding-8B embeddings, with normalized rows.The variance evaluation uses the same 2018–2022 period, making the exercise an in-sample descriptive ranking.
- Representation sensitivity: Representation sensitivity varies output width, model family, and encoder vintage while holding the main evaluation design fixed where possible.The BGE comparison changes model family, pretraining, and effective input length jointly; EttaX vintage comparisons are descriptive sensitivities.
6 Results
The news-only allocation ranks unusually low in in-sample variance across four prespecified feasible populations, while remaining above the covariance-informed sample GMV. Its allocation is diversified across many names, relies substantially on cross-sector certificate credit, changes moderately across expanding information cutoffs, and varies descriptively across representations.
- Main in-sample variance ranking: 0.69%–1.33% of feasible portfolios have variance no greater than the Qwen3-Embedding-8B news-only allocation across four prespecified capped populations.The variance percentile is the share of feasible reference portfolios with variance no greater than the candidate’s; lower is better.
- Main in-sample variance ranking: The low ranking persists under matched and concentrated populations, including a matched mean effective N of 24.23 versus 23.65 for the candidate.Only 0.89% of matched-law draws and 1.33% under the more concentrated law attain equally low or lower variance.
- Main in-sample variance ranking: Equal risk weights rank between 21.1% and 28.6% across the same four populations, well above the news-only allocation under every reference law.The benchmark is inverse-volatility capital weighting expressed in normalized risk-weight coordinates.
- Portfolio characteristics and conventional benchmarks: The news-only standardized variance is 8.3% below equal risk but 35.6% above the ex post long-only sample GMV.Returns evaluate the allocation after construction; they do not determine its weights.
- Allocation anatomy: The allocation assigns positive weight to 38 of 52 firms, with largest holding 12.1% and effective number of names 23.65.Pair credits depend on both counterpart weight and squared transport separation, so the largest position need not have the largest pair credit.
- Allocation anatomy: Cross-sector pairs account for 87.1% of certificate credit and within-sector pairs 12.9%, while sector panels remain diagnostic rather than causal explanations.The allocation’s pair terms identify which weighted separations the optimizer uses.
- Evolution across expanding information cutoffs: Across expanding 2018–2022 article cutoffs, effective N stays between 22.95 and 24.56, largest weight ranges from 8.3% to 12.1%, and turnover from 13.2% to 26.8%.The largest reallocation occurs from 2018 to 2019; the design cannot isolate a COVID effect because the entire accumulated article distribution changes at once.
- Representation sensitivity: Representation sensitivity is descriptive: moderate compression is close to the primary representation, extreme compression changes the geometry materially, and the ladder does not support monotone model-size ordering.The matched EttaX vintage contrasts also do not establish causal vintage effects or representation invariance.
7 Discussion and Limitations
The paper presents a bounded positive implication: information-derived separation yields a one-sided portfolio-variance restriction and a zero-slack allocation rule without cross-asset return covariance. Its empirical interpretation is limited by maintained transmission restrictions, representation and comparison-set choices, implementation gaps, and in-sample survivor-panel evaluation.
- Contribution: The certificate maps observed Wasserstein separation into a one-sided portfolio-variance restriction, while the zero-slack rule minimizes the certified upper bound without estimating cross-asset return covariance.The resulting allocation lies near the first percentile across four prespecified reference populations, whereas equal risk weights rank materially higher.
- Identification: Identification depends on an antilipschitz carrier, bounded firm-specific slack, one coherent joint exposure law, and a return bridge that text data do not identify.The certificate is therefore valid under declared restrictions rather than being an unconditional implication of text geometry.
- Representation dependence: Representation choices can materially alter measured distances and allocations, including through non-monotone width compression and unmatched vintage outcomes.The primary representation was prespecified, but the evidence does not establish representation invariance.
- Comparison-set design: Percentile rankings depend on four prespecified Dirichlet-capped reference populations, whose concentration parameter and cap determine what “low” means.Matching effective portfolio size addresses the immediate concern that concentrated candidates are compared only with diffuse portfolios.
- Portfolio implementation: Capital-weight implementation requires marginal volatility scales, and the practical gap between standardized and capital-weight allocations remains unquantified.Cross-asset covariance is not required, but the standardized exercise does not quantify this implementation gap.
- External validity: The evidence comes from a coverage-screened 52-firm survivor panel using the same 2018–2022 period for information geometry and return evaluation, so it is descriptive and in-sample.Prospective claims require point-in-time construction and genuinely out-of-sample return evaluation.
8 Conclusion
The paper addresses the difficulty of estimating return covariance matrices by using firm-level information distributions to construct a coherent portfolio-level risk certificate and allocation rule. In the 52-firm panel, the Qwen3-Embedding-8B allocation ranked unusually low in-sample, while the result remained descriptive and conditional on maintained restrictions.
- Problem: The paper asks what observed differences between firms’ information distributions can certify about portfolio risk before cross-asset covariance is estimated.The resulting certificate is a one-sided restriction on admissible portfolio variance, not a point estimate.
- Method: Multi-firm transport dispersion supplies the sharp certificate, while a weighted pairwise certificate provides the computationally simpler decision rule used empirically.At zero firm-specific slack, the common carrier scale changes certified variance reduction but not the normalized allocation based on observed W2 geometry.
- Takeaway: The framework turns distribution-valued firm information into a coherent portfolio-level risk bound and an implementable allocation rule without expected returns or cross-asset return covariance.Capital implementation still requires marginal volatility scales.
- Results: 0.69%–1.33%: the Qwen3-Embedding-8B news-only allocation’s variance percentiles across four prespecified capped reference populations, with equal risk weights ranking higher in every comparison.Its standardized variance was 8.3% below the equal-risk benchmark but 35.6% above the ex post long-only sample GMV.
- Limitations: The empirical evidence is descriptive and in-sample, and the certificate remains conditional on maintained transmission restrictions that the text data do not identify.The paper does not establish prospective performance or covariance-free replication of the covariance-informed optimum.
A Formal verification
The formal record machine-checks the paper’s identities, bounds, convexity, existence, and pairwise-to-multifirm relationships under explicit assumptions. It establishes algebraic guarantees, while leaving empirical proxies, economic restrictions, and realized-return claims outside formal validation.
- Scope: The formal record checks the displayed identities and bounds under their stated assumptions, without validating empirical or economic components.The unchecked components include article-embedding proxies, carrier-and-slack restrictions, distance estimation, and chronological evaluation.
- Main guarantees: Theorem 1 composes the carrier-and-slack W2 floor with the coherent joint-law variance cap, conditional on transmission, slack, joint-law, and integrability hypotheses.It establishes neither calibration of those inputs nor a realized total-return or ex-ante empirical bound.
- Risk bound: Equation (18) bounds total standardized variance by one minus certificate credit plus δ, treating δ as an assumed residual budget without deriving residual orthogonality.The wrapper applies when systematic standardized variance is bounded by one minus the certificate credit and aggregate residual contribution is at most δ.
- Optimization: Theorem 4 proves existence of a certified-variance minimizer on every nonempty compact feasible subset and preserves the pointwise variance certificate.The result does not establish uniqueness or unconditional convexity for arbitrary distance matrices.
- Main guarantees: Theorem 3 gives an attainment-free, sharper aggregate certificate by subtracting one weighted root-mean-square slack radius rather than two radii per pair.It composes the aggregate carrier floor with a multi-marginal variance envelope.
- Main guarantees: Proposition 2 is machine-checked as an exact greatest-element statement, making the multi-marginal bound unimprovable within the coherent-joint-law class.Attainment is supplied as a premise, and the result does not identify which joint law the data realize.
- Optimization: Theorem 5 establishes convexity only under the stated curvature property of the squared distance matrix, implying neither minimizer uniqueness nor superior realized variance.Whether an empirical matrix satisfies the curvature condition remains a separate numerical question.
- Connections: Proposition 1 identifies the pairwise covariance-envelope correction as the two-asset case of the multifirm construction by exact machine-checked equality.This establishes that the two formulations describe one object rather than merely compatible objects.
B Barycentre representation of dispersion
The barycentre representation equates a weighted multimarginal coupling problem with an infimum over laws of weighted Wasserstein-2 distances. The equivalence is proved in both directions without requiring either infimum to be attained.
- Representation: For laws P1, . . . , Pn on H with finite second moments, Proposition 3 states the barycentre representation.The representation rewrites the relevant dispersion quantity through a law Q on H.
- Forward direction: Given a multimarginal coupling, its random weighted barycentre induces a law Q, yielding the direction from the joint-coupling infimum to the barycentre infimum.The pointwise weighted-variance identity supplies the bridge between the two representations.
- Reverse direction: Conversely, near-optimal couplings of each Pi with a common Q can be glued through their conditional laws to construct a multimarginal coupling.Conditioning on the common Q-distributed variable applies the pointwise weighted-variance minimum.
- Conclusion: The two infimum values agree after taking expectations and minimizing over couplings and Q, without assuming an optimal barycentre or joint plan exists.Only equality of infimum values is required for the paper’s argument.
C Data and W2 artifact construction
The empirical construction aligns article-embedding distributions with return data for 52 firms, using a defined W2 artifact and balanced annual cloud sizes. The canonical roster records the priced firms and their sectors.
- Data: 53 firms begin the empirical universe, but return alignment leaves 52 firms with 1207 common daily observations.WBA is the only embedding-covered name without a return series.
- Embeddings: Each article has metadata and a 4,096-dimensional Qwen3-Embedding-8B vector, with rows normalized to unit Euclidean norm and Euclidean chord distance as ground cost.The retained representation is article-level rather than a single pooled firm vector.
- W2 artifact: The persisted artifact is rooted W2 using balanced quadratic transport, rooted normalization, and ground exponent one; the certificate boundary squares it once.Regression tests enforce this transition between the stored artifact and the certificate.
- Sampling: Balanced assignment uses equal empirical cloud sizes after date restriction and URL-hash ordering, with common sizes of 6, 19, 32, 64, and 128 from 2018 through 2022.This rule retains all 52 priced firms rather than deleting firms with thin coverage.
- Roster: Table 5 lists the canonical priced W2 roster used by the certificate and diagnostics.The table contains ticker symbol, firm name, and sector label, sorted by sector then symbol.
D Supplementary empirical results
The supplementary comparison reports standardized in-sample variance relative to the full-sample long-only GMV benchmark. The news-only allocation is presented as a descriptive zero-slack maximizer of the news certificate.
- Figure 4: Figure 4 compares standardized in-sample variance against the full-sample long-only GMV benchmark.The figure provides the visual comparison corresponding to the conventional benchmark values reported in Table 2.
- Figure 4: The news-only allocation is a descriptive zero-slack maximizer of the news certificate.This characterizes the allocation’s construction in the supplementary comparison.
D.1 Complete embedding-representation sensitivity
The sensitivity analysis compares news-only allocations across seven base representations and three EttaX encoder vintages on a common 52-firm panel, using fixed portfolio inputs and relative-GMV measures. Moderate Qwen3 truncation does not reject a zero relative-variance difference, while selected fixed-width and vintage comparisons differ or fail equivalence testing.
- Design: The sensitivity tables hold priced firms, standardized-return covariance, reference laws, caps, and bootstrap scheme fixed while changing only the representation.This isolates representation changes in the reported comparisons.
- Metrics: Relative-GMV index contrasts are candidate-minus-reference differences, so negative estimates favor the candidate allocation.The index sets long-only sample GMV to 100; values above 100 indicate percentage excess variance.
- Interpretation: A dash indicates that the news-only allocation violates the fixed law’s cap rather than that the cap was relaxed.The common return bootstrap re-standardizes returns and recomputes the GMV denominator in each draw.
- Design: Table 6 evaluates representation sensitivity on the common 52-firm, 1,207-date standardized-return panel across seven base representations and three matched EttaX vintages.The table reports fixed-law allocation percentiles and relative variance, where lower values are better.
- Results: The full-width-to-1,024 Qwen3-8B contrast includes zero, so moderate truncation does not reject a zero difference in relative variance.At fixed 1,024 width, the Qwen3-4B comparison and the 64-coordinate Qwen3-8B stress cell differ after Holm adjustment.
- Results: None of the EttaX vintage intervals satisfies the stated equivalence bound.The EttaX rows use a benchmark-calibrated post-specified interval comparison.