Do LLM “Crowds” Produce Investment Signals? An Empirical Test

The integration of artificial intelligence into algorithmic trading has ignited a race to transform generative text into systematic alpha. A new paper written by Steven Edwards empirically investigates whether constructing a synthetic consensus using large language models can simulate information aggregation dynamics or if it merely acts as a sophisticated echo chamber. By utilizing an expansive framework to evaluate portfolio construction across distinct synthetic mandates, the study challenges whether generative agents can truly democratize the wisdom of crowds within highly efficient capital markets.

The core methodology establishes a robust experimental architecture designed to isolate genuine cross-sectional predictability from systematic factor exposures. The author queries an ensemble of 100 philosophically distinct investor archetypes to generate stock selections, aggregating the results into a consensus strategy. Crucially, the backtesting window is implemented strictly after the model’s structural data limit to neutralize structural look-ahead bias. To map performance drivers, the resulting allocations are subjected to an asset pricing asset-decomposition framework using standard systematic risk premiums augmented with a momentum factor.

The empirical findings reveal a striking operational paradox between absolute performance and informational diversity. On one hand, a market-capitalization-weighted allocation of the consensus tickers achieves a remarkable 29.5% annualized return, yielding a statistically significant 9.65% annualized alpha that survives rigorous factor decomposition. On the other hand, ensemble methods exhibit rapid convergence, with as few as 25 personas capturing 88% of the full universe’s composition—mirroring a single, unprompted neutral baseline by 92%. This extreme overlap confirms a systemic violation of the independence condition required for true collective intelligence.

Ultimately, the apparent outperformance of the LLM crowd is an artifact of regime-specific beta rather than authentic security selection skill. The structural convergence toward high-beta, mega-cap technology equities happened to perfectly align with an intensive, AI-driven market re-rating during the specific out-of-sample window. For institutional allocators, the takeaway is clear: free-recall prompting fails to generate novel investment signals and instead maps existing financial media prominence. Quantitative managers should pivot their LLM utilization away from standalone stock picking and toward processing unstructured, specialized data at scale to unlock true analytical leverage.

Authors: Steven Edwards

Title: Do LLM “Crowds” Produce Investment Signals? An Empirical Test

Link: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=6543239

Abstract:

This paper tests whether aggregating stock selections across a large, philosophically diverse ensemble of large language model (LLM) personas can produce investment signals beyond passive benchmark exposure. One hundred distinct investor personas, spanning value, growth, momentum, ESG, quantitative, and contrarian philosophies, were constructed using GPT-4o and each asked to select fifty US-listed equities. The fifty most frequently mentioned stocks formed a consensus portfolio, backtested across three weighting schemes over a 561 trading-day period beginning January 2024, strictly after the model’s October 2023 training cutoff to eliminate look-ahead bias. The market capitalisation weighted consensus portfolio returned 29.5% annualised versus 15.3% for the S&P 500 ETF (SPY), with a statistically significant annualised alpha of approximately 10% surviving a six-factor model comprising the Fama-French five factors augmented with momentum (t = 3.27, p = 0.001). However, constituent analysis reveals that all fifty portfolio stocks are S&P 500 members, and a neutral single-prompt control portfolio shares 92% of constituents with the 100-persona ensemble. Bootstrap sampling demonstrates convergence to approximately 88% of the full-ensemble composition by N = 25 personas. The evidence suggests LLM persona ensembles reproduce media coverage prominence rather than generating independent investment signals, consistent with a violation of the independence condition required for genuine crowd wisdom. The residual factor-adjusted alpha most likely reflects concentrated exposure to large-cap technology equities during an AI-driven market re-rating rather than genuine stock selection skill; replication across regimes is required.

As always, we present several interesting figures and tables:






Notable quotations from the academic research paper:


“Can a synthetic crowd of LLM investor personas simulate the information aggregation dynamics observed in real investor populations? […] One hundred investor personas with distinct investment philosophies and calibrated stochasticity parameters were constructed and queried via the OpenAI GPT-4o API. Each persona independently selected fifty US-listed equities, producing a universe of 1,059 unique tickers. Three weighting schemes were evaluated over a 561-trading-day out-of-sample period, with six-factor decomposition applied to isolate genuine alpha from systematic exposures. A bootstrap ensemble size analysis assessed whether the persona apparatus produces stable selections, a statistical test evaluated whether portfolio composition differs from a random draw from the S&P 500, and a neutral single-prompt control portfolio was constructed to isolate the marginal contribution of persona diversity.


[…] Constituent analysis and bootstrap results are consistent with the echo chamber hypothesis: the consensus portfolio reproduces the most media-prominent US equities, shares 92% of its constituents with a single neutral LLM call, and converges to approximately 88% of its final composition with as few as 25 personas. At the same time, the market capitalisation weighted portfolio generates statistically significant annualised alpha of approximately 10% that survives six-factor adjustment. Whether this reflects genuine return-relevant information encoded in the LLM training corpus or a regime-specific artefact of the evaluation period cannot be determined from a single market cycle.


This paper tests whether a large ensemble of philosophically diverse LLM investor personas can replicate the information aggregation dynamics of genuine crowd wisdom. On the question of information aggregation, the evidence is clear: it cannot. The independence condition is violated by construction, the consensus portfolio shares 92% of its constituents with a single neutral LLM call, and bootstrap analysis confirms that 25 personas provide essentially the same result as 100. LLMs trained on the same corpus produce correlated selections regardless of persona framing, and the resulting portfolio reproduces the most media-prominent equities in the US market.


On the question of performance, the picture is more nuanced. The market capitalisation weighted portfolio generates statistically significant alpha of approximately 10% that survives six- factor adjustment. Whether this reflects genuine return-relevant information in the training corpus, an incomplete factor model, or a regime-specific artefact cannot be determined from a single market cycle.”


Are you looking for more strategies to read about? Sign up for our newsletter or visit our Blog or Screener.

Do you want to learn more about Quantpedia Premium service? Check how Quantpedia works, our mission and Premium pricing offer.

Do you want to learn more about Quantpedia Pro service? Check its description, watch videos, review reporting capabilities and visit our pricing offer.

Do you want algorithmic access to the full Quantpedia database via the API? Subscribe to Quantpedia Pro, ask for an API key, and explore the in/out-of-sample statistics, source academic papers, and code snippets — ideal for quantitative research, systematic trading workflows, and AI model training.

Are you looking for historical data or backtesting platforms? Check our list of Algo Trading Discounts.


Or follow us on:

Facebook Group, Facebook Page, Telegram, Twitter, Linkedin, Medium or Youtube

Share onRefer to a friend
Subscription Form

Subscribe for Newsletter

 Be first to know, when we publish new content
logo
The Encyclopedia of Quantitative Trading Strategies

Log in

SUBSCRIBE TO NEWSLETTER AND GET:
- bi-weekly research insights -
- tips on new trading strategies -
- notifications about offers & promos -
Subscribe
QuantPedia
Privacy Overview

This website uses cookies so that we can provide you with the best user experience possible. Cookie information is stored in your browser and performs functions such as recognising you when you return to our website and helping our team to understand which sections of the website you find most interesting and useful.