Blog · · Ashutosh Agarwal

Idea Generation in the Age of AI: Where Hedge Fund Ideas Come From

13Fs are stale, screens converge on the same names, and a conference pitch is crowded the moment it is delivered. Where analysts actually find non-consensus ideas now, and how AI agents mine podcasts and run screens nobody else has.

TL;DR

  • The classic channels for hedge fund idea generation are either stale, crowded, or both, and crowding is now measurable enough that a PM can price it into the bar they set for your pitch
  • The portfolio still demands a steady stream of new names, and the analyst is the one on the hook for it week after week
  • Two sources remain largely untapped: the podcast corpus, which is unstructured enough that almost nobody mines it at scale, and agentic artifacts, screens built from your own criteria that rerun themselves every morning

A hedge fund portfolio is a living system that constantly demands attention. Names get stale, a thesis breaks, management turns over, the market moves in a direction nobody modeled, and the PM is always looking to upgrade a portion of the book. They rely on analysts for a steady stream of new names, which makes idea generation one of the analyst's most important functions, alongside deep research and the pitch itself.

So analysts wake up each morning thinking about where the next idea comes from and how to pitch it to maximize the odds of getting an allocation. More exposure to your ideas means more chance of a large bonus at the end of the year. But when you count the number of names that have to be scanned each week to surface a few good ones, the math is unforgiving, even for a specialist.

A sector analyst covering twenty to twenty-five names sifts roughly twenty raw ideas for every one worth taking to a PM. That is about one pitchable idea a week, which sounds manageable until you annualize it: hundreds of touches, since the same name gets revisited and a live watchlist keeps growing. Call it forty pitches a year, ten to fifteen names that actually get sized if the year goes well, and three to five new high-conviction names in the book. The numbers move with the structure of the fund, its concentration and its turnover, but the shape holds. Analysts are permanently on the hook for new ideas with real upside.

A bad idea rarely makes it to the PM at all, and anything obvious or crowded faces a very high bar. So where do the differentiated ideas come from? Today your channels are probably the same as everyone else's, and that is exactly the problem.

Where hedge fund ideas come from today

Here is the honest version of the classic idea sources, what each is still good for, and where each one breaks down.

Channel Good for Where it breaks down
13F filings Seeing what respected funds owned, and building a picture of who is accumulating Filed 45 days after quarter end, so an early-quarter trade can be 135 days old before you read it. The aggregations, like the Goldman VIP basket, are read by everyone
Conference pitches A fully worked thesis, delivered free, by someone who has done the work The entire room hears it at once, so it is non-consensus by definition for about a day. Our Sohn 2026 scorecard tracks how those pitches actually performed
Quantitative screens Narrowing a large universe to something workable, fast Everyone runs similar fundamental filters against the same Capital IQ, Bloomberg and FactSet fields, so the output converges on the same shortlist
Sell-side research Sector context, model cross-checks, and knowing what consensus believes Thirty-plus firms land in every inbox before the open. A useful research input, not a differentiated idea
Alternative data Tracking a thesis you already hold, between prints Adoption reached roughly 90% of funds in 2025, up from 62% in 2023. Broad access means the edge now sits in how you use it, not in having it
Expert networks Vetting an idea and filling one specific gap the documents cannot $500 to $1,500 an hour, and most of your competitors hold the same seats. Better for confirmation than for discovery
Idea dinners Pressure-testing a view against smart peers Structurally a crowding machine. The first mover profits precisely because everyone else piles in behind them
Podcasts Operator-level detail on niche end markets, often more candid than a monitored call Unstructured, and there are 120M+ episodes. Unusable by hand, which is exactly why it is not yet crowded
Agentic artifacts A screen built on your own qualitative and quantitative criteria that reruns itself daily You still have to know what to care about. The agent runs your thinking, it does not supply it

Alternative data is the clearest example of the pattern that eventually catches every channel. In 2023 it was an edge for the 62% who had it and knew what to do with it. By 2025, at roughly 90% adoption, it became the cost of not being surprised by what better-informed traders already knew.

The rest of the list has the same problem in a slower form. Everyone attends the same conferences, reads the same notes, and screens the same fields, which produces the herding you can watch in real time whenever a crowded name breaks. PMs know it, and they set a much higher bar for anything that has already reached them through a traditional channel.

Crowding is measurable now, which raises the bar on your pitch

Goldman publishes a Hedge Fund VIP basket of the fifty stocks that appear most often among funds' top-ten holdings, which means the industry's crowding is not a vibe, it is a published list. Amazon has shown up as a top-ten holding for around 120 funds, Meta for roughly 80, Microsoft for about 75.

Pitching a name on that list is hard for two separate reasons.

The first is the question you cannot dodge: what do you know that the rest of the market does not? When 120 funds already hold the name, the answer has to be specific, provable, and genuinely absent from the sell-side view. That is a much harder standard than being right about the business.

The second is the crowding itself. Even if the thesis clears the bar, the name carries an exit-liquidity problem that has nothing to do with fundamentals, and it gets scrutinized at every market event where funds head for the door at the same time.

It pays to bring ideas that are under the radar. And the harder point is this: an analyst does not just need one non-consensus idea, they need a repeatable system that produces them week after week.

Two sources are still largely untapped. One is mining the podcast corpus for alpha. The other is building your own agentic artifacts, screens that encode your criteria and deliver fresh names every time you log in.

Why podcasts work for idea generation

Podcasts are not acknowledged the way alternative data is, but parts of Wall Street already treat them as a primary research input.

In April 2025, Millennium hired portfolio manager Steve Schurr from Balyasny on a package reported at roughly $100 million. His team had reportedly generated $250 million at Balyasny in 2024, and his process is described as mining overlooked sources, podcast transcripts among them, while hunting names with fewer than three research teams covering them. The method is obscurity by design.

Matt Reustle put numbers on it in July 2025 on Business Breakdowns, in an episode titled "Alpha In Podcasts?" He looked at the 200-plus companies the show had covered since 2021. The aggregate basket lagged the S&P, so podcast beta is clearly not the trade. The interesting part sat in the tails, where AppLovin outperformed by roughly 540% and went on to rise 713% in 2024.

His conclusion is the useful one. The alpha had less to do with marquee interviews and more to do with industry commentary on niche shows. He pointed to an oilfield services podcast where working operators described conditions on the ground that diverged sharply from the sell-side's picture of the sector. That is the shape of a real edge: people who do the work, describing what they see, months before it shows up in a print.

The catch is that you cannot listen your way to it. You need to mine it at scale, know which questions are worth asking, and be able to interrogate the whole body at once rather than one episode at a time.

Searching the corpus instead of listening to it

Managers know there is signal in podcasts. The problem is scale.

There are millions of shows, hundreds of millions of hours of recorded conversation, and thousands of new episodes published every day. No analyst, and no team of analysts, can keep up with that. The bottleneck is not access, it is the ability to search, filter, compare and retrieve the handful of passages that matter to your book across every episode at once.

That is the problem Matterfact Podcast Intelligence was built to solve. Instead of listening at 2x for an hour hoping to catch one useful detail on a name you own, you ask a question and query the corpus. The agent retrieves the relevant passages, compares how a theme is discussed across hundreds of sources, and hands back the hours so you can spend them on the part that pays, which is deciding what the signal means.

For idea generation specifically, this changes the shape of the week. Rather than waiting for the next conference, an analyst can interrogate 120M+ episodes on a Tuesday morning and pull out where operators disagree with the sell side on the real cost of new GPU clusters, or which end markets a specialty chemicals executive quietly stopped talking about.

Ask a question about a name everyone thinks they already understand, and what comes back is not a summary. It is the argument, split into the two sides that people who work in the industry are actually having.

Ask: From podcasts, find out all the debates about Nvidia's moat

Where the debate on Nvidia's moat actually splits

The moat holds

  • CUDA and the surrounding libraries are the switching cost, not the silicon
  • Rack-scale interconnect is where challengers are furthest behind
  • An annual cadence keeps the performance-per-watt gap open

The moat is narrowing

  • Hyperscaler in-house accelerators absorb the steady-state inference work
  • Inference buyers are far more price-sensitive than training buyers
  • A handful of customers carry the revenue, and they know it

Synthesized from 57 episodes across 41 shows, with the speaker and timestamp behind every line

That is one query. The answer comes back in a form you can take into a morning meeting, and you can see which shows are worth your attention before you ever press play.

Idea generation without prompt engineering

Most analysts are not prompt engineers, and the blank prompt box is where AI for research quietly fails. We wrote about how to write prompts for investment research, and then we pre-engineered the prompts so nobody has to start from scratch.

Playbooks are packaged, customizable prompts that run in one click, and they work across podcasts, market data, filings and performance history rather than on any single source.

Ten of them are built specifically for idea generation. Each one is a screen with a thesis attached, not just a filter.

Idea Generation playbooks

Surface candidates. Quantitative screen, then qualitative thesis.

  • Custom Quantitative Screen: Define your own criteria, P/E, ROE, growth, FCF yield, then narrow qualitatively to a shortlist.
  • Quality Compounder Screen: High ROIC, strong FCF conversion, low debt, a long reinvestment runway. Buffett-style compounding.
  • GARP: Growth at Reasonable Price: PEG below 1, double-digit growth, healthy returns, without paying a high-flier multiple.
  • Deep Value / Distressed Screen: Cheap on every metric, plus a credible recovery path. Filters out the value traps.
  • Earnings Revision Inflection: The first up-revision after a streak of downgrades. The turn is the alpha.
  • Margin Inflection Screen: Gross-margin trough plus operating leverage about to flip. The earnings J-curve.
  • SOTP Dislocation Screen: Multi-segment companies trading more than 20% below sum-of-parts, with a breakup catalyst.
  • Net Cash / Balance Sheet Cheap: Net cash above 30% of market cap and the operating business is profitable. Almost free optionality.
  • Capital Return Yield: Buyback plus dividend yield above 8%, FCF coverage above 1.2x. The compounding income screen.
  • Cyclical Trough Mean-Reversion: At a trough multiple versus its own history, with mid-cycle EPS power signalling upside.

Those ten sit inside a library of more than two hundred playbooks covering diligence, modeling, valuation, catalysts and forensics, so the name that clears this stage has somewhere to go next.

Open one and you get sensible defaults, the few things worth overriding, and the exact prompt it is about to send. Nothing runs behind your back.

Quality Compounder Screen (Idea Generation playbook)

High ROIC, FCF conversion, low debt, long reinvestment runway. Buffett-style compounders.

Default thresholds: ROIC > 15% (5y avg), FCF/NI > 80%, net debt/EBITDA < 1.5x, EPS CAGR > 8% (5y), gross margin > 40%.

Universe: SP500 (S&P 500 constituents), NASDAQ100 (Nasdaq 100 constituents), DOW30 (Dow Jones Industrial Average 30 constituents).

Runs monthly on the 1st at 8:00 AM ET by default.

Prompt sent:

Run a quality-compounder screen on NASDAQ100. Thresholds: ROIC > 15% sustained over the last 5 years, FCF / net income > 80%, net debt / EBITDA < 1.5x, EPS CAGR > 8% over 5 years, gross margin > 40%. Then narrow qualitatively: durability of the moat, reinvestment runway, management's capital allocation record, and what would have to break for the compounding to stop. Return a ranked shortlist with the case for each name and the one figure that would falsify it.

The thresholds are common property. What you ask the agent to do with the survivors is not, and that is where the differentiation starts.

So treat every playbook as a starting point you are meant to edit. Tweak the structured prompt until it reflects how you actually think about a sector, or skip the library and describe what you are looking for in plain English. Ask for signs of a widening moat, or for fundamentals deteriorating faster than the multiple suggests, and the agent turns that into a screen and does the heavy lifting.

Then you overlay whatever else matters: quantitative filters, profitability thresholds, balance-sheet constraints, liquidity floors.

Turn the screen into a live artifact

If you like what comes back, the workflow does not have to end there. You can turn it into a live artifact that reruns the screen and feeds you names every time you log in.

S&P 500 top movers

A few clicks and you have an agent running a screen you built from your own criteria. Nobody else has that exact screen, which is the whole point when it is time to pitch. We walked through how one of these gets built from a single prompt in an earlier post, and the same approach carries into deep diligence on the name once it clears.

A weekly system, not a lucky week

The analysts who consistently bring the PM something new are not luckier, they run a loop. Here is the version worth stealing.

Run this every week:

  • Run your own screen, built on your criteria, not the standard fields everyone else filters on
  • Query the podcast corpus on your sectors and read where operators disagree with the sell side
  • Check the new names against a crowding list before you get attached to any of them
  • Take the survivors into deep diligence, with the bear case written down first
  • Turn whatever worked into an artifact that reruns itself, so next week starts further along

The point is not to replace the classic channels. It is to add two that your competitors are not systematically mining.

None of this replaces judgment. Deciding which thesis matters, which bear argument is real, and how hard to push is still the part that gets rewarded, and it is still entirely yours. What changes is where you start. If the analyst next to you is screening the same fields against the same database and hearing the same pitch at the same dinner, the ideas you both bring will look alike, and only one of you will be able to answer the question about what the market is missing.

That is the whole game in idea generation: get to the name before the crowd does, and be able to show your work when the PM asks how you found it.

FAQ

Where do hedge funds get stock ideas?

Mostly from a standard set of channels: quantitative screens, sell-side research, 13F filings and their aggregations, investment conference pitches, expert network calls, alternative data and peer idea dinners. Each is useful, and each is used by nearly everyone, which is why ideas sourced this way tend to arrive at the PM already crowded. The channels that are still under-mined are unstructured ones, particularly the podcast corpus, and custom agentic screens built on criteria specific to one analyst.

What makes a trade crowded, and why do PMs care?

A crowded name is one that many funds hold in size at the same time. Goldman publishes a Hedge Fund VIP basket of the fifty stocks appearing most often among funds' top-ten holdings, so crowding is measurable rather than anecdotal. PMs care for two reasons: it is very hard to explain what you know that 120 other funds do not, and crowded names sell off together whenever funds head for the exit at once, independent of fundamentals.

How do analysts find non-consensus stock ideas?

By sourcing from places the rest of the market is not systematically reading. In practice that means going where the data is unstructured enough to be inconvenient, such as long-form podcast commentary from operators in niche end markets, and building screens that encode your own qualitative criteria rather than the standard fundamental fields available in every terminal.

Are podcasts actually useful for investment research?

For discovery, yes, if you mine them at scale rather than listen episode by episode. Matt Reustle's July 2025 Business Breakdowns analysis of 200-plus covered companies found the aggregate basket lagged the S&P while the tails carried everything, with AppLovin outperforming by roughly 540%. The signal was concentrated in niche industry commentary, not marquee interviews, which is precisely the layer no analyst has the hours to cover by hand.

Can AI agents generate stock ideas on their own?

They can generate candidates, run screens continuously, and surface names that match criteria you defined, including qualitative ones expressed in plain English. They cannot tell you what is worth caring about. The criteria, the judgment about which candidate deserves four weeks of diligence, and the thesis you take to the PM remain the analyst's work.

How is this different from AlphaSense or Tegus?

Those are document and transcript libraries, and they are good ones. The gap is what happens after retrieval. Matterfact runs agents over the corpus and over your own data, turns a research question into a screen, and stands the result up as a live artifact that reruns on its own schedule rather than as a search result you have to repeat by hand.