Alex Yarosh Get Free Snapshot

Source record

@tjrobertson52 TikTok profile avatar

@tjrobertson52

2025-08-14

Your ChatGPT prompts = $$$. Companies are paying millions to know what you're typing. Here's the wild industry you didn't know existed ๐Ÿ‘€ #ChatGPT #AIData #PromptData #TechNews...

Share source record
Tiktokexcerpt onlyen2

Source Text

i

People like me wanna know what you're typing into ChatGPT. In fact, you wanna know it so badly that there's now a multi million dollar industry around collecting this prompt data. So I wanna talk about why this data is so valuable, why it's so hard to get, and what it can mean for the future.

Every month, more and more people are using large language models to make big decisions. A lot of these decisions are purchase decisions. I think we're quickly entering a world where most purchase decisions are at least somewhat influenced by large language models.

And so it's no surprise that businesses are willing to invest a lot of money to be recommended by these large language models. This is currently the No. 1 thing we're helping our clients with.

And if you want to increase your visibility in the large language models, you need to know what sources these large language models are using to make their recommendations. Currently, the only way to get these horses is to run the prompts through the large language models and look at what websites they're referring to before they come back with the recommendation.

The problem is we actually have no idea what people are typing into ChatGPT. Currently we're just guessing. We guess a bunch of prompts that someone might type into ChatGPT.

We look at those sources, and then we try to get our clients mentioned and recommended on those sources. Or we create better sources and hope that the large language models use our sources instead.

This is somewhat similar to the process we do in traditional SEO with keyword research. But the big difference is we actually have a pretty good idea of what search terms people type into search engines. And this is because Google shares that data.

ChatGPT does not share that data.

Now I think in the future, more and more of these AI searches are gonna be done in Google using Google's AI mode. And while Google does share that data, they currently only share searches that are performed many times. This is largely for privacy reasons.

If a lot of people search for the same term, they can aggregate that data into a single data point. So you can't extract any personal information about the person who made the search. However, the typical prompt is much longer and more unique and unlikely to be typed multiple times.

So as things stand, it seems unlikely that we'll get direct access to what people are typing into these large language models.

Naturally, wherever there's demand, there's gonna be businesses popping up to capitalize on that demand. So now of course there are a series of Chrome extensions that will capture this data and then sell it. Some of this is done ethically.

The person installing the Chrome extension understands that they're selling their data, and in exchange they might get access to the larger data set. But some of this is being done without the users knowledge, and that's definitely unethical.

So I'm sure in the next year we're gonna see startups finding more and more sophisticated ways to capture this data or simulate this data. But ethics aside, the tools that have this data are just gonna be way more useful than tools that don't. Well, I'll do my best not to support any companies that are harvesting this data unethically for my clients.

I am looking for tools that have this data cause as it is with all forms of marketing, without data, we're really just guessing.

Source Intelligence

i

The source frames AI-visibility research as guessing buyer prompts, running them through LLMs, and inspecting referenced websites.

2 related signals ยท Prompt research / AI sources / Risk/avoid / prompt data ethics

  • Build prompt-simulation workflows that record prompt assumptions, cited pages, and source opportunities, while labeling prompt-volume uncertainty clearly.
  • Avoid prompt-data vendors or extensions that lack informed consent, clear data provenance, and transparent methodology.

Questions this source answers

i

What is this source mainly about?

The source frames AI-visibility research as guessing buyer prompts, running them through LLMs, and inspecting referenced websites.

What should an operator take from it?

Build prompt-simulation workflows that record prompt assumptions, cited pages, and source opportunities, while labeling prompt-volume uncertainty clearly.

Which topics does it connect to?

This source is connected to Prompt research / AI sources, Risk/avoid / prompt data ethics.

What public evidence supports the record?

People like me wanna know what you're typing into ChatGPT. In fact, you wanna know it so badly that there's now a multi million dollar industry around collecting this prompt data. So I wanna talk about why this data is so valuable, why it's so hard to get, and what it can mean for the future...