BetaMaShop is in public beta. We improve it continuously, and your feedback shapes what comes next.
MaShop/Blog/Tools/In the Weights Reveals What AI Models Remember Abo…
ToolsJune 22, 2026
Read · 5 min
llm memorization · training data

In the Weights Reveals What AI Models Remember About You

In the Weights, built by two ex-OpenAI staffers, scores how deeply a person is baked into AI model training data. Here is how it works.

A new website called In the Weights turns a slightly unsettling question into a shareable number: do the major AI models actually know who you are, purely from what they absorbed during training? Built by two former OpenAI employees, the tool queries a spread of large language models and clusters their answers into a single strength score that estimates how deeply your identity is baked into the models themselves. Mozart and Shakespeare sit near the ceiling. Most of us sit far lower, and finding out exactly where has turned the site into a fast-spreading curiosity.

Key takeaways
  • In the Weights, made by ex-OpenAI staffers Thomas Dimson and Joey Flynn, scores how well AI models recall a person from training data alone.
  • It queries several models, including Grok, Gemini, GPT versions, Claude and Llama, then clusters the descriptions into one strength score.
  • Scores run up to 996, with figures like Macaulay Culkin at 988 near the top of the leaderboard.
  • The creators flag real caveats: models hallucinate, and typos or common names tend to lower scores.

The project was reported by The Decoder and TechCrunch, both of which ran their own names through it. The framing TechCrunch reached for, an AI-centric vanity search, captures the appeal. It is the 2026 version of typing your name into Google to see what comes back, except the question is no longer what the web says about you. It is what the models have memorized.

What In the Weights actually measures

The premise rests on a technical distinction that is easy to miss. When you ask a modern chatbot about a person, it can answer in two very different ways. It can reach for a tool, run a web search, and summarize what it finds. Or it can answer from memory, drawing on patterns encoded in its parameters during training without looking anything up. In the Weights cares only about the second kind of answer.

Those parameters are the weights the site is named after. They are the billions of numerical values that a model adjusts as it trains, and collectively they store what the model has learned. If a person shows up reliably when a model answers from those weights alone, it means the training process treated that person as relevant enough to retain. As the creators put it, showing up means the model considered you relevant enough during training to recall without tools like web search. The score is, in effect, a measure of how much of you survived compression into the model.

That distinction matters because it separates fame on the open web from fame inside a model. Plenty of people have a strong Google presence built on recent pages and social profiles, yet barely register in the weights because their footprint is too thin or too recent to have been memorized. The reverse also happens. A long-dead composer with centuries of text written about him can dominate the weights while having no living web presence to speak of. The tool, in other words, is not a popularity contest in the usual sense. It rewards a specific kind of presence: sustained, text-heavy coverage that a training run was likely to encounter many times.

It helps to think of a trained model as a lossy compression of an enormous body of text. During training the system squeezes a vast corpus into a fixed number of parameters, and like any lossy process it keeps what recurs and discards what does not. People who appear thousands of times across the corpus get preserved in sharp detail. People who appear a handful of times get blurred or dropped entirely. A strength score is a rough readout of where you fall on that spectrum, which is why it can diverge so sharply from how visible you feel online today.

Who built In the Weights and why

The site comes from Thomas Dimson and Joey Flynn, both of whom worked at OpenAI. According to TechCrunch, they arrived at the company through its acquisition of their design startup, Global Illumination. That background shows in the product. The tool is less a research instrument than a polished consumer toy, the kind of thing engineered to be screenshotted and posted.

Dimson has been candid about the motivation. He told TechCrunch that the project grew out of a belief that Google vanity searches are the wrong objective in 2026 as more traffic moves to LLMs. The argument is that as people increasingly ask chatbots rather than search engines, the question of whether a model knows you starts to matter as much as where you rank on a results page. He also acknowledged the more primal hook. The reception, he said, has been insane, because the tool strikes a nerve of wanting to see if you live forever. Being embedded in the weights is a strange new form of permanence, and people want to know whether they made the cut.

"Google vanity searches are the wrong objective in 2026 as more traffic moves to LLMs."Thomas Dimson, to TechCrunch

How the scoring works under the hood

The mechanism is straightforward enough to explain without a math degree, which is part of why the results feel trustworthy at a glance. The site sends a structured prompt to each model it tests, asking something close to: who is this person, give up to ten results, each with a short description and a confidence rating. It then collects those answers across models and groups the descriptions that point to the same individual, rolling everything into a single strength score.

Clustering is the clever part. A single famous name can map to several different people, and a raw count of mentions would muddle them together. By grouping answers that describe the same individual, the tool can tell apart two people who share a name and attribute the score to the right one, at least when the models give it enough to work with. The confidence ratings the models attach to each guess feed into how strongly a given identity registers, so a model that answers tentatively contributes less than one that answers with conviction.

The range tops out at 996. The Decoder reports that figures like Mozart and Shakespeare land at or near the maximum, which fits intuition: these are people written about so extensively that essentially every general-purpose model has absorbed them many times over. On the public leaderboard, TechCrunch notes Macaulay Culkin sitting at 988 with Luciano Pavarotti close behind, a reminder that durable cultural fame translates cleanly into model memory.

The models queried span the major labs. TechCrunch lists Grok, Gemini, multiple GPT versions, Claude and Llama, along with several lesser-known systems. Querying many models rather than one is deliberate. A person memorized by every model is more deeply embedded in the collective training corpus than someone who only shows up in a single system, and combining the answers smooths out the quirks of any individual model. It also guards against a single model's idiosyncrasies skewing the result, since each lab trains on a different mix of data and makes different choices about what to keep.

Note

A high score does not mean the information a model recalls about you is accurate. It only means the model is confident it knows who you are. Confidence and correctness are different things, and In the Weights measures the former.

What ordinary scores look like

For everyone who is not Taylor Swift, the numbers come back far smaller, and the journalists who tested the tool were refreshingly honest about it. At The Decoder, the author scored 175 and a colleague landed at 262, modest figures that reflect a real but limited footprint in the training data. At TechCrunch, writer Anthony Ha did better, scoring 641, which the site placed in roughly the top six percent.

Those mid-range scores are arguably the most interesting part of the exercise. They show that being a working journalist with a long byline history puts you meaningfully into the weights without making you a household name. The gap between a 175 and a 641 is the gap between someone the models half-remember and someone they can describe with some confidence. It is a granular look at a kind of digital presence that did not used to be measurable, and it scales in a way that feels intuitively right: the more durable text exists about you, the higher you climb.

The caveats the creators put up front

To their credit, Dimson and Flynn do not oversell the precision. They flag several ways the score can mislead, and those caveats double as a useful primer on how language models handle facts about people.

The first is hallucination. Models can fabricate biographical details with total confidence, so a high score can sit on top of invented facts. TechCrunch hit exactly this: one model, GPT-5.4 Mini (2026), wrongly flagged Anthony Ha as an ambiguous name form that could refer to multiple people, even though it is a specific person's name. The tool surfaces that the model has an opinion; it does not guarantee the opinion is right.

The second is sensitivity to spelling. Typos lower scores, because a misspelled query no longer matches the patterns the model learned. The third is the common-name problem. People with names shared by many others tend to score worse, since the model's knowledge gets split across all the people who share the name and no single identity dominates. The Decoder also notes that smaller models make it harder to show up, because a model with fewer parameters has less room to memorize the long tail of people. A model with billions of parameters still has to budget that capacity, and the budget goes to the names it saw most. The tool reportedly tests systems down to relatively small models, where only the most prominent figures register at all.

Taken together, these caveats are a reminder that the score is a signal, not a verdict. It tells you something genuine about how the models encode you, but it bundles in noise from name collisions and spelling, plus the simple fact that the systems sometimes make things up. Treating the number as precise would be a mistake the creators themselves warn against.

What In the Weights says about training data and memorization

Beyond the vanity-search fun, the site is a surprisingly accessible window into a serious topic: how much large language models memorize, and about whom. Researchers have spent years studying the fact that these models do not merely learn general patterns; they also retain specific strings and facts from their training data, sometimes verbatim. That memorization is what lets a model answer biographical questions without a search, and it is also at the heart of ongoing debates about copyright and privacy.

In the Weights makes that abstract property tangible. When you see that a model can describe a mid-tier public figure in detail with no web access, you are looking at the direct consequence of how these systems are built. Everything they know from memory came from somewhere in the training corpus, which for general-purpose models means a large slice of the public internet plus books and other published text. A score is a rough proxy for how much of the public record about you was swept into that corpus and preserved.

There is a privacy edge to this that the tool's playful framing partly obscures. For a celebrity, being deeply embedded in the weights is just fame by another name. For a private individual who happens to have a visible online trail, it raises thornier questions about what models retain, whether that retention is accurate, and whether anyone consented to it. The site does not resolve those questions, but by putting a number on personal memorization it makes them harder to ignore. People who would never read a paper on training-data extraction will happily check their own score, and in doing so they bump into the same underlying issue. A surprising score, high or low, is often the first time someone realizes that a model has formed a view of them at all.

Why a vanity score is a sign of the times

It would be easy to dismiss In the Weights as a clever distraction, and on one level that is exactly what it is. But the framing Dimson chose is worth taking seriously. The shift he points to, from search engines toward chatbots as the default way people look things up, is real and accelerating, and it changes what online visibility even means. If a growing share of questions about a person or a business gets answered by a model speaking from memory rather than by a ranked list of links, then being known by the models becomes its own form of presence, distinct from search ranking and not controllable through the usual tactics.

That is the quietly important idea hiding inside a toy. We have spent two decades optimizing for how we appear in search results. In the Weights suggests a second axis is forming, one that measures how we appear in the models themselves, and unlike a search index, the weights cannot be edited on request. A new page can change your Google ranking within days. Changing what a model has already memorized requires retraining or fine-tuning that you do not control, which makes model memory a far stickier kind of reputation than the search era ever produced.

That stickiness is already spawning a cottage industry. A discipline sometimes called generative engine optimization has grown up around the goal of shaping how chatbots describe a brand or a person, mirroring the search-optimization playbook of an earlier era. In the Weights does not promise to move your score, and the creators are careful not to pitch it as a growth tool, but it gives that emerging field a crude scoreboard. If you can measure how present you are in the models, you can at least start to reason about whether anything you publish moves the needle over time, even if the feedback loop is slow and the models only update when they are retrained.

It is worth holding onto some skepticism about how meaningful any single number really is. A model's willingness to describe you confidently says nothing about whether people are actually asking about you, and a leaderboard topped by long-dead composers is a clue that the metric tracks textual saturation more than present-day relevance. The tool is most valuable as a conversation starter, a way to make people look directly at the fact that these systems carry a stored impression of millions of individuals. That impression was assembled without anyone asking, and now there is at least a rough way to see your own place in it.

For anyone thinking about brand reputation or the broader question of how AI systems mediate knowledge, that is worth watching. You can read more of our coverage of the tools reshaping how people work with AI on the MaShop blog, but the simplest takeaway is also the most human one: a lot of people just want to know whether the machines remember them, and now there is a number for it.

Comments 0

0 / 4000Your email stays private.
No comments yet. Be the first.

Keep reading picked for you.

Describe it. MaShop builds it.

Commerce apps and websites from one sentence. No card to start.

Start building