惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

MyScale Blog
MyScale Blog
博客园 - 司徒正美
A
About on SuperTechFans
Vercel News
Vercel News
H
Hackread – Cybersecurity News, Data Breaches, AI and More
爱范儿
爱范儿
I
InfoQ
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
博客园_首页
Google DeepMind News
Google DeepMind News
T
Tailwind CSS Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
F
Fortinet All Blogs
S
SegmentFault 最新的问题
阮一峰的网络日志
阮一峰的网络日志
D
Docker
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
G
Google Developers Blog
Stack Overflow Blog
Stack Overflow Blog
M
MIT News - Artificial intelligence
Jina AI
Jina AI
H
Help Net Security
量子位
IT之家
IT之家

Show HN

GitHub - astefanutti/shaderbang: Shebang for Shaders Show HN: Generate Claude Code Workflows using Spec Driven Development approach Show HN: AI agents for UK GDAD PCF roles and their skills The Two Pillars: Mixer Mode and Meta-Software in the Reorganization of Software Work After AI GitHub - JaiCode08/teleport-env What 1,000+ Harness Experiments Taught Me About Self-Improving Agents Show HN: Liiists, a Markdown-first, iOS and CLI list app SwiperTab – Get this Extension for 🦊 Firefox (en-US) GitHub - kouhxp/fftext: Summarize, explain, fact-check, or translate any text, URL, or file. No GPU. No cloud. One command GitHub - sweetpad-dev/sweetpad: Develop Swift/iOS projects using VSCode GitHub - dogmaticdev/IRON: IRON a.k.a. Intermediate Representation Object Notation is a Interpreter/Database that is used to create Programming Languages. GitHub - sjhalani7/vaen: Package your AI coding harness into a portable .agent file, and share it across repos, teams, & the community without ever having to copy-paste instructions, skills, MCP config, or secrets. Show HN: Gandalf the Grader Show HN: Citadeld – replay any CI failure locally from a single file GitHub - tdortman/cuSBF: High-Performance GPU Super Bloom Filter coral-ai/claude-code-token-xray at main · Coral-Bricks-AI/coral-ai GitHub - ulyssestenn/funes: Funes is a Git-based framework for LLM-managed knowledge work: an AI Librarian ingests raw sources, builds an interlinked Markdown knowledge base, and uses it to produce cited reports, analyses, and other outputs. GitHub - ThatXliner/gah: Git Add Hunk, built for agents to use GitHub - harmont-dev/harmont-cli: Command-line client for the Harmont CI platform GitHub - brooksmcmillin/mcp-authflow: OAuth 2.0 Authorization Server framework for MCP servers GitHub - javaid-codes/audit-supply-chain-agents GitHub - amorey/gochan: A small library of common channel architectures for Go, inspired by Rust GitHub - arifozgun/OpenGem: Free, Open-Source AI API Gateway with Gemini, OpenAI & Anthropic Compatibility in 1 file GitHub - Pranesh950/BioPetals: 🌸 Run BIOxAI models at home, BitTorrent-style. Fine-tuning and inference up to 10x faster than offloading GitHub - cnguyen14/bounty-doctor: Diagnose a GitHub bounty issue before you waste hours: detects honeypot scam repos, AI-bot attempt swarms, and stale contests. Show HN: CoreMCP – MCP Server for On-Prem DBs Show HN: KittyHTML – Render HTML/CSS as an inline image in your terminal GitHub - bingud/filemat: Web-based file manager Show HN: TruthLens – Free multi-signal deepfake image detector GitHub - apexlocal-jz/claude-usage-tray: Windows system-tray app showing your Claude Code rate-limit usage at a glance. Zero deps, ~300 lines of PowerShell. Cross-IDE (works regardless of VS Code, Cursor, plain terminal).
Who
heterodoxjed · 2026-06-23 · via Show HN

A small experiment in language-model memory

When a language model is trained, what it absorbs about the world gets baked into its weights — the billions of numbers that hold everything it knows. A person is in the weights if a model can recall them on its own, without looking them up — but that's rarely all-or-nothing, so intheweights.com asks 13 models how sure each is, scoring its confidence from 0 to 100. (Combine a person's 13 confidences and the site gives them one strength score; we work with the 13 underlying numbers.) We ran 291 real people through it — from household names to the genuinely obscure — and looked for patterns.

How the 13 models score one person

Each model's confidence runs from 0 ("never heard of them") to 100 ("sure who they are"). It's not yes-or-no, and it's not one verdict — for the same person, the 13 answers can land anywhere from 0 to 100. Here's one person — a Georgian judoka:

The bars are the 13 models' confidence. Some are sure; several have no idea who he is — pooled together, those make up his strength. Now do this for 291 people.

More looked up → better known, but loosely

Each dot is one of 185 people. Across the bottom: how often they're looked up — their monthly Wikipedia pageviews. Up the side: how well the 13 models know them, their 13 confidences averaged into one score. It climbs — household names sit near the top — but loosely: among the rarely-looked-up, the models know some people surprisingly well and draw a blank on others just as obscure.

Pageviews are on a log scale — each step right means about 10× more monthly views, so the famous and the obscure fit on one chart. Hover any dot for who it is, its score, and a Wikipedia preview; click through to the article.

Some models are far more confident than others

Average each model's confidence across the people we ran, and the averages run from one model that's sure of almost everyone (about 90 out of 100, top) down to one that's blank on almost everyone (about 18, bottom). So how high a person scores depends a lot on which model you ask, not just on who the person is.

Bar length = that model's average confidence across those people, 0–100. Hover a model's name for its size and knowledge cutoff.

But they mostly rank people the same way

Scoring people high and ranking the same people high are two different things — a model can hand out higher numbers across the board but still put the same people on top. So set the overall levels aside and compare orderings: line up each model's people from best-known to least-known. By that test the models match moderately well: a typical pair ranks the people about 0.65 alike, where 1 would mean identical orderings and 0 means no relation at all. So they mostly agree on who is better-known than whom, even where they disagree on the exact scores. The clear exception is the smallest model, whose ordering barely matches the others.

Your line of work barely matters — except for athletes

You might expect some kinds of people to live in the weights more than others. Mostly they don't: split by occupation, recognition is strikingly flat — six of the seven groups sit within a few points of each other, at broadly similar fame. The one clear exception is athletes, who lag the rest. So apart from sport, what you did for a living barely changes whether the models know you.

Average recognition (0–100) per occupation, across the people we ran; their fame is broadly similar across groups, so this isn't just a fame gap. Athletes (orange) stand out.

A shared name costs you a little

Does sharing your name with other notable people hurt? We compared German footballers whose full name is theirs alone on Wikidata to ones whose exact name is shared by several other notable people, matched on fame — so the only real difference is the name:

The short version

Whether a model "knows" a person tracks how often the world looks them up — loosely. A shared name hurts a bit. And while the 13 models broadly agree on who ranks as better- or lesser-known, they differ enormously in how confident they are overall — so whether a borderline person counts as "in the weights" still comes down to which model you ask.

Every number here describes these 291 people — a deliberately wide spread from famous to obscure, not a random or representative sample. Read them as comparisons, not rates.