




















I started thinking about this from a narrow problem: how to improve the answers a language model gives.
When someone says an AI answer is "bad," the word does not help much. Bad how? It might have invented a fact, missed what the user wanted, answered at the wrong level of abstraction, refused something harmless, complied with something it should have refused, lost the thread from three turns ago, or solved the benchmark-shaped version of the task instead of the real one.
Those are different failures, and they usually need different fixes. Until you split "bad" into kinds, you can collect thumbs-downs and average scores forever without learning much, because you are measuring something you have not yet understood.
The more I sat with this, the less AI-specific it felt. A domain becomes workable when you learn which distinctions matter, and the structure that holds those distinctions is usually dismissed as paperwork.
A taxonomy is not a filing system. It is a theory of what differences matter.
The cleanest job a taxonomy can do is prediction.
The periodic table is the obvious example. It does not just list the elements; it arranges them so that recurring properties line up, and the arrangement was reliable enough that Mendeleev could leave gaps and describe the properties of elements nobody had discovered yet. That is the rare case where a classification organizes known facts so well that it points toward unknown ones.
Most taxonomies never manage anything like that, and it is easy to undervalue the work they do if the periodic table is the only standard you hold them to.
Some taxonomies exist to coordinate. The International Classification of Diseases gives hospitals, governments, researchers, and insurers a shared way to record illness and death. A diagnosis code is crude next to a person's actual condition, but the crudeness is partly the point: thousands of institutions have to record similar cases in similar ways, or the data cannot be compared at all.
Triage categories do something different. They allocate. When more people need help than there are hands or supplies to help them, the categories exist to decide what happens first, not to understand any one patient in full.
And some taxonomies govern. A platform's categories for spam, harassment, hate speech, self-harm, and scams shape far more than a moderation report: they determine which posts get queued for review, what moderators are trained to notice, which appeals are possible, and what the platform can later claim about its own behavior. A harm with no category is hard to count. A harm filed under the wrong one is hard to understand.
Separating these jobs explains why some debates go in circles.
Take the DSM, the manual of mental disorders. A standing criticism is that its diagnoses do not cut reality at clean joints: conditions overlap, boundaries move between editions, symptoms vary from person to person. Fair enough, if we are judging the DSM as a predictive taxonomy, something aspiring to the condition of the periodic table. But the DSM spends most of its life doing other jobs. It gives clinicians and researchers a shared language, tells insurers and hospitals what counts as a diagnosable and reimbursable condition, and shapes which research gets funded and which kinds of suffering become legible to institutions. A category can be scientifically messy and still carry enormous administrative force. The first question to ask about a taxonomy is not whether it is true but what it is good for.
None of these systems is merely describing something. Triage changes who gets treated first; a diagnosis can unlock insurance or flatten a person into criteria; a platform's harm categories decide what gets enforced. The labels become part of the reality people have to live inside.
Ian Hacking called one version of this the looping effect. Classify people a certain way — as hyperactive children, as refugees, as a personality type — and they respond to the classification: adopting it, resisting it, organizing around it, until the category describes a population it helped produce. James Scott told a similar story about states in Seeing Like a State, and Bowker and Star about infrastructure in Sorting Things Out, but the shared observation is that classification is never only bookkeeping. Naming the kinds is already a decision about what a system can see and what it is allowed to do.
People often ask whether a taxonomy is true. Sometimes that is the right question; when the job is prediction, truth is the standard. But a taxonomy built for reimbursement, policy, moderation, shelf placement, or emergency response should be judged by how well it does that work and by what damage it causes along the way.
A system built to coordinate millions of cases will flatten the individual case. That is the price of the job, not a defect in the design. The ICD has to treat a particular illness as an instance of a type, or health data stops being comparable across hospitals and countries. The same flattening that serves an insurer can do real harm in a therapist's office, where the work is to see one person clearly. Similar operation, different job, opposite verdict.
This is why the wrong taxonomy can be worse than no taxonomy. With no categories you are in fog, and at least you know you are in fog. With bad categories you get the comfort of precision before the distinctions have been earned. A team invents five labels for model failures and starts charting them, and the chart looks like progress even if the labels are renamed vibes. A product team decides users churn for five reasons, and from then on every messy reason has to fit one of the five. The map stops reporting the territory and starts overruling it.
Expertise is partly the ability to live between those two dangers: too few categories, and too much faith in the ones you have.
Beginners lack distinctions and see only the gross category: the code is broken, the essay is bad, the model failed, the patient is sick. Intermediates tend to have the opposite problem. They have learned the official taxonomy and begun treating every case as an instance of it, with a confidence the boxes have not earned. Experts use the taxonomy while seeing its seams. They know which distinctions are load-bearing, which are teaching devices, which are political compromises, and which are probably temporary. They have not rejected the map; they just remember what it was made for.
This suggests a useful way into a new field. The beginner asks, "What should I read?" A better question is, "What distinctions do insiders make that outsiders miss?" The answer often lives less in the textbook than in the workflow: what a practitioner checks first, which cases they refuse to lump together, which failure modes have earned names, and, most revealing, which categories exist because they track reality and which exist because an institution needs to process the world at scale.
Linnaean ranks let naturalists coordinate observations across the world for two centuries. Then evolutionary thinking showed that some of those tidy categories were useful handles rather than faithful pictures of ancestry. That is the common lifecycle: a taxonomy helps people see for a while, and then, if the field keeps moving, the same taxonomy starts getting in the way. The best one does not pretend to be reality. It is useful enough to act with, explicit about the job it is doing, and easy to revise when its distinctions stop working.
Which brings me back to the model answers. What I needed was never a better scoring system; it was a theory of which differences in "bad" matter. A hallucinated fact and an answer at the wrong altitude can earn the same thumbs-down while implicating entirely different parts of the system, and no amount of averaging will surface that. Once the kinds have names, a pile of complaints turns into cases, cases turn into patterns, and patterns turn into decisions about what to change. That is why taxonomy work looks like administration from the outside and feels like power from the inside.
A taxonomy is a theory of what differences matter. To improve one is to improve what you are capable of noticing.
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。