惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Microsoft Azure Blog
Microsoft Azure Blog
GbyAI
GbyAI
P
Proofpoint News Feed
Engineering at Meta
Engineering at Meta
Recent Announcements
Recent Announcements
L
LangChain Blog
B
Blog
阮一峰的网络日志
阮一峰的网络日志
Microsoft Security Blog
Microsoft Security Blog
博客园 - 【当耐特】
M
MIT News - Artificial intelligence
D
Docker
WordPress大学
WordPress大学
J
Java Code Geeks
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
The GitHub Blog
The GitHub Blog
博客园 - 叶小钗
Last Week in AI
Last Week in AI
Stack Overflow Blog
Stack Overflow Blog
有赞技术团队
有赞技术团队
MyScale Blog
MyScale Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
MongoDB | Blog
MongoDB | Blog
博客园 - Franky

Forbes - Innovation

Why Do Humans Have Fingerprints? Hint: It’s Not What You Think Booking.com Confirms Data Breach, Reservation PIN Codes Changed Why Major News Sites Are Blocking The Internet Archive’s Wayback Machine iPhone Fold Release Date: New Report Details Frustrating Apple News Comet Tracker: How To See Pan-STARRS And Three Planets On Wednesday NYT Mini Crossword Today: Tuesday, April 14 Hints And Answers Today’s NYT Strands Hints, Spangram, Answers: Tuesday, April 14 (It’s A Little Unclear) Today’s Wordle #1760 Hints And Answer For Tuesday, April 14 Most Of The Microplastics In Urban Air Come From Tires Today’s Wordle #1759 Hints And Answer For Monday, April 13 NYT Mini Crossword Today: Monday, April 13 Hints And Answers NYT Pips Today: Hints, Answers And Walkthrough For Monday, April 13 The YC Chief Who Codes 10,000 Lines A Day Has A Simple Secret Samsung Expands One UI 8.5 Beta To More Galaxy Owners Why You Should Stop Using Your iPhone If It’s On This List Chamath Says Firms That Treat AI As A Strategy Hand Rivals Their Edge 3 Unexpected Habits Of Secure Couples, By A Psychologist The First Lamp That Folds Your Clothes Samsung’s Disappointing Price Update For Galaxy Phone Buyers 3 Subtle Signs Someone Is Falling In Love With You, By A Psychologist Do Mantis Shrimp See More Colors Than Humans? A Biologist Explains NYT Connections Answers Explained For Monday, April 13 (#1,037) NYT Connections Hints Today: Monday, April 13 Clues And Answers (#1,037) LEGO Luigi & Mach 8 (72050) Review: 2026’s Best Set Yet? Marc Andreessen Says AI Productivity Will Trigger A Hiring Boom 3D Printing Is The Ultimate Hack To Reduce Household Spending Apple iPhone Fold: Striking Design Revealed In Leaked Photos Apple Smart Glasses: New Leak Reveals A Major Design Twist To Beat Meta Tested: The AI Coming To The Rivian R2 Quordle Hints Today: Monday, April 13 Clues And Answers
​Great AI Systems Need A Human Touch
Ambarish Majumdar · 2026-06-03 · via Forbes - Innovation

Ambarish Majumdar is a Marketing Science Partner at Meta, where he uses his SME knowledge in Marketing Science to better AI models.

getty

​As enterprises accelerate AI adoption, most conversations begin with model selection, infrastructure and deployment speed. Those decisions matter, but they are rarely what determine whether AI succeeds in production. The real difference between a system that looks impressive in a demo and one that creates lasting business value is human judgment.

That human layer comes from subject-matter experts (SMEs).

Large language models can generate fluent, fast responses, but fluency should not be mistaken for quality. A model can sound confident while giving the wrong recommendation, missing a compliance issue or creating customer friction. In enterprise environments, quality is rarely just about correctness—it is about trust, context and decision-making.

Anthropic captured this clearly in its writing on evaluations: Evals are not the final checkpoint before launch. They are the foundation of responsible AI development.

My experience working with AI models that power ad ranking and relevance for Bing at Microsoft underscores a critical truth: Automated systems, no matter how sophisticated, cannot fully replace human judgment.

Consider a client case I worked on. Our AI models served advertisements for a popular SUV when users searched for the SUV's model name. Internal quality systems flagged these placements as high-fidelity matches; after all, the keyword and the ad were lexically aligned. However, a human review of the user sessions told a different story. The overwhelming search intent behind that term was geographic—users were looking for information about a city of the same name, not a vehicle. What the algorithm scored as a precise match was, in reality, an irrelevant ad experience that eroded user trust and wasted advertiser spend.

This example illustrates a broader principle: AI excels at pattern recognition and scale, but it often lacks the contextual reasoning needed to interpret intent. Because of this, human evaluation remains an essential layer in any system where relevance is the measure of success.​​

SMEs Define What Quality Looks Like

Every AI system is optimizing toward something. In business settings, that “something” cannot be defined by benchmarks alone. SMEs establish the target.

They answer practical questions such as: Is this recommendation trustworthy? Would a customer support agent actually use this answer? Does this output reflect policy, compliance and operational reality?

In healthcare, quality may mean safety and clarity. In finance, it may mean risk reduction and policy alignment. In customer support, it may mean whether the answer resolves the issue instead of simply sounding helpful. Without SME input, teams often optimize for what is easiest to measure instead of what matters most.

Stress Testing Is Not Evaluation

Many organizations assume stress testing is enough. It is not.

Stress testing pushes a model into difficult or extreme scenarios to expose failure points. It helps identify where the system breaks and which edge cases create risk. That is useful, but it does not define quality.

A model passing stress tests does not automatically mean it performs well in everyday business operations. Stress testing shows where failure happens. Evaluation defines what success should look like. The two are related, but they are not interchangeable.

Peer Review Is Not Evaluation Either

Peer review is also valuable, but it should not be confused with a true eval system.

A few reviewers looking at outputs and sharing opinions creates useful feedback, but it is often inconsistent and subjective. One reviewer may approve something another would reject. Strong evaluations require structure. That means clear scoring criteria, repeatable standards, consistent measurement across reviewers and alignment to business outcomes, not personal preference.

In the end, evaluation has to turn judgment into something measurable.

Why Multiple SMEs Matter

One reviewer is rarely enough. Enterprise AI touches multiple functions, and each team sees quality differently.

For example, a legal reviewer may focus on compliance risk, or a product leader may focus on usability. A support leader may focus on customer experience, while an operations leader may focus on efficiency and reliability. If only one perspective is used, important failure modes get missed. Multiple SMEs create stronger signals because real business decisions are never one-dimensional. Trustworthy AI requires a broader lens.

Evals Help Teams Climb The Right Hill

Most AI improvement happens through small iterations—prompt changes, workflow adjustments, retrieval improvements and model updates. This process resembles climbing a hill. You take one step at a time and try to move toward better performance.

But those steps only work if your team knows which direction is actually uphill.

Without strong evals, teams often optimize for speed instead of reliability. Because of this, regressions go unnoticed and surface-level improvements look like real progress. Alternatively, with strong SME-led evals, improvements are often tied to business outcomes. That means change can become easier to validate and progress can become repeatable instead of accidental.

Evals provide direction, not just reporting.

Human Evals First, Automated Evals At Scale

The strongest AI teams have to start with human evaluation. SMEs define the standards, identify failure patterns and establish what quality means. Only after that foundation exists should automated evaluations take over at scale.

Auto evals can then be used for detecting model drift, catching regressions across releases, monitoring consistency across workflows and preserving trust as systems evolve. Automation does not replace SME judgment. It protects it.

Conclusion

The companies that want to succeed with AI don't need the largest models or the fastest deployment cycles. They need to know how to define quality and improve it with discipline.

Great AI systems need a human touch because trust is still built by people, not models.​​​


Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?