惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

宝玉的分享
宝玉的分享
博客园_首页
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
S
SegmentFault 最新的问题
爱范儿
爱范儿
量子位
大猫的无限游戏
大猫的无限游戏
博客园 - 三生石上(FineUI控件)
酷 壳 – CoolShell
酷 壳 – CoolShell
P
Proofpoint News Feed
博客园 - 司徒正美
Microsoft Azure Blog
Microsoft Azure Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
美团技术团队
Attack and Defense Labs
Attack and Defense Labs
V
Vulnerabilities – Threatpost
L
LINUX DO - 热门话题
月光博客
月光博客
J
Java Code Geeks
A
Arctic Wolf
Apple Machine Learning Research
Apple Machine Learning Research
Cyberwarzone
Cyberwarzone
D
DataBreaches.Net
P
Privacy & Cybersecurity Law Blog
D
Docker
Scott Helme
Scott Helme
S
Schneier on Security
S
Securelist
A
About on SuperTechFans
P
Proofpoint News Feed
V
V2EX
Know Your Adversary
Know Your Adversary
T
Tailwind CSS Blog
I
Intezer
T
Tenable Blog
Last Week in AI
Last Week in AI
MyScale Blog
MyScale Blog
T
The Blog of Author Tim Ferriss
T
The Exploit Database - CXSecurity.com
GbyAI
GbyAI
IT之家
IT之家
G
GRAHAM CLULEY
Hugging Face - Blog
Hugging Face - Blog
Project Zero
Project Zero
The Register - Security
The Register - Security
C
CERT Recently Published Vulnerability Notes
小众软件
小众软件
aimingoo的专栏
aimingoo的专栏
云风的 BLOG
云风的 BLOG
C
Cybersecurity and Infrastructure Security Agency CISA

Salesforce

How We Protect Our Data as Customer Zero Scale Your MRR: Subscription Management For Small Business Streamlining Commerce Media Ad Inventory Management 12 AI Sales Strategies for Startups That Actually Work Sell Smarter: Ecommerce Metrics To Track For Your Small Business Shop Apply the Orchestration Density Framework to Your Next Automation Decision Wait, Black Friday Planning In Spring? It’s Time to Start Holiday Promotions AI-First Operations, One Process at a Time How BCU Is Transforming Banking Service with Agentforce Salesforce Headless 360: What the Agent Consumer Means for Your Integration Architecture Meet Customers Where They Are: Agentforce Contact Center Now Offers WhatsApp Voice 11 Free Lead Generation Tips for Small and Growing Businesses SFR-VibeTrain: The Agent That Trains Agents Strengthening Salesforce Security Against AI-Driven Threats Join Us in the Community Hub at Connections 2026 The Best Way To Build AI Agents That Customers Trust 5 Ways AI is Changing the Communication Game For Startups Trust in the Era of Agents: Highlights from the 2nd Annual Trusted AI Impact Report You Can Be an Agentic Enterprise No Matter What Size Business How to Make Your Email Marketing Accessible for Everyone What is Headless? Don’t Lose Your Head, SMBs: It’s a Good Thing Architect the Future UI: Slack as Your Agentic Surface Point of Sale Innovations to Modernize the Shopper Experience Governing the Agentic Enterprise at Scale with MuleSoft Omni Gateway How to Cut Service Time with Case Routing Automation 5 Tips for Marketers to Get Started with Salesforce Flow No One is Vibe Coding Trade Promotion Management 7th Edition State of Sales Report: 3 Growth Trends for Startups and SMBs How the Architect Vista Brought Architectural Thinking to Life at TDX 2026 5 Steps to Develop an Architect Mindset With AI Why AI Isn’t Replacing Developers, It’s Empowering Them 10 Ways to Make Your AI Agent a Better Communicator AI in Design 2025: What Real Use Taught Us. How Salesforce Personalization Learns Which Offers Drive Revenue The 4-Step Guide to Salesforce Agent and Application Development Get Ready for Connections 2026: Top Sessions and New Reveals 8 Ways AI Agents Are Evolving in 2026 4 Principles to Make the Right Salesforce UI Decisions Apprentice Journey Shines a Light on Talent Pathways at Salesforce Agent Script: The Control Plane for Agentic Decisions Scaling the Agentforce Life Sciences Ecosystem to Drive the Future of Pharma and MedTech Unlocking Unstructured Data: Building AI-Powered Support Triage with Data 360 Asking For a Friend: What Are Rich Communication Services (RCS)? 5 Slack Shortcuts For Small Teams 195% ROI In Field Service? Here’s How They Did It Submit Your Architect Session: The Dreamforce 2026 Call for Participation Is Open What Is an AI Assistant for Small Business? B2C Commerce April release: Transforming the B2C developer experience with agents Meet Your 24/7 Prospecting Partner — And 5 More Stand-Out Features In Our April Release 11+ Small Business YouTube Channels You Need to Follow Today Stop Treating Disputes Management Like IT Tickets Limitless Service: A New Operating Model for Growth in the Agentic Era What is Transactional Reconciliation in Email and SMS Marketing? How to Prepare for National Small Business Week (2026) How to Design a High-Scale Multi-Cloud Incident Journey 10 Ways An AI CRM Can Amplify Your Startup Vibe Code Better Agents with Agentforce Free vs. Paid CRM: Which is Right for Your Business? Salesforce Customer Success Awards 2026: Lead Era of the Agentic Enterprise 12 Free Webinars for Small Business Owners (2026) The Agentforce Life Sciences Consultant Certification Maximize Growth: The Power of Partnerships for SMBs Salesforce AI Research at ICLR 2026 Data Protection For Small Business: How To Safeguard Yourself Beyond 100K Tokens: Evaluating AI Agents in Long-Context Software Engineering Top 32 Small Business Tools To Try Today 6 New Innovations Redefining Salesforce Development How to Win the Battle for Attention in the Agentic Email Inbox Mastering the Eisenhower Matrix: Prioritize Like a Pro 10 Signs It’s Time To Upgrade Your CRM and How To Get Started (2026) Celebrating 10 Years of Financial Services Innovation How SMBs Can Gain An Edge With Agentic AI: Key Trends From Our Marketing Report How MAN Truck & Bus is Shaping the Future of Sales with Salesfive Introducing the Future of Salesforce Data Protection: Backup & Recover Next Stop Making AI Slop – Build a Foundation for Authentic AI Content 5 Tips to Help Marketers Navigate AI Email Summaries Data Sharing: Is it Safe? Is it Secure? Everything You Need to Know Small Business Week Readiness: 4 Things to Do Before May (2026) How Salesforce Employees Make an Impact During Earth Month AI Agents Are Advancing Rapidly… Is Your Testing Strategy Keeping Up? What Is Microproductivity and Why Is It Helping So Many Teams? TDX 2026 Roundup: Agentforce Edition 3 Ways Salesforce Connections Has Boosted My Career AI Agents Don’t Just Answer‌ — ‌They Act. Do You Have a Governance Strategy? The Future of MedTech Field Execution is Agentic 5 Steps to Prepare Your Data For an AI CRM From Break/Fix to Profit Engine: Aftermarket Service for Robot OEMs Building Trusted Human-Agent Collaboration: A Practical Framework ISV Strategy for the Salesforce Summer ’26 Release 5 Email Marketing Tips for Small Business Commerce Shops Should You Give Your AI Agent a Human Name? Creating Pathways into AI for People with Disabilities The Organized Chaos of Upfronts: 3 Hurdles Impacting Your Yield Trying to Scale Beyond ‘One-Off’ AI Tasks? You’re Probably Using the Wrong Interface What is Cost Per Lead (CPL)? The Case for Unified CCaaS and CRM — And Why the Data Makes It Clear In the New Era of AI, You Need to Win Over Both Humans and Agents What Is SPIN Selling? A Way to Build Trust With Your Customers G2 Crowns Salesforce as Best Financial Services Software Hidden Insights: The Guide to Tableau For Small Business Owners
Why Technical Accuracy is the Wrong Metric for Agent Success
Denise Marti · 2026-05-13 · via Salesforce

Imagine you’re looking for a specific bakery in a maze of narrow city streets. Your maps app is technically flawless, giving real-time directions and rerouting perfectly. But you’re still lost, squinting at your phone and second-guessing every corner. Your cognitive load is through the roof. Eventually, you close the app and ask a local for help.

This is workflow breakage in action. The system was correct, but it failed to meet you where you were. It’s a reminder that an agent can be right and still fail the user.

We see this same friction in our technical logs. Accuracy is trending at 92%. Task completion is at 89%. Engagement is a steady 94%. By every technical metric, the agent is a success. But then, usage drops. Users return to their old, manual workflows. When asked why, they don’t say the agent was wrong. They say, “It just doesn’t work for me.”

This is the silent adoption killer: the gap between system success and user success.

To build agents people actually use, we have to stop asking if the agent is accurate and start asking if it’s worth using. This shift requires a new kind of infrastructure that treats quality as a rigorous release gate centered on human outcomes.

Here’s what we’ll cover:

Spot the gap between system metrics and lived experience
Diagnose failures using the three severity tiers
Evaluate quality at scale with heuristics
Turn quality into a release gate, not an afterthought
Putting these principles into practice

Spot the gap between system metrics and lived experience

The metrics problem in AI is subtle. Accuracy, task completion, and engagement are the standard markers of a healthy system. But telemetry and lived experience often tell two different stories. A log might show a successfully completed task with no errors, while the user feels they have wasted their time.

System success and user success aren’t the same. When we analyzed transcripts of interactions that appeared successful in the logs, we found that nearly one in seven actually failed the user. These are the phantom successes. While the data flags a win, the user experiences a loss and quietly abandons the agent and reverts to old workflows because the interaction was too slow, too wordy, or too difficult to verify.

To find these gaps, you have to look beyond the numbers and identify where the agent is technically right and practically wrong. This happens when the system understands the task but fails to understand the human context, and closing this gap starts with a simple shift in perspective. Instead of asking if the agent is accurate, ask if it provides enough value to replace the user’s current way of working.

Back to the top

Take a Deeper Dive

Diagnose failures using the three severity tiers

When a user says an agent doesn’t work, that feedback is often too broad to be actionable. To fix the experience, you have to go deeper into the root cause. At Salesforce, we use a severity tiering system to categorize failures based on how they impact the user’s ability to move forward. This allows us to prioritize the most damaging issues first while keeping an eye on the friction that causes long-term abandonment.

P0: Trust-breaking failures
These are critical errors where the agent provides a wrong answer or entirely fails to respond. A common example is a legal agent giving a confident but outdated answer about contract terms, and because the legal risk increases with every single wrong answer, the user loses faith in the entire system. P0 failures are loud and must be addressed immediately to maintain the security and trust foundation of the product.

P1: Intent failures
In these cases, the agent might be technically accurate, but it isn’t useful enough to help the user complete their task. This happens when an agent ignores inputs, misunderstands intent, or fails to clarify a vague request. These failures block adoption because they force the user to do the mental heavy lifting that the AI was supposed to handle.

P2: Friction and overhead
P2 failures are the most subtle because the agent provides the correct information but delivers it in a way that creates friction. The answer is accurate, but the format or timing is so inconvenient that the user still has to exert extra effort to make the data useful. This looks like returning a long, unformatted block of text on a mobile device or providing a link instead of a summary. While these don’t break the system, they do increase cognitive load and reduce productivity. Over time, these small frustrations accumulate and lead users back to their old workflows.

Back to the top

Evaluate quality at scale with heuristics

To define what a high-quality experience looks like, we use heuristics. In design, a heuristic is a practical rule of thumb used to evaluate the usability of an interface. Think of these as experience checks that help design and engineering teams catch failures before they reach the user. By evaluating transcripts against these 11 benchmarks, you can pinpoint exactly why an interaction feels off, even when the system logs claim success.

These heuristics are grouped by the severity tiers they typically trigger:

P0: Trust-breaking checks

  • Factual & Reliable: Are the agent’s answers grounded in actual context or data?

P1: Understanding checks 

  • Effective: Does the agent actually complete the task, not just respond correctly?
  • Memory & UI Context: Does the agent use information that’s already available in the chat or UI?
  • Teachable: Does the agent adjust when corrected or redirected?
  • Responsive: Does the agent ask for clarification when the request is ambiguous?
  • Trusted: Does the agent stay within safe and appropriate limits?
  • Decisive: Does the agent give a clear next step, rather than just explaining and leaving the user hanging?

P2: Friction and overhead checks

  • Consistent: Does the agent maintain the same tone, terms, and behavior across turns?
  • Helpful: Does the agent’s response actually move the user toward completing their task?
  • Conversational: Does the agent use language that feels natural and easy to work with?
  • Approachable: Does interacting with the agent feel effortless?

Using these heuristics manually for every conversation is impossible at scale. To operationalize this, we’ve moved toward an “LLM-as-judge” pipeline for evaluations. By running conversations against it, we can reduce the time it takes to review a session from 30 minutes to just a few seconds. This allows us to maintain a high bar for quality across thousands of interactions without removing the human designer or builder from the strategic loop.

Back to the top

Turn quality into a release gate, not an afterthought

Defining quality is only half the battle. To truly protect the user experience, these heuristics must become part of your infrastructure. At Salesforce, we treat severity tiers as release gates:

  • P0 blocks deployment. This is the dealbreaker. If we see a factual error or safety violation at any stage, we stop. There is no scaling through a hallucination and if the agent is live, we pull it.
  • P1 blocks scale. This is the misunderstanding. If a pilot reveals that the agent is leaving users hanging, we do not open the doors to more people. We refuse to scale a misunderstanding.
  • P2 blocks adoption. This is the overhead. Users rarely complain about friction, they simply stop coming back. This is where adoption dies quietly.

P0 failures are loud and easy to find, but P1 and P2 failures are the ghost failures that cause a product to fail slowly. By encoding these as gates, we ensure we are shipping an agent users will actually keep using, rather than just accurate code.

By making quality a gate, you ensure that the responsibility for the user experience is shared across design, engineering, and product management. This moves the conversation beyond whether the code works to whether the experience is ready for a human.

Operationalizing this process goes beyond checking boxes. It builds a system that values time and trust. While automation tools like LLM-based evaluations help us maintain this high bar at scale, the ultimate judgment remains human. Instead of removing the designer or builder, we provide them with better data to make strategic decisions on how to optimize the experience to help humans reach an outcome.

Back to the top

Putting these principles into practice

If you’re ready to apply these principles to your own agentic workflows, start small. You can begin closing the gap between your logs and the lived experience of your users without a complex pipeline.

Take these three steps today:

  1. Review 50 “successful” sessions. Look past the green checkmarks in your logs to read the actual transcripts.
  2. Tag the friction. Use the 11 heuristics to identify where the agent understood the task but failed the human.
  3. Ask the key question. Would you choose this AI workflow over the way you do the task today?

If the answer is anything less than a clear “yes,” you’ve found your starting point for improvement. By shifting your focus from technical accuracy to human value, you move beyond building tools that work and begin building agents that people truly trust.

Back to the top