惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - 聂微东
Forbes - Security
Forbes - Security
IT之家
IT之家
P
Privacy International News Feed
宝玉的分享
宝玉的分享
小众软件
小众软件
Google DeepMind News
Google DeepMind News
美团技术团队
G
GRAHAM CLULEY
T
Tor Project blog
Recorded Future
Recorded Future
I
Intezer
C
Cyber Attacks, Cyber Crime and Cyber Security
D
Darknet – Hacking Tools, Hacker News & Cyber Security
The Hacker News
The Hacker News
Hugging Face - Blog
Hugging Face - Blog
A
About on SuperTechFans
Scott Helme
Scott Helme
WordPress大学
WordPress大学
F
Full Disclosure
D
Docker
G
Google Developers Blog
C
CXSECURITY Database RSS Feed - CXSecurity.com
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Cyberwarzone
Cyberwarzone
The Last Watchdog
The Last Watchdog
V
V2EX
www.infosecurity-magazine.com
www.infosecurity-magazine.com
NISL@THU
NISL@THU
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
Security Latest
Security Latest
Recent Commits to openclaw:main
Recent Commits to openclaw:main
Recent Announcements
Recent Announcements
P
Palo Alto Networks Blog
L
LINUX DO - 热门话题
V
Visual Studio Blog
B
Blog RSS Feed
Microsoft Security Blog
Microsoft Security Blog
博客园 - 叶小钗
N
Netflix TechBlog - Medium
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
量子位
腾讯CDC
H
Heimdal Security Blog
博客园 - 【当耐特】
Simon Willison's Weblog
Simon Willison's Weblog
P
Privacy & Cybersecurity Law Blog
S
Securelist
Vercel News
Vercel News
J
Java Code Geeks

Amplitude

Beyond the Rate: Retail Banking's New Competitive Front How NS Prevented €1.8M in Revenue Loss Through Experimentation Go from Product Launch to Insight to Action in Minutes What Makes a Good vs Bad North Star Metric The Role of Feature Management in Successful Product Development Cohort Retention Analysis: Reduce Churn Using Customer Data 7 Steps to Measuring the Success of a Feature 14 Best Product Management Tools for 2026 (Plus Tips from Senior PMs) The Definitive Guide to Behavioral Cohorting Meet the Winners of the 2026 Amplitude AI Impact Awards Beyond Last-Touch Attribution: Find Out Which Interactions Really Matter Agent Connectors Are Better Together Agents That Act on What Actually Happened How Square Used Amplitude to Enhance the Seller Experience and Power Growth Migrating Analytics Platforms Without The Chaos Wanted Lab Grows Sign-Ups by 150% & Builds Experimentation Culture How to Balance Inference Cost and User Experience for Agents Introducing Zoning Insights: Web Intelligence at a Glance Five best practices for getting started with AI agents 24 Quarters at #1. Here’s What’s Next. How We Built a Product That Tells Us What To Build Next: Inside Amplitude Wave Looking Beyond Campaign Metrics: 7 Marketing Success Stories AI Evals for Product Managers: A Beginner’s Guide to Getting Started The Builder Skills Library Introducing Agent Connectors in Amplitude Understand How AI Thinks, Get Better Results How We Redesigned Amplitude Docs for Agents and Made Everyone an Author AI Broke Your Experimentation Program. Here’s How to Fix It. Every Stuck User Is a Support Ticket Waiting to Happen Tracing the Sale: Connect Behavior to Conversions with Persisted Properties Building CLI Agents: It’s What You Don’t Give Them That Counts Three Tips for Better Prompts in Amplitude Global Agent How AI Took the Data Analyst’s Job, and Created a Better One Default Prompts Are Tanking Your Agent’s Retention Optimizing Core Web Vitals with Amplitude’s Global Agent Don’t Ask Global Agent Anything, Ask These Three Things How We Built a Design Agent at Amplitude with Claude Managed Agents and Cloudflare The Problem with Chasing Churn How Hostinger Achieved a 20%+ Conversion Lift Through Experimentation How STAGE Streams Smarter by Putting Data at the Center Building the Validation Stack for AI Product Development Making AI Analytics Safe for Financial Services Teams Amplitude Heatmaps Update: More Reliable Screenshots and Accurate Placement Most Teams Ship Agent Personalities by Accident. We Didn’t. What I Learned Pointing a Ralph Loop at My Product for a Week How Mercado Libre Scales Decision Making with AI Claude Cowork for PMs: 5 Playbooks to Get Started How ACKO Drove 13% More Conversions & 50% Drop in Calls with GenAI Agents Just Made Your Feature Launch Channel Smarter Homegrown FinOps Tools: How AI “Build” Beat “Buy” for Us in <1 Year Introducing The Amplitude Quickstart Series Rebuilding Session Replay’s Delivery Layer to Be Lighter on Your Page The Eval Signal That Predicts 3x Agent Retention Agents Write Code. Fixing It Is Still On You. Amplitude and Statsig Partnership 5 Agent Skills to Automate Your Weekly Product Review Amplitude Plug and Play: New AI Plugin in Claude and Cursor Marketplaces Introducing Amplitude Wizard CLI: Set Up Amplitude from Your Codebase Making AI Search Count (and Convert) How VEED Evolved Its AI Search Strategy What’s New with Amplitude Agents Effortless Support at Scale: Making Human Support More Human AI Week 2026: Upleveling All Together Amplitude AI Builders: Paul Hultgren Chats about AI Assistant Dashboard Dread to AI-Driven Decisions: How Tira Rebuilt Its Analytics Workflow Your Product Deserves a Better Support Agent How Cisco Systems Accelerated Adoption by 20% Through Data Innovation
Putting A Number On AI Quality
Fionn O'Raghallaigh · 2026-06-29 · via Amplitude

The Economist Group is best known for its newspaper that has existed since 1843. In 1946, the Economist Intelligence Unit was set up to answer questions Economist readers were asking. Today EIU helps businesses, financial firms and governments to understand how the world is changing and how that creates opportunities to be seized and risks to be managed. EIU’s team of analysts cover nearly every country in the world with analysis, forecasts and indicators to back decisions. They create a lot of content. And data. It can be hard to navigate.

Enter Lens. In March, we added the multi-turn AI research assistant Lens to Viewpoint, EIU’s website. It allows our analyst, strategist and risk customers to ask about the fiscal outlook or compare political risk across the five biggest economies in South America and get an answer sourced entirely from EIU content. Getting Lens to a good enough quality was hard work. We worked with our analysts on evaluations. We went back and forth until we were happy. By the time we launched we knew Lens was good. But a feature in the wild is a different beast and pinpointing the areas to improve was the next challenge.

Reading the room

Product analytics showed us what users did. It couldn't tell us whether the AI answer was any good. A user who asked three questions in a session might have been getting excellent answers and going deeper. Or they might have been rephrasing the same question because the first two answers were wrong. The data looked the same either way. We opened sessions and read them. But not fast enough to act on it. First impressions matter.

From samples to signal

Lens sessions landed into Amplitude with signals attached out of the box: intent classification, outcome tracking, quality scores. Getting there took one engineer and some back and forth, but once it was running, reading a trace stopped being archaeology. I could open a session, see each turn, see what the agent retrieved and how it responded, and see how it scored. Before that, our data insights lead, our platform engineers and I were drawing conclusions from handfuls of sessions. A few examples, a hunch. Now we're looking at the same evidence across all of them.

Filtering by behavior

The capability I lean on most is addictive. Pull every session matching a behavioral signal and drill in.

Early on, the request complexity classification surfaced a concern about how Lens handled ambiguous questions. I filtered to those sessions, found two areas that worried me, and we now track both programmatically. That whole loop, from hunch to confirmed pattern to standing metric, took an afternoon.

It also catches things that look like problems and are not. Around half of our session outcomes were labelled "Clarification Requested," which initially read as a defect. Drilling in showed the agent was ending answers with a follow-up question, a deliberate design choice. But when we looked at what users did next, the pattern was clear enough. They weren't going deeper, they were going elsewhere. We tweaked the prompt. Without the ability to interrogate the label, we might have spent a sprint fixing behavior that was working as intended.

A scorecard instead of a vibe

We all know AI outputs are inherently non-deterministic, which is a long way of saying that viewing a handful of sessions and calling it quality assurance doesn't scale. We wanted a number.

Amplitude's standard signals gave us a floor on day one. Custom evals let us go further, scoring sessions against our own definition of good, including rubrics built from ground-truth Q&A pairs our analysts wrote.

By Amplitude's measure, Lens now holds a 96.9% task success rate, and weekly task failures fell 84% as we worked through the issues the data surfaced. Those gains came from ordinary engineering. The measurement is what tells us where to point it.

Evals alone tell you whether an answer was correct. Quality signals connected to real product usage tell you whether the product is getting better at its job.

What comes next

Agent interactions in Amplitude arrive as decomposed events, the same shape as any other product event. That opens the door we care most about, connecting Lens quality to engagement across Viewpoint as a whole. Does a strong Lens experience deepen how teams use the platform? As we add structured data and chart visualisation to Lens, and as other parts of The Economist Group explore whether the architecture fits their own products, that connected view is how we will judge the work.

None of this is about watching individual users. It is about holding our AI products to the same standard of evidence we hold our analysis to. Our clients pay us for rigor. The tools we use to improve their experience should have some too.

Ready to use AI to transform your product?

Amplitude Agents help you understand your users more easily than ever.

Get started now