惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

P
Proofpoint News Feed
WordPress大学
WordPress大学
S
Schneier on Security
Recent Commits to openclaw:main
Recent Commits to openclaw:main
AWS News Blog
AWS News Blog
The Cloudflare Blog
酷 壳 – CoolShell
酷 壳 – CoolShell
博客园 - 叶小钗
NISL@THU
NISL@THU
T
Tor Project blog
L
Lohrmann on Cybersecurity
D
Darknet – Hacking Tools, Hacker News & Cyber Security
博客园 - 聂微东
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
月光博客
月光博客
Microsoft Azure Blog
Microsoft Azure Blog
P
Proofpoint News Feed
G
GRAHAM CLULEY
博客园_首页
K
Kaspersky official blog
GbyAI
GbyAI
P
Privacy & Cybersecurity Law Blog
freeCodeCamp Programming Tutorials: Python, JavaScript, Git & More
AI
AI
爱范儿
爱范儿
Cloudbric
Cloudbric
MongoDB | Blog
MongoDB | Blog
Martin Fowler
Martin Fowler
aimingoo的专栏
aimingoo的专栏
I
InfoQ
腾讯CDC
O
OpenAI News
F
Full Disclosure
P
Privacy International News Feed
K
KPMG report finds enterprise disconnect between AI and its ROI | CIO
Webroot Blog
Webroot Blog
Forbes - Security
Forbes - Security
MyScale Blog
MyScale Blog
L
LangChain Blog
H
Help Net Security
C
CERT Recently Published Vulnerability Notes
C
Cisco Blogs
人人都是产品经理
人人都是产品经理
S
Security @ Cisco Blogs
T
Tenable Blog
H
Hackread – Cybersecurity News, Data Breaches, AI and More
N
News and Events Feed by Topic
博客园 - 三生石上(FineUI控件)
Attack and Defense Labs
Attack and Defense Labs
Apple Machine Learning Research
Apple Machine Learning Research

Amplitude

Beyond the Rate: Retail Banking's New Competitive Front How NS Prevented €1.8M in Revenue Loss Through Experimentation Go from Product Launch to Insight to Action in Minutes What Makes a Good vs Bad North Star Metric The Role of Feature Management in Successful Product Development Cohort Retention Analysis: Reduce Churn Using Customer Data 7 Steps to Measuring the Success of a Feature 14 Best Product Management Tools for 2026 (Plus Tips from Senior PMs) The Definitive Guide to Behavioral Cohorting Meet the Winners of the 2026 Amplitude AI Impact Awards Beyond Last-Touch Attribution: Find Out Which Interactions Really Matter Agent Connectors Are Better Together Agents That Act on What Actually Happened How Square Used Amplitude to Enhance the Seller Experience and Power Growth Migrating Analytics Platforms Without The Chaos Wanted Lab Grows Sign-Ups by 150% & Builds Experimentation Culture How to Balance Inference Cost and User Experience for Agents Introducing Zoning Insights: Web Intelligence at a Glance Five best practices for getting started with AI agents 24 Quarters at #1. Here’s What’s Next. How We Built a Product That Tells Us What To Build Next: Inside Amplitude Wave Looking Beyond Campaign Metrics: 7 Marketing Success Stories AI Evals for Product Managers: A Beginner’s Guide to Getting Started The Builder Skills Library Introducing Agent Connectors in Amplitude Understand How AI Thinks, Get Better Results How We Redesigned Amplitude Docs for Agents and Made Everyone an Author AI Broke Your Experimentation Program. Here’s How to Fix It. Every Stuck User Is a Support Ticket Waiting to Happen Tracing the Sale: Connect Behavior to Conversions with Persisted Properties Building CLI Agents: It’s What You Don’t Give Them That Counts Three Tips for Better Prompts in Amplitude Global Agent How AI Took the Data Analyst’s Job, and Created a Better One Default Prompts Are Tanking Your Agent’s Retention Optimizing Core Web Vitals with Amplitude’s Global Agent Don’t Ask Global Agent Anything, Ask These Three Things How We Built a Design Agent at Amplitude with Claude Managed Agents and Cloudflare The Problem with Chasing Churn How Hostinger Achieved a 20%+ Conversion Lift Through Experimentation How STAGE Streams Smarter by Putting Data at the Center Building the Validation Stack for AI Product Development Making AI Analytics Safe for Financial Services Teams Amplitude Heatmaps Update: More Reliable Screenshots and Accurate Placement Most Teams Ship Agent Personalities by Accident. We Didn’t. What I Learned Pointing a Ralph Loop at My Product for a Week How Mercado Libre Scales Decision Making with AI Claude Cowork for PMs: 5 Playbooks to Get Started How ACKO Drove 13% More Conversions & 50% Drop in Calls with GenAI Agents Just Made Your Feature Launch Channel Smarter Homegrown FinOps Tools: How AI “Build” Beat “Buy” for Us in <1 Year Introducing The Amplitude Quickstart Series Rebuilding Session Replay’s Delivery Layer to Be Lighter on Your Page The Eval Signal That Predicts 3x Agent Retention Agents Write Code. Fixing It Is Still On You. Amplitude and Statsig Partnership 5 Agent Skills to Automate Your Weekly Product Review Amplitude Plug and Play: New AI Plugin in Claude and Cursor Marketplaces Introducing Amplitude Wizard CLI: Set Up Amplitude from Your Codebase Making AI Search Count (and Convert) How VEED Evolved Its AI Search Strategy What’s New with Amplitude Agents Effortless Support at Scale: Making Human Support More Human AI Week 2026: Upleveling All Together Amplitude AI Builders: Paul Hultgren Chats about AI Assistant Dashboard Dread to AI-Driven Decisions: How Tira Rebuilt Its Analytics Workflow Your Product Deserves a Better Support Agent How Cisco Systems Accelerated Adoption by 20% Through Data Innovation
Putting A Number On AI Quality
Fionn O'Raghallaigh · 2026-06-29 · via Amplitude

The Economist Group is best known for its newspaper that has existed since 1843. In 1946, the Economist Intelligence Unit was set up to answer questions Economist readers were asking. Today EIU helps businesses, financial firms and governments to understand how the world is changing and how that creates opportunities to be seized and risks to be managed. EIU’s team of analysts cover nearly every country in the world with analysis, forecasts and indicators to back decisions. They create a lot of content. And data. It can be hard to navigate.

Enter Lens. In March, we added the multi-turn AI research assistant Lens to Viewpoint, EIU’s website. It allows our analyst, strategist and risk customers to ask about the fiscal outlook or compare political risk across the five biggest economies in South America and get an answer sourced entirely from EIU content. Getting Lens to a good enough quality was hard work. We worked with our analysts on evaluations. We went back and forth until we were happy. By the time we launched we knew Lens was good. But a feature in the wild is a different beast and pinpointing the areas to improve was the next challenge.

Reading the room

Product analytics showed us what users did. It couldn't tell us whether the AI answer was any good. A user who asked three questions in a session might have been getting excellent answers and going deeper. Or they might have been rephrasing the same question because the first two answers were wrong. The data looked the same either way. We opened sessions and read them. But not fast enough to act on it. First impressions matter.

From samples to signal

Lens sessions landed into Amplitude with signals attached out of the box: intent classification, outcome tracking, quality scores. Getting there took one engineer and some back and forth, but once it was running, reading a trace stopped being archaeology. I could open a session, see each turn, see what the agent retrieved and how it responded, and see how it scored. Before that, our data insights lead, our platform engineers and I were drawing conclusions from handfuls of sessions. A few examples, a hunch. Now we're looking at the same evidence across all of them.

Filtering by behavior

The capability I lean on most is addictive. Pull every session matching a behavioral signal and drill in.

Early on, the request complexity classification surfaced a concern about how Lens handled ambiguous questions. I filtered to those sessions, found two areas that worried me, and we now track both programmatically. That whole loop, from hunch to confirmed pattern to standing metric, took an afternoon.

It also catches things that look like problems and are not. Around half of our session outcomes were labelled "Clarification Requested," which initially read as a defect. Drilling in showed the agent was ending answers with a follow-up question, a deliberate design choice. But when we looked at what users did next, the pattern was clear enough. They weren't going deeper, they were going elsewhere. We tweaked the prompt. Without the ability to interrogate the label, we might have spent a sprint fixing behavior that was working as intended.

A scorecard instead of a vibe

We all know AI outputs are inherently non-deterministic, which is a long way of saying that viewing a handful of sessions and calling it quality assurance doesn't scale. We wanted a number.

Amplitude's standard signals gave us a floor on day one. Custom evals let us go further, scoring sessions against our own definition of good, including rubrics built from ground-truth Q&A pairs our analysts wrote.

By Amplitude's measure, Lens now holds a 96.9% task success rate, and weekly task failures fell 84% as we worked through the issues the data surfaced. Those gains came from ordinary engineering. The measurement is what tells us where to point it.

Evals alone tell you whether an answer was correct. Quality signals connected to real product usage tell you whether the product is getting better at its job.

What comes next

Agent interactions in Amplitude arrive as decomposed events, the same shape as any other product event. That opens the door we care most about, connecting Lens quality to engagement across Viewpoint as a whole. Does a strong Lens experience deepen how teams use the platform? As we add structured data and chart visualisation to Lens, and as other parts of The Economist Group explore whether the architecture fits their own products, that connected view is how we will judge the work.

None of this is about watching individual users. It is about holding our AI products to the same standard of evidence we hold our analysis to. Our clients pay us for rigor. The tools we use to improve their experience should have some too.

Ready to use AI to transform your product?

Amplitude Agents help you understand your users more easily than ever.

Get started now