惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

爱范儿
爱范儿
Forbes - Security
Forbes - Security
Help Net Security
Help Net Security
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
Hacker News: Ask HN
Hacker News: Ask HN
TaoSecurity Blog
TaoSecurity Blog
IT之家
IT之家
Microsoft Azure Blog
Microsoft Azure Blog
云风的 BLOG
云风的 BLOG
博客园 - 司徒正美
B
Blog
阮一峰的网络日志
阮一峰的网络日志
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Google DeepMind News
Google DeepMind News
Microsoft Security Blog
Microsoft Security Blog
The Register - Security
The Register - Security
美团技术团队
C
CERT Recently Published Vulnerability Notes
I
Intezer
C
Cybersecurity and Infrastructure Security Agency CISA
Google Online Security Blog
Google Online Security Blog
B
Blog RSS Feed
PCI Perspectives
PCI Perspectives
C
Cyber Attacks, Cyber Crime and Cyber Security
Threat Intelligence Blog | Flashpoint
Threat Intelligence Blog | Flashpoint
MyScale Blog
MyScale Blog
S
Securelist
Recorded Future
Recorded Future
Know Your Adversary
Know Your Adversary
Security Archives - TechRepublic
Security Archives - TechRepublic
M
MIT News - Artificial intelligence
C
Check Point Blog
T
Threat Research - Cisco Blogs
博客园 - Franky
P
Proofpoint News Feed
人人都是产品经理
人人都是产品经理
U
Unit 42
F
Fortinet All Blogs
S
Security @ Cisco Blogs
The GitHub Blog
The GitHub Blog
Apple Machine Learning Research
Apple Machine Learning Research
MongoDB | Blog
MongoDB | Blog
The Hacker News
The Hacker News
酷 壳 – CoolShell
酷 壳 – CoolShell
F
Full Disclosure
aimingoo的专栏
aimingoo的专栏
大猫的无限游戏
大猫的无限游戏
D
Docker
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
L
LINUX DO - 热门话题

Amplitude

Beyond the Rate: Retail Banking's New Competitive Front How NS Prevented €1.8M in Revenue Loss Through Experimentation Go from Product Launch to Insight to Action in Minutes What Makes a Good vs Bad North Star Metric The Role of Feature Management in Successful Product Development Cohort Retention Analysis: Reduce Churn Using Customer Data 7 Steps to Measuring the Success of a Feature 14 Best Product Management Tools for 2026 (Plus Tips from Senior PMs) The Definitive Guide to Behavioral Cohorting Putting A Number On AI Quality Meet the Winners of the 2026 Amplitude AI Impact Awards Beyond Last-Touch Attribution: Find Out Which Interactions Really Matter Agent Connectors Are Better Together Agents That Act on What Actually Happened How Square Used Amplitude to Enhance the Seller Experience and Power Growth Migrating Analytics Platforms Without The Chaos Wanted Lab Grows Sign-Ups by 150% & Builds Experimentation Culture How to Balance Inference Cost and User Experience for Agents Introducing Zoning Insights: Web Intelligence at a Glance Five best practices for getting started with AI agents 24 Quarters at #1. Here’s What’s Next. How We Built a Product That Tells Us What To Build Next: Inside Amplitude Wave Looking Beyond Campaign Metrics: 7 Marketing Success Stories AI Evals for Product Managers: A Beginner’s Guide to Getting Started The Builder Skills Library Introducing Agent Connectors in Amplitude Understand How AI Thinks, Get Better Results How We Redesigned Amplitude Docs for Agents and Made Everyone an Author AI Broke Your Experimentation Program. Here’s How to Fix It. Every Stuck User Is a Support Ticket Waiting to Happen Tracing the Sale: Connect Behavior to Conversions with Persisted Properties Building CLI Agents: It’s What You Don’t Give Them That Counts Three Tips for Better Prompts in Amplitude Global Agent How AI Took the Data Analyst’s Job, and Created a Better One Default Prompts Are Tanking Your Agent’s Retention Optimizing Core Web Vitals with Amplitude’s Global Agent Don’t Ask Global Agent Anything, Ask These Three Things How We Built a Design Agent at Amplitude with Claude Managed Agents and Cloudflare The Problem with Chasing Churn How Hostinger Achieved a 20%+ Conversion Lift Through Experimentation How STAGE Streams Smarter by Putting Data at the Center Making AI Analytics Safe for Financial Services Teams Amplitude Heatmaps Update: More Reliable Screenshots and Accurate Placement Most Teams Ship Agent Personalities by Accident. We Didn’t. What I Learned Pointing a Ralph Loop at My Product for a Week How Mercado Libre Scales Decision Making with AI Claude Cowork for PMs: 5 Playbooks to Get Started How ACKO Drove 13% More Conversions & 50% Drop in Calls with GenAI Agents Just Made Your Feature Launch Channel Smarter Homegrown FinOps Tools: How AI “Build” Beat “Buy” for Us in <1 Year Introducing The Amplitude Quickstart Series Rebuilding Session Replay’s Delivery Layer to Be Lighter on Your Page The Eval Signal That Predicts 3x Agent Retention Agents Write Code. Fixing It Is Still On You. Amplitude and Statsig Partnership 5 Agent Skills to Automate Your Weekly Product Review Amplitude Plug and Play: New AI Plugin in Claude and Cursor Marketplaces Introducing Amplitude Wizard CLI: Set Up Amplitude from Your Codebase Making AI Search Count (and Convert) How VEED Evolved Its AI Search Strategy What’s New with Amplitude Agents Effortless Support at Scale: Making Human Support More Human AI Week 2026: Upleveling All Together Amplitude AI Builders: Paul Hultgren Chats about AI Assistant Dashboard Dread to AI-Driven Decisions: How Tira Rebuilt Its Analytics Workflow Your Product Deserves a Better Support Agent How Cisco Systems Accelerated Adoption by 20% Through Data Innovation
Building the Validation Stack for AI Product Development
Eric Metelka · 2026-05-14 · via Amplitude

A lot has happened in a year in the world of experimentation. A year ago my company, Eppo, which offered warehouse-native experimentation, was bought by Datadog. A year later and my company, Amplitude, is welcoming Statsig, its customers, and its brand to its platform.

The team at Statsig built a strong product. They recognized early that engineers had a need for better tools for rollouts and to understand the value of what they were shipping. They developed a builder-first approach to feature flags, experiments, metrics, and rollout controls, that clearly resonated in the market.

At Amplitude, we believe, just like Statsig does, that experimentation is core infrastructure and a foundational part of how products get built. This is even more important in an AI world. Partnering with Statsig is an opportunity to accelerate a shared vision for the future of product development.

How building products has changed

The bottleneck in product development has moved. It used to be writing code. But now, with a majority of developers using AI coding tools, code generation is only getting faster. PMs write code while designers build and ship full UX flows. The code barrier to getting something built has fully collapsed.

But the gap between shipping a new feature and knowing that it’s good for users has actually gotten wider. Teams are shipping faster than ever, and while the volume of changes going out the door has exploded, the infrastructure to validate those changes hasn't kept pace. Existing bottlenecks in the experimentation process compound when shipping velocity increases.

With non-deterministic products like LLMs, it has become even harder to determine if you’re shipping the right thing. Whether you’re working on a chatbot, a recommendation engine, or something else, non-deterministic outputs give you a different response every time. Unit tests can’t give you the confidence you need. Experimentation can.

Additionally, the people building these products aren't necessarily the same people who ran experiments five years ago. The number of people capable of writing code or shipping new features has exploded, but the number who deeply understand how to validate those features has not. Modern experimentation tooling needs to support a much broader range of AI builders.

Building the validation stack for AI product development

Internally, we’re thinking about what the “2.0” of experimentation needs to become.

Version 1.0 is a known loop: ship with feature flags, measure impact with experiments, understand usage with analytics. That loop still works. But teams building AI products need another layer of validation and rigor. You need offline evaluation, live experimentation, and continuous monitoring working together.

The starting point for Experimentation 2.0 is offline evals. Instead of manually checking a few outputs and hoping for the best, you run prompts and models through thousands of labeled test cases before anything reaches production. The goal is to catch regressions early and avoid surprises in production.

Say you’re running an AI support ticket classifier. You have a prompt that triages tickets to billing, technical support, or sales. You update the prompt to handle edge cases better. Is the new version actually better? Offline evals let you run both versions against a labeled dataset of a thousand tickets, score them against graders (including LLM-as-a-Judge for cases where string matching doesn’t work), and see exactly where the new version wins and where it regresses. You iterate on this loop rapidly before any user sees the change.

From there, you move to progressive rollout with gradual deployment and instant rollbacks, tied to service metrics, business KPIs, and LLM-specific observability signals. If latency spikes or error rates climb, the system responds before the issue spreads.

Then onto online experimentation. A/B tests on live traffic with statistical confidence. Shadow-mode evals that grade model output against production scenarios without exposing users to risk. Every rollout should measure impact, not just reduce risk.

Running through this entire 2.0 loop is LLM observability, which gives you real-time logging, monitoring, and anomaly alerting in a single view alongside business metrics and user engagement. When something goes wrong with your AI product, you shouldn’t need four dashboards to figure out where.

Amplitude + Statsig will get there faster

Statsig and Amplitude were already building toward the same future; one where flags, experiments, and analytics aren’t separate products you have to stitch together, but layers in a single system that covers the full product development lifecycle.

This partnership accelerates that vision. Amplitude has been building out Agent Analytics to connect observability and evals with product analytics, while Statsig’s roadmap has been focused on building capabilities like AI Configs for controlling prompts and model parameters without redeploying, and an MCP server integration that embeds experimentation directly into AI coding workflows.

We’re continuing to invest in both platforms with a focus on maintaining the existing Statsig platform across cloud and warehouse deployments and supporting current customers through the transition. We’re also building a shared roadmap that moves both platforms forward together.

Experimentation at the speed of shipping

A year ago, no one knew how the evaluation loop needed to change for probabilistic products. Now we do. AI coding assistants generate more changes than any team can manually validate. LLM-powered products introduce non-deterministic behavior that demands continuous evaluation and validation. The cost of shipping a bad change keeps climbing as products get more complex.

The teams that will outperform with AI aren’t necessarily the ones shipping the most features, but the ones learning what worked and feeding that answer back into the next decision. This creates a feedback loop that accelerates product velocity.

Amplitude spent years making experimentation faster and more accessible. Statsig spent years making it more powerful and more developer-native. Together, we’re building the validation layer that closes the gap between shipping and understanding value.