惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

cs.CV updates on arXiv.org
cs.CV updates on arXiv.org
博客园_首页
M
MIT News - Artificial intelligence
月光博客
月光博客
WordPress大学
WordPress大学
Google DeepMind News
Google DeepMind News
Y
Y Combinator Blog
The Cloudflare Blog
D
Docker
阮一峰的网络日志
阮一峰的网络日志
L
LangChain Blog
Engineering at Meta
Engineering at Meta
Last Week in AI
Last Week in AI
Vercel News
Vercel News
MyScale Blog
MyScale Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Martin Fowler
Martin Fowler
U
Unit 42
Stack Overflow Blog
Stack Overflow Blog
A
About on SuperTechFans
The Register - Security
The Register - Security
B
Blog
Recorded Future
Recorded Future
J
Java Code Geeks
Recent Announcements
Recent Announcements
Microsoft Security Blog
Microsoft Security Blog
H
Help Net Security
F
Fortinet All Blogs
B
Blog RSS Feed
Project Zero
Project Zero
The Hacker News
The Hacker News
T
Threatpost
D
Darknet – Hacking Tools, Hacker News & Cyber Security
L
LINUX DO - 热门话题
Jina AI
Jina AI
宝玉的分享
宝玉的分享
云风的 BLOG
云风的 BLOG
AWS News Blog
AWS News Blog
G
Google Developers Blog
GbyAI
GbyAI
S
Securelist
T
Tenable Blog
博客园 - 【当耐特】
Security Latest
Security Latest
人人都是产品经理
人人都是产品经理
T
Tor Project blog
Latest news
Latest news
P
Proofpoint News Feed
T
The Blog of Author Tim Ferriss

PostHog's RSS Feed

Training our own AI models - PostHog From 270GB RAM to 5GB: Moving local flag evaluation from Django to Rust The best analytics stack for vibe-coded apps The do's and don'ts of minimum viable product marketing - PostHog The best MCP servers for startups, by workflow 4,063 errors closed without a human opening PostHog – here's what we learned - PostHog PostHog Code and the self-driving product - PostHog Why attacking your competitors online is dumb - PostHog The best real-time analytics platforms for developers, compared DuckDB vs ClickHouse: Why we use both at PostHog - PostHog PostHog's next chapter - PostHog Making Claude Cowork actually useful - PostHog PostHog vs Matomo in-depth tool comparison You're doing lifecycle emails wrong Untangling Tokio and Rayon in production: From 2s latency spikes to 94ms flat The best HIPAA-compliant A/B testing tools - PostHog A beginner's guide to testing AI agents - PostHog I hate the standup bot (so I built an agent to do it for me) - PostHog The best CDPs for developers, compared The best error tracking tools for developers, compared The best feature flag software for developers, compared 7 best session replay tools for mobile apps 7 best free open source business intelligence tools right now 7 best free and open source LLM observability tools PostHog vs LogRocket in-depth tool comparison The most popular PostHog alternatives, compared Open source (and self-hosted) session replay tools - PostHog The 9 best GA4 alternatives for apps and websites - PostHog PostHog vs Google Analytics 4 in-depth tool comparison How we built automatic clustering for LLM traces - PostHog The 7 best HIPAA-compliant analytics tools 8 best open source analytics tools you can self-host - PostHog The best product analytics tools for startups, compared PostHog vs FullStory in-depth tool comparison The best in-app survey tools for product teams, compared The 7 best mobile app analytics tools PostHog vs Hotjar in-depth tool comparison The 8 best free and open-source feature flag services - PostHog The 5 best free and open-source A/B testing tools - PostHog The best mobile app A/B testing tools, compared What is a feature flag? Feature Flags vs Remote Config vs A/B Testing PostHog is now available in Vercel’s v0 The best Heap alternatives & competitors, compared PostHog vs Heap in-depth tool comparison PostHog vs Pendo in-depth tool comparison PostHog × Vercel: feature flags, minus the plumbing Your logs' final destination is in GA. You always end up here anyway Behind the scenes of a PostHog hackathon - PostHog The most popular Mixpanel alternatives & competitors, compared PostHog vs Mixpanel in-depth tool comparison The 9 best GDPR-compliant analytics tools How we use Logs at PostHog The best web analytics tools for developers, compared You product data just got a job: Workflows is now out App onboarding: How to fix drop-off points Meet Logs (beta) – logs with all the tools you’re already using Why small teams crush tiger teams How we built user behavior analysis with multi-modal LLMs (in 5 not-so-easy steps) - PostHog The best Contentsquare alternatives & competitors, compared 8 learnings from 1 year of agents – PostHog AI - PostHog Why we killed our AI product assistant Workflows graduate to beta! Product data, meet automation The best Rollbar alternatives & competitors, compared Workflows are now in Alpha and I already broke mine - PostHog I've consistently underestimated how important communication is as a CEO - PostHog How we made feature flags even faster and more reliable The best session replay tools for developers, compared What I learned attending my first ever hackathon - PostHog Did you know AI is answering our community questions? - PostHog How not to be boring - PostHog We built an internal tool to generate changelog images for social media - PostHog What we built at our windswept Mykonos hackathon - PostHog How we built our onboarding email flow (with actual performance data) - PostHog We're building a better PostHog community by closing our public Slack - PostHog Introducing Notebooks for PostHog - PostHog Why we've launched PostHog user surveys - PostHog How we made feature flags faster and more reliable - PostHog In-depth: ClickHouse vs Redshift - PostHog Introducing HouseWatch: An open-source toolkit for ClickHouse - PostHog Introducing HogQL: Direct SQL access for PostHog - PostHog What we built at our sun-kissed Aruba hackathon - PostHog In-depth: ClickHouse vs BigQuery - PostHog In-depth: ClickHouse vs Elasticsearch - PostHog HogMail #22: Why do companies over-hire?" - PostHog Our simpler goal: Help engineers to be better at product - PostHog In-depth: ClickHouse vs Snowflake - PostHog HogMail #21: Avoiding the "Product Death Cycle" - PostHog Sunsetting Kubernetes support for PostHog - PostHog Why 'Product Engineer' is the most fun role I've had in tech - PostHog HogMail #20: Why do startups fail? - PostHog The best Google Optimize alternatives for apps and websites - PostHog Array 1.43.0: Massive performance improvements! - PostHog In-depth: ClickHouse vs Druid - PostHog HogMail #19: Which meetings should you kill? - PostHog CEO diary: The things I learned in 2022 - PostHog The essential tools used by product engineers - PostHog HogMail #18: What can SaaS learn from the New York Times? - PostHog What is a product engineer? - Product Engineer Handbook - PostHog Array 1.42.0: Get beta features via our roadmap! - PostHog HogMail #17: The personal traits that can't be taught - PostHog
Stop AI slop: Run evals with LLM-as-a-Judge - PostHog
Cleo Lant · 2026-01-15 · via PostHog's RSS Feed

Every time your AI product generates text, code, or images, it's being judged.

Not against some complex scoring matrix or your internal metrics, but by a user who's tired, distracted, and one bad output away from a final verdict:

  • "This helped me."
  • "This wasted my time."
  • "This is AI slop and now I don't trust you."

If you’re shipping anything LLM-powered in production, you need a simple way to answer: “Is this AI model doing what I want it to?”

That's what evaluations are for.

PostHog evaluations use LLM-as-a-judge to automatically score generative AI outputs against criteria like relevance, helpfulness, or toxicity.

How it works:

  • Write a short evaluation prompt
  • Choose a sampling rate (0.1% – 100%)
  • Define pass/fail criteria
  • Optionally add property filters to narrow which generations get evaluated

To prevent false negatives, N/A is used when the evaluation prompt is not relevant to the LLM generation. For example, a "mathematical accuracy" evaluation would apply the N/A label to responses that contain no math.

Running evals with AI enables you to batch test hundreds or thousands of traces, then apply human judgement to investigate pass/fail samples. To help you get started, we included five pre-built templates:

TemplateWhat it checksBest for
RelevanceWhether the output addresses the user's inputCustomer support bots, Q&A systems
HelpfulnessWhether the response is useful and actionableChat assistants, support bots, productivity tools
JailbreakAttempts to bypass safety guardrailsSecurity-sensitive applications, apps with PII
HallucinationMade-up facts or unsupported claimsRAG systems, knowledge bases
ToxicityHarmful, offensive, or inappropriate contentUser-facing applications

You can also create custom evals to suit the specific use cases of your AI features, and get a temperature check on user sentiment (more on that later).

AInception

Problem 1: Manual review doesn’t scale

LLM observability tools capture the inputs, outputs, latency, tokens, costs, and errors associated with AI workflows. This makes it simple for engineers to review generations and traces, and hunt for "AI slop".

Slop (a disguting yet accurate term) is any output from an LLM that feels generic, low quality, or just plain wrong.

The problem with manual review is that it doesn't scale. Suppose looking through one complex trace takes an engineer ~15 minutes:

  • 10 traces = 2.5 hours
  • 100 traces = half a work week
  • 10,000 traces = existential dread

Since the average AI product has tens or hundreds of thousands of generations occurring per day, there's no way to review them all and maintain sanity.

Problem 2: Margin of error affects your margins

In January 2024, a user convinced DPD's delivery chatbot to start swearing and criticizing the company. It wrote a poem calling itself "a useless chatbot" and DPD "a customer's worst nightmare." 1.3 million views later, the bot was disabled.

Around the same time, Air Canada's chatbot told a bereaved customer they could retroactively apply for bereavement fares (a policy that didn't actually exist). The airline argued the chatbot was "a separate legal entity." A tribunal disagreed and ordered them to pay $812 plus fees.

This might not sound like a math problem to a product engineer, but it definitely does to legal and finance.

Beyond monitoring for hallucinations and brand disasters, evals are a handy tool to define what "good AI" looks like for your product.

Good output or bad output? That depends on the task. An evaluation configured for a meme generator would pass content that an eval for a scientific research assistant would defintely fail.

Luckily, the best practices for writing evals are simple:

  • Set the domain expertise ("you are a world class sommelier" or "you are are evaluating whether a user is attempting to manipulate an LLM")
  • Be specific about pass/fail criteria
  • Include examples of good vs bad, and edge cases when relevant
  • Keep prompts concise and specific (avoid trying to evaluate multiple things in one shot)

Here's a template you can use:

LLMs fail in unpredictable ways. Using one LLM to judge another will sometimes produce bizarre results. Keep humans in the loop to verify the judge isn't also hallucinating. Your evaluation criteria will drift as you discover new failure modes in production.

Examples of AI slop you can catch with evals:

  • Fake product capabilities and integrations (nightmare for sales and support)
  • Creepy name overuse: "Hey Daniel 😊 That's a great point Daniel, I've got you Daniel."
  • Made-up refund policies, cancellation terms, or upgrade rules
  • Off-brand responses that don't match your voice
  • Lazy outputs, ignoring instructions or dropping context

Evals are primarily used to prevent negative outputs or regressions, but you can also use them to search for positive signals.

Examples of positive signals you can catch with evals:

  • Users discovering creative use cases for AI features you didn't anticipate (potential feature gap)
  • Happy users who might become community champions or case studies (informal NPS)
  • Power users hitting rate limits (upsell opportunity)
  • Feature discovery moments: "Wait, this can do X?" (onboarding gaps)

evals in PostHog

Run multiple evals in parallel to spot-check different behaviors.

Evals are unit tests for your AI product. And like all product data, if you measure it, you can improve it.

But evals alone aren’t enough. A model can “work” and still fail to earn a habit.

This is important because AI-native products have a retention problem. Generous free tiers and easy cancellation attracts "AI tourists" – they extract value, then disappear.

When you connect eval results to real user behavior, you can see which AI behaviors actually affect retention, where users get stuck, and what’s worth fixing next.

The AI product improvement loop:

1. AI Observability shows what your AI is doing

  • See inputs, outputs, latency, tokens, costs, errors
  • Summarize LLM traces and events for quick debugging
  • Run evals to batch test for issues and opportunities

2. Session Replay shows what users see when they interact with AI

  • Compare the front-end user journey with the trace log
  • Watch how users react to poor outputs. Do they retry? Rage-click? Navigate elsewhere?

3. Product Analytics connects AI quality to business metrics

  • Track how AI feature usage correlates with retention, expansion, and revenue
  • Identify which AI features have the worst eval scores and the highest usage? (fix those first)

AI product improvement loop

If you're already using AI Observability in PostHog, you can start creating evaluations right away. Your first 100 evaluation runs are on us. After that, you'll need to use your LLM API key. Evals count as regular LLM events (100K events included on our free tier).

Try evaluations in PostHog