惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

爱范儿
爱范儿
博客园_首页
U
Unit 42
Apple Machine Learning Research
Apple Machine Learning Research
云风的 BLOG
云风的 BLOG
MongoDB | Blog
MongoDB | Blog
美团技术团队
H
Help Net Security
G
Google Developers Blog
B
Blog RSS Feed
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
aimingoo的专栏
aimingoo的专栏
Google DeepMind News
Google DeepMind News
J
Java Code Geeks
M
MIT News - Artificial intelligence
腾讯CDC
IT之家
IT之家
Vercel News
Vercel News
C
Check Point Blog
博客园 - 三生石上(FineUI控件)
Last Week in AI
Last Week in AI
I
InfoQ
博客园 - 司徒正美
A
About on SuperTechFans

Snorkel AI

Building AI-Native Systems for Federal Infrastructure: A Conversation with Rezaur Rahman Code World Models and AutoHarness for LLM Agents Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory Building FinQA: An Open RL Environment for Financial Reasoning Agents How Tool Discipline Let a 4B Model Outsmart a 235B Giant on Financial Tasks Coding agents don’t need to be perfect, they need to recover Closing the Evaluation Gap in Agentic AI SlopCodeBench: Measuring Code Erosion as Agents Iterate Introducing the Snorkel Agentic Coding Benchmark 2026: The year of environments Part V: Future Direction and Emerging Trends in Rubric-Based AI Evaluation The self-critique paradox: Why AI verification fails where it’s needed most Chat With the Terminal-Bench Team | Snorkel AI Intelligence per watt: A new metric for AI’s future Terminal-Bench 2.0: Raising the bar for AI agent evaluation Snorkeling in RL environments Introducing SnorkelSpatial: A Benchmark for LLM Spatial Reasoning Scaling Trust: Rubrics in Snorkel's Quality Process Evaluating Multi-Agent Systems in Enterprise Tool Use Evaluating Coding Agents with Terminal-Bench 2.0 Parsing isn’t neutral: why evaluation choices matter The science of rubric design The right tool for the job: An A-Z of rubrics Data quality and rubrics: how to build trust in your models Building the benchmark: inside our agentic insurance underwriting dataset Evaluating AI agents for insurance underwriting LLM observability: key practices, tools, and challenges Anthropic Claude + AWS: revolutionizing pharma data analytics with Snorkel AI Data-centric development of an enterprise AI agent with Snorkel Building the data development platform for specialized AI
Content filtering breakthrough: Snorkel client reaches 96...
Gabe Smith · 2024-03-26 · via Snorkel AI

The world of social media moves fast, which poses a challenge for those who need to efficiently and accurately filter social media content. Snorkel AI recently worked with a large social media management that faced just this kind of challenge.

They needed a more effective model for tagging profiles according to whether or not they linked to adult content. Their existing model fell short, and the tools at their disposal proved insufficient.

This was a serious concern. They had a significant partnership that hinged on improving their ability to accurately classify adult profiles.

Enter Snorkel AI. Our mission is to democratize AI by making it easier for enterprises to build and deploy machine learning models. We were ready to help.

I recently talked with Matt Casey, data science content lead at Snorkel AI, about this case. You can watch the full interview (embedded below), but I’ve summed up the main points here.

The content filtering challenge: hitting a ceiling with no way through

The client was in a challenging situation. Their existing model achieved a recall of about 85% in identifying adult-oriented profiles. An impending partnership demanded a model with a recall in the upper nineties.

The platform they used before turning to Snorkel AI presented several roadblocks. First, it was not conducive to quick iterations, a key requirement given the client’s strict timeline and the fast-paced nature of social media content.

Second, their existing tool made the process of labeling new data cumbersome and slow. Compounding this problem, the client had no labeled data to begin with. Even if they did, their existing platform didn’t offer an easy way to incorporate labeled data into the existing model.

In short, they were stuck. They had a deadline looming, a goal to meet, and no clear path to meeting that goal before the clock ran out.

So, they reached out to us.

Snorkel AI’s solution: Snorkel Flow

We introduced the client to Snorkel Flow, our AI data development platform. The platform amplifies the impact of subject matter experts (SMEs) to scale and streamline the data labeling process.

Snorkel Flow’s programmatic labeling process starts with labeling functions—essentially programmable rules to label data. Snorkel Flow users can build labeling functions according to various data features—from continuous variable thresholds to vector embedding clusters. In this case, the client’s labeling functions were primarily substring-based, focusing on identifying specific keywords in the data.

This resulted in an unusually high number of labeling functions. By the end of the project, the client’s users had created 160 separate labeling functions. Some Snorkel Flow projects can use as few as ten labeling functions, but this keyword approach allowed them to cover many edge cases specific to adult content.  

The results: a content filtering model above target and on time

Before turning to Snorkel Flow, the customer projected that the project would take six months. They would have had to manually label tens of thousands of profiles to lift model performance to the level needed. And they didn’t have six months to spare.

Instead, the client achieved a recall of 96% using Snorkel Flow In just three weeks. The quick turnaround was particularly impressive considering the absence of any labeled data at the onset, and particularly valuable because it allowed them to complete their project ahead of the deadline dictated by their pending partnership.

Ongoing success and future plans

The client continues to use our platform independently. They recently revalidated their model to account for data drift (which is significant in the world of social media and adult content) and found that it remained more accurate than they expected.

The client updated or removed a small number of labeling functions and exported a new version of the model to keep its recall high. I want to note that this would not be so easy with manual labeling. Reinvestigating the data and updating problematic labels could have taken human labelers several days—perhaps weeks—of cumulative labor. Our client completed this task in a couple of hours.

The outputs of this model have become central to the client’s data lake, powering downstream analytics and recommendation models. This model isn’t just a standalone solution: it’s a key piece that enables many other operations within the company.

Looking ahead, the client plans to expand their use of Snorkel Flow to other projects. We’re excited to continue supporting them in their machine learning journey.

Snorkel Flow: accelerating AI data development

This case study underscores the transformative power of machine learning in improving content filtering. Through Snorkel Flow, the client was able to drastically improve their adult content labeling model, meet their partnership requirements, and set the stage for future success.

As machine learning continues to evolve, we’re excited to see how it will further revolutionize content filtering and other critical business operations.

Ready to accelerate AI development?

Deploy production AI and ML applications 10-100x faster with Snorkel’s experts, using our proprietary technology.

Request a demo

Gabe smith: zero to 96% recall content filtering classifier in just 3 days!