惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

AWS News Blog
AWS News Blog
N
Netflix TechBlog - Medium
Hugging Face - Blog
Hugging Face - Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
博客园 - Franky
Apple Machine Learning Research
Apple Machine Learning Research
H
Help Net Security
博客园 - 聂微东
Security Archives - TechRepublic
Security Archives - TechRepublic
酷 壳 – CoolShell
酷 壳 – CoolShell
Cisco Talos Blog
Cisco Talos Blog
人人都是产品经理
人人都是产品经理
I
Intezer
C
CERT Recently Published Vulnerability Notes
博客园 - 三生石上(FineUI控件)
Simon Willison's Weblog
Simon Willison's Weblog
Project Zero
Project Zero
The Register - Security
The Register - Security
F
Full Disclosure
Scott Helme
Scott Helme
Cyberwarzone
Cyberwarzone
U
Unit 42
P
Proofpoint News Feed
Engineering at Meta
Engineering at Meta
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
NISL@THU
NISL@THU
T
Tor Project blog
A
Arctic Wolf
Y
Y Combinator Blog
I
InfoQ
M
MIT News - Artificial intelligence
P
Privacy & Cybersecurity Law Blog
Know Your Adversary
Know Your Adversary
Google DeepMind News
Google DeepMind News
Vercel News
Vercel News
V
Vulnerabilities – Threatpost
C
Cybersecurity and Infrastructure Security Agency CISA
Latest news
Latest news
量子位
N
News and Events Feed by Topic
T
The Blog of Author Tim Ferriss
H
Hackread – Cybersecurity News, Data Breaches, AI and More
Hacker News - Newest:
Hacker News - Newest: "LLM"
博客园_首页
S
Securelist
L
LINUX DO - 最新话题
T
Troy Hunt's Blog
The Cloudflare Blog
The GitHub Blog
The GitHub Blog
cs.CV updates on arXiv.org
cs.CV updates on arXiv.org

Snorkel AI

Building AI-Native Systems for Federal Infrastructure: A Conversation with Rezaur Rahman Code World Models and AutoHarness for LLM Agents Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory Building FinQA: An Open RL Environment for Financial Reasoning Agents How Tool Discipline Let a 4B Model Outsmart a 235B Giant on Financial Tasks Coding agents don’t need to be perfect, they need to recover Closing the Evaluation Gap in Agentic AI SlopCodeBench: Measuring Code Erosion as Agents Iterate Introducing the Snorkel Agentic Coding Benchmark 2026: The year of environments Part V: Future Direction and Emerging Trends in Rubric-Based AI Evaluation The self-critique paradox: Why AI verification fails where it’s needed most Chat With the Terminal-Bench Team | Snorkel AI Intelligence per watt: A new metric for AI’s future Terminal-Bench 2.0: Raising the bar for AI agent evaluation Snorkeling in RL environments Introducing SnorkelSpatial: A Benchmark for LLM Spatial Reasoning Scaling Trust: Rubrics in Snorkel's Quality Process Evaluating Multi-Agent Systems in Enterprise Tool Use Evaluating Coding Agents with Terminal-Bench 2.0 Parsing isn’t neutral: why evaluation choices matter The science of rubric design The right tool for the job: An A-Z of rubrics Data quality and rubrics: how to build trust in your models Building the benchmark: inside our agentic insurance underwriting dataset Evaluating AI agents for insurance underwriting LLM observability: key practices, tools, and challenges Anthropic Claude + AWS: revolutionizing pharma data analytics with Snorkel AI Data-centric development of an enterprise AI agent with Snorkel Building the data development platform for specialized AI LLM-as-a-judge for enterprises: evaluate model alignment at scale Why GenAI evaluation requires SME-in-the-loop for validation and trust Research spotlight: is long chain-of-thought structure all that matters when it comes to LLM reasoning distillation? Why enterprise GenAI evaluation requires fine-grained metrics to be insightful LLM alignment techniques: 4 post-training approaches Research spotlight: Is intent analysis the key to unlocking more accurate LLM question answering? Why enterprises should embrace LLM distillation Retrieval-augmented generation (RAG) failure modes and how to fix them What is large language model (LLM) alignment? Databricks + Snorkel Flow: integrated, streamlined AI development How LLM evaluation drives better models in Snorkel Flow Unlock proprietary data with Snorkel Flow and Amazon SageMaker LLM evaluation in enterprise applications: a new era in ML Snorkel AI joins the AWS ISV Accelerate Program and launches Snorkel Flow Availability in AWS Marketplace AI data development: a guide for data science projects SnorkelCon 2024: Inaugural Snorkel AI user conference gathers leaders from 30+ Fortune 500 companies Snorkel Flow 2024.R3: Supercharge your AI development with enhanced data-centric workflows Explore the new GenAI Evaluation Suite: Snorkel 2024.R3 New NLP features in Snorkel Flow 2024.R3 Enterprise data compliance and security review: Snorkel Flow 2024.R3 How a global financial services company built a specialized AI copilot accurate enough for production Task Me Anything: innovating multimodal model benchmarks Alfred: Data labeling with foundation models and weak supervision RAG: LLM performance boost with retrieval-augmented generation Call center AI for customer experience management: a case study New GenAI features, data annotation: Snorkel Flow 2024.R2 How data slices transform enterprise LLM evaluation Meta’s Llama 3.1 405B is the new Mr. Miyagi, now what? Meta’s new Llama 3.1 models are here! Are you ready for it? Data-centric AI with Snorkel and MinIO Weak supervision for non-categorical applications + superalignment Snorkel AI signs strategic collaboration agreement with AWS to help enterprises cross the demo-to-production chasm AI alignment made simple: innovative solutions for businesses How does the Snorkel Flow label model work? Vision language models: how LLMs boost image classification Long context models in the enterprise: benchmarks and beyond How to build production-grade RAG retrieval with Snorkel Flow How Bonito helps fine-tune specialized LLMs faster than ever Walking safely before building flying saucer seatbelts: introducing Enterprise Alignment Role-based access controls in Snorkel Flow secure enterprise data Accelerating AI development in manufacturing with Snorkel Flow and AWS SageMaker How ROBOSHOT boosts zero-shot foundation model performance Discover what’s new in Snorkel Flow: Flexible data and LLM connectivity, secure data controls, and more! Faster than ever document intelligence with new Snorkel Flow FM-first workflow The art of data development for Enterprise LLMs Crossing the demo-to-production chasm with Snorkel Custom How Snorkel topped the AlpacaEval leaderboard (and why we're not there anymore) CRFM's HELM and enterprise LLM evaluation beyond accuracy How we achieved 89% accuracy on contract question answering Five sessions not to miss at Google Cloud Next 24 Content filtering breakthrough: Snorkel client reaches 96% recall in 3 days Here's how Snorkel Flow + Google AI built an enterprise-ready model in a day Snorkel teams with Microsoft to showcase new AI research at NVIDIA GTC How Skill-it! enables faster, better LLM training Fine-tuned representation models boost LLM systems. Here's how Enterprise GenAI to surge in 2024: survey results Large language model training: how three training phases shape LLMs LoRA: Low-Rank Adaptation for LLMs LLM distillation demystified: a complete guide Enterprises must shift their focus from models to data in AI development Insurance’s GenAI revolution: a business perspective Scaling human preferences in AI: Snorkel's programmatic approach Building better enterprise AI: incorporating expert feedback in system development “Fall in love with your data”—Snorkel AI’s Enterprise LLM Summit Why QBE Ventures invested in Snorkel AI New benchmark results demonstrate value of Snorkel AI approach to LLM alignment Retrieval augmented generation (RAG): a conversation with its creator Snorkel Flow 2023.R4: enhanced UI + PDF and Databricks tools How Snorkel Flow users can register custom models to Databricks Stanford professor discusses exciting advances in foundation model evaluation
What is specialized GenAI evaluation, and why is it so critical to enterprise AI?
Shane Johnson · 2025-03-06 · via Snorkel AI

The purpose of GenAI evaluation within an enterprise is to ensure custom AI assistants and copilots respond in a way that meets business requirements and expectations.

Or, perhaps more specifically, that they’re responding the same way an experienced employee would. The challenge is that what makes a response correct is determined by company standards, policies, guidelines and so on. The same judgment expected of experienced employees, especially subject matter experts (SMEs) and/or those in customer-facing roles, must be demonstrated by GenAI assistants and copilots too.

We refer to these expectations as acceptance criteria. They are the characteristics which SMEs consider when determining whether or not a response is acceptable. However, it simply isn’t practical to have SMEs review every single response to see if it meets all of their acceptance criteria.

The problem with enterprise GenAI evaluation is that there is a gaping hole when it comes to applying SME acceptance criteria, evaluating GenAI applications within a business context.

There are plenty of standard LLM benchmarks – MMLU PRO, IFEval, MBPP EvalPlus, MATH, GPQA Diamond and many others. However, in practice, they’re simply used by providers to prove that their latest model is better than competing ones at core capabilities such as coding and reasoning. They’re not particularly helpful for enterprise AI teams when it comes to evaluating their GenAI applications.

It should come as no surprise then to see the emergence of GenAI evaluation platforms for enterprises. They’re more helpful than LLM benchmarks, but they place a strong emphasis on the use of out-of-the-box (OOTB) evaluators – and it’s not enough. Yes, they’re helpful in identifying structural errors such as inefficient retrieval. However, these general evaluators can’t help enterprises understand whether or not their GenAI applications are responding as they should, and if they meet the requirements for production deployment.

Production requires confidence, and confidence requires specialized GenAI evaluators.

Specialized evaluators

Simply put, an evaluator is a function which checks to see if a response (along with the prompt and context) meets a specific acceptance criteria. In this way, an evaluator acts as a proxy for SMEs – allowing AI teams to run automatic, comprehensive and trustworthy evaluations at scale (vs. asking SMEs to review every response one at a time).

There are OOTB evaluators included in every GenAI evaluation platform, often implemented with LLM-as-a-Judge (LLMAJ). However, while these evaluators can help identify structural errors such as poor chunk relevance, they’re not a proxy for SMEs – and can’t determine whether or not a response is acceptable to the business.

The most critical acceptance criteria are based specifically on the domain, business or use cases. Here are few examples which can’t be addressed by OOTB evaluators:

  • [Domain] Adhering to industry regulations
  • [Business] Consistency with brand guidelines
  • [Use case] Following established best practices

What these examples are really enforcing for depends on the context. For example, when building an AI assistant to help with customer service, there may be best practices such as not repeating questions and finishing conversations by asking if there is anything else you can help with. There may be brand guidelines which prohibit certain language or industry regulations which constrain what information and be requested or shared.

Regardless, OOTB evaluators which evaluate structural correctness (e.g., instruction and context adherence) are not enough to provide enterprises with the confidence needed to move forward let alone identify where an AI assistant/copilot is failing to meet business requirements and expectations.

This is why specialized evaluators are required too.

However, creating specialized evaluators isn’t as simple as writing an LLM prompt, and it can’t be done without the help of SMEs.

Prompt engineering

Creating specialized LLMAJ evaluators is, in part, a prompt engineering exercise. It will almost certainly require multiple iterations before the prompt is consistently inducing the correct judgment from an LLM – and input from SMEs whose judgment it’s trying to replicate will be necessary. However, the only way to know it’s doing this is by incorporating reference prompt, context and response triplets. It’s the best way to validate the correctness of an LLMAJ evaluator. Validation requires ground truth.

It’s critical that LLMAJ evaluators align with SMEs. Otherwise, what good are they?

SME alignment

This may be the most critical aspect of evaluation. Because specialized evaluators are created to act as a proxy for SMEs, part of the validation process must include comparing LLMAJ judgments with SME judgments.

If an LLMAJ evaluator produces judgments which match a small amount of ground truth, that’s a positive signal. However, in practice, especially during the initial stages of development, enterprises often find that LLMs and SMEs are far from aligned in terms of judgment – and it’s important to understand where they’re misaligned and the degree to which they are.

The easiest way to do this is to run an evaluation which generates results from the LLMAJ evaluator and assign a high-value subset of the evaluation data to SMEs for expert judgment, which includes their rationale. Next, compare their judgment with that of the LLMAJ. If their alignment is around 50%, that’s not good. It means the LLMAJ evaluator needs improvement, and that will require input from SMEs. If it’s around 90%, then the LLMAJ evaluator can be used as a reliable proxy for SMEs.

Speed, scale and confidence

At the end of the day, the path to production for enterprise AI runs through specialized evaluation.

By creating automated proxies for SMEs, AI teams can accelerate the evaluation process exponentially without sacrificing quality. And it scales AI adoption by enabling them to not only run trustworthy evaluations frequently, but to standardize on a framework which can be used to support evaluation of all GenAI applications – whether assistants or copilots, employee-facing or customer-facing.

Finally, it’s necessary to provide the confidence needed for production deployment. All too often, when enterprises find themselves in the news for less than positive reasons regarding AI, it’s because GenAI applications were not properly evaluated – and the risk was neither recognized nor mitigated.

Specialization is but one aspect of enterprise GenAI evaluation. Stay tuned as we’ll discuss another critical aspect in our next blog – insight.