惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

V
Visual Studio Blog
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
T
The Blog of Author Tim Ferriss
宝玉的分享
宝玉的分享
The Register - Security
The Register - Security
D
Docker
The Cloudflare Blog
A
About on SuperTechFans
Microsoft Security Blog
Microsoft Security Blog
Recent Announcements
Recent Announcements
月光博客
月光博客
B
Blog RSS Feed
博客园 - 【当耐特】
The GitHub Blog
The GitHub Blog
B
Blog
IT之家
IT之家
美团技术团队
Engineering at Meta
Engineering at Meta
C
Check Point Blog
云风的 BLOG
云风的 BLOG
Last Week in AI
Last Week in AI
G
Google Developers Blog
MongoDB | Blog
MongoDB | Blog
Microsoft Azure Blog
Microsoft Azure Blog
S
SegmentFault 最新的问题
V
V2EX
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
Apple Machine Learning Research
Apple Machine Learning Research
U
Unit 42
H
Help Net Security
雷峰网
雷峰网
人人都是产品经理
人人都是产品经理
博客园 - 司徒正美
Stack Overflow Blog
Stack Overflow Blog
博客园 - Franky
PCI Perspectives
PCI Perspectives
J
Java Code Geeks
CTFtime.org: upcoming CTF events
CTFtime.org: upcoming CTF events
M
MIT News - Artificial intelligence
腾讯CDC
A
Arctic Wolf
C
CERT Recently Published Vulnerability Notes
量子位
C
CXSECURITY Database RSS Feed - CXSecurity.com
Latest news
Latest news
钛媒体:引领未来商业与生活新知
钛媒体:引领未来商业与生活新知
The Hacker News
The Hacker News
有赞技术团队
有赞技术团队
Schneier on Security
Schneier on Security
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻

Snorkel AI

Building AI-Native Systems for Federal Infrastructure: A Conversation with Rezaur Rahman Code World Models and AutoHarness for LLM Agents Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory Building FinQA: An Open RL Environment for Financial Reasoning Agents How Tool Discipline Let a 4B Model Outsmart a 235B Giant on Financial Tasks Coding agents don’t need to be perfect, they need to recover Closing the Evaluation Gap in Agentic AI SlopCodeBench: Measuring Code Erosion as Agents Iterate Introducing the Snorkel Agentic Coding Benchmark 2026: The year of environments Part V: Future Direction and Emerging Trends in Rubric-Based AI Evaluation The self-critique paradox: Why AI verification fails where it’s needed most Chat With the Terminal-Bench Team | Snorkel AI Intelligence per watt: A new metric for AI’s future Terminal-Bench 2.0: Raising the bar for AI agent evaluation Snorkeling in RL environments Introducing SnorkelSpatial: A Benchmark for LLM Spatial Reasoning Scaling Trust: Rubrics in Snorkel's Quality Process Evaluating Multi-Agent Systems in Enterprise Tool Use Evaluating Coding Agents with Terminal-Bench 2.0 Parsing isn’t neutral: why evaluation choices matter The science of rubric design The right tool for the job: An A-Z of rubrics Data quality and rubrics: how to build trust in your models Building the benchmark: inside our agentic insurance underwriting dataset Evaluating AI agents for insurance underwriting LLM observability: key practices, tools, and challenges Anthropic Claude + AWS: revolutionizing pharma data analytics with Snorkel AI Data-centric development of an enterprise AI agent with Snorkel Building the data development platform for specialized AI LLM-as-a-judge for enterprises: evaluate model alignment at scale Why GenAI evaluation requires SME-in-the-loop for validation and trust Research spotlight: is long chain-of-thought structure all that matters when it comes to LLM reasoning distillation? Why enterprise GenAI evaluation requires fine-grained metrics to be insightful What is specialized GenAI evaluation, and why is it so critical to enterprise AI? LLM alignment techniques: 4 post-training approaches Research spotlight: Is intent analysis the key to unlocking more accurate LLM question answering? Why enterprises should embrace LLM distillation Retrieval-augmented generation (RAG) failure modes and how to fix them What is large language model (LLM) alignment? Databricks + Snorkel Flow: integrated, streamlined AI development How LLM evaluation drives better models in Snorkel Flow Unlock proprietary data with Snorkel Flow and Amazon SageMaker LLM evaluation in enterprise applications: a new era in ML Snorkel AI joins the AWS ISV Accelerate Program and launches Snorkel Flow Availability in AWS Marketplace AI data development: a guide for data science projects SnorkelCon 2024: Inaugural Snorkel AI user conference gathers leaders from 30+ Fortune 500 companies Snorkel Flow 2024.R3: Supercharge your AI development with enhanced data-centric workflows Explore the new GenAI Evaluation Suite: Snorkel 2024.R3 New NLP features in Snorkel Flow 2024.R3 Enterprise data compliance and security review: Snorkel Flow 2024.R3 How a global financial services company built a specialized AI copilot accurate enough for production Task Me Anything: innovating multimodal model benchmarks Alfred: Data labeling with foundation models and weak supervision RAG: LLM performance boost with retrieval-augmented generation Call center AI for customer experience management: a case study New GenAI features, data annotation: Snorkel Flow 2024.R2 How data slices transform enterprise LLM evaluation Meta’s Llama 3.1 405B is the new Mr. Miyagi, now what? Meta’s new Llama 3.1 models are here! Are you ready for it? Data-centric AI with Snorkel and MinIO Weak supervision for non-categorical applications + superalignment Snorkel AI signs strategic collaboration agreement with AWS to help enterprises cross the demo-to-production chasm AI alignment made simple: innovative solutions for businesses How does the Snorkel Flow label model work? Vision language models: how LLMs boost image classification Long context models in the enterprise: benchmarks and beyond How to build production-grade RAG retrieval with Snorkel Flow How Bonito helps fine-tune specialized LLMs faster than ever Walking safely before building flying saucer seatbelts: introducing Enterprise Alignment Role-based access controls in Snorkel Flow secure enterprise data Accelerating AI development in manufacturing with Snorkel Flow and AWS SageMaker How ROBOSHOT boosts zero-shot foundation model performance Discover what’s new in Snorkel Flow: Flexible data and LLM connectivity, secure data controls, and more! Faster than ever document intelligence with new Snorkel Flow FM-first workflow The art of data development for Enterprise LLMs Crossing the demo-to-production chasm with Snorkel Custom How Snorkel topped the AlpacaEval leaderboard (and why we're not there anymore) CRFM's HELM and enterprise LLM evaluation beyond accuracy How we achieved 89% accuracy on contract question answering Five sessions not to miss at Google Cloud Next 24 Content filtering breakthrough: Snorkel client reaches 96% recall in 3 days Here's how Snorkel Flow + Google AI built an enterprise-ready model in a day Snorkel teams with Microsoft to showcase new AI research at NVIDIA GTC How Skill-it! enables faster, better LLM training Fine-tuned representation models boost LLM systems. Here's how Enterprise GenAI to surge in 2024: survey results Large language model training: how three training phases shape LLMs LoRA: Low-Rank Adaptation for LLMs LLM distillation demystified: a complete guide Enterprises must shift their focus from models to data in AI development Insurance’s GenAI revolution: a business perspective Scaling human preferences in AI: Snorkel's programmatic approach “Fall in love with your data”—Snorkel AI’s Enterprise LLM Summit Why QBE Ventures invested in Snorkel AI New benchmark results demonstrate value of Snorkel AI approach to LLM alignment Retrieval augmented generation (RAG): a conversation with its creator Snorkel Flow 2023.R4: enhanced UI + PDF and Databricks tools How Snorkel Flow users can register custom models to Databricks Stanford professor discusses exciting advances in foundation model evaluation
Building better enterprise AI: incorporating expert feedback in system development
Chris Glaze · 2024-01-30 · via Snorkel AI

Enterprises that aim to build valuable, useful generative AI applications must view them from a systems-level. While large language models form the core of these applications, they exist as part of an ecosystem that connects numerous components, each of which plays a vital role in the end-users experience.

I recently discussed some of my work on generative AI (GenAI) applications in a talk called “Data Development for GenAI: A Systems Level View” at Snorkel AI’s Enterprise LLM Summit. My talk focussed on the importance of understanding the larger ecosystem in which large language models (LLMs) exist and how fine-tuning all system components with expert feedback can improve application performance.

You can watch the entire talk on our YouTube page, but I’ve summarized the main points below.

LLM application ecosystems

LLMs don’t exist in a vacuum. Used correctly, they form the foundation of an application ecosystem where data feeds various system components both upstream and downstream.

This ecosystem can include:

  • Input data pre-processing
  • Metadata, such as document tags
  • Embedding spaces, eg for vector databases
  • Document retrieval systems
  • LLM prompt templates
  • The LLM itself
  • LLM response detection algorithms (eg hallucinations, private company information)
  • Front-end components to facilitate user interaction

The pre-processing stage cleans and prepares data for the LLM. Next steps may include embedding models and a retrieval systems to find appropriate passages to inject as context in the final prompt. Then, the model itself ingests the prompt and yields a response, but the process isn’t over. The post-processing stage refines the response. This stage might, for instance, detect hallucinations to ensure the AI’s generated content is coherent and sensible before delivering the final output to the user.

Each of these components plays a crucial role in the system’s performance. However, their performance isn’t static. They require fine-tuning based on expert feedback

Incorporating expert feedback into the fine-tuning process has previously been a challenge—primarily due to scalability issues. AI projects often demand vast amounts of data, and experts rarely want to review and annotate each record individually.

Image4

Case study: improving a RAG system for a global bank

A team of researchers and engineers at Snorkel recently completed a collection of enterprise GenAI projects, including one that improved a retrieval augmented generation (RAG) system for a top global bank. This system handled complex, unstructured financial documents and answered business-critical questions about them.

Through a combination of programmatic data development techniques, we fine-tuned every component of the RAG system. The result was a significant accuracy boost, with only a minimal amount of time required from subject matter experts (SMEs) to explain how to interpret document language and where to look for key pieces of information.

Document chunking and tagging
Image3

First, we had to improve how the application chunked and tagged documents. For example, the retrieval model struggled with questions about dates. Dates may seem like a simple concept, but these documents often referenced dates in subtle ways the model missed.

We consulted with an SME to create a custom date extractor. This new extractor, developed in only four hours, had a 99 macro average F1 score.

We also discovered that the retrieval system struggled with legal definitions. With some more SME help, we developed a model to identify these legal terms. With about four hours of work, we achieved a 93 macro average.

These two support components set the stage to help our users find the information they needed when they needed it.

Fine tuning the embedding space

Next, we moved to the embedding space. The off-the-shelf embedding model we started with treated key document sections as being too similar than they were from an SME’s point of view, making them difficult to differentiate. We asked the SME on the project to manually annotate a small number of documents with explanations as to why they made their data annotations. We then adapted the expert’s logic and intuition to label a larger training set that we used to fine-tune our original embedding model. The result was a much clearer distinction between relevant and irrelevant sections of the document.

Image1

Improving chunk retrieval

We also improved the chunk retrieval process. The original algorithm used fixed parameters, which proved too rigid. We instead made the process adaptive to the question and the distribution of relevance scores. As a result, our system pulled all of the most relevant chunks that fit within the LLM’s context window.

Results

By fine-tuning all the components of the RAG system, we achieved a 54-point increase in question-answering accuracy in just three weeks. This improvement required fewer than 10 hours from the SME—who spent more time verifying the system’s output than annotating documents.

Image2

The role of expert feedback in AI development

Expert feedback has always been a vital component of valuable AI systems. It’s even more important in GenAI systems. In our case study, the insights, logic, and intuition the SME provided guided how we developed and fine-tuned different components. However, incorporating expert feedback into the AI development process historically posed a significant challenge. Their time is valuable, and data labeling can be onerous.

Our approach maximizes the impact of SME involvement through scalable tooling that multiplies expert logic and intuitions without the need for extensive manual annotation. This not only accelerates the development process but also enhances the effectiveness and accuracy of the resulting AI models.

Expert input + systems-level thinking: the key to enterprise AI

Taking a systems-level perspective and incorporating expert input leads to stronger, more valuable, and more useful GenAI applications. Expert feedback plays a crucial role in this process, and scalable solutions allow data scientists to efficiently incorporate this feedback while minimizing the burden on SMEs.

As we continue to develop methods, we look forward to seeing further improvements in the accuracy and effectiveness of LLM-backed and GenAI applications.

More Snorkel AI events coming!

Snorkel has more live online events coming. Look at our events page to sign up for research webinars, product overviews, and case studies.

If you're looking for more content immediately, check out our YouTube channel, where we keep recordings of our past webinars and online conferences.