惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

I
InfoQ
D
DataBreaches.Net
Engineering at Meta
Engineering at Meta
GbyAI
GbyAI
Martin Fowler
Martin Fowler
Security Latest
Security Latest
Cisco Talos Blog
Cisco Talos Blog
MongoDB | Blog
MongoDB | Blog
D
Darknet – Hacking Tools, Hacker News & Cyber Security
IT之家
IT之家
cs.AI updates on arXiv.org
cs.AI updates on arXiv.org
S
Security Affairs
www.infosecurity-magazine.com
www.infosecurity-magazine.com
博客园_首页
L
LINUX DO - 最新话题
Know Your Adversary
Know Your Adversary
S
Schneier on Security
The Last Watchdog
The Last Watchdog
Attack and Defense Labs
Attack and Defense Labs
T
Tenable Blog
G
GRAHAM CLULEY
Y
Y Combinator Blog
P
Palo Alto Networks Blog
L
LINUX DO - 热门话题
Hugging Face - Blog
Hugging Face - Blog
W
WeLiveSecurity
C
Cybersecurity and Infrastructure Security Agency CISA
aimingoo的专栏
aimingoo的专栏
博客园 - 司徒正美
The Register - Security
The Register - Security
T
The Exploit Database - CXSecurity.com
MyScale Blog
MyScale Blog
M
MIT News - Artificial intelligence
Cyberwarzone
Cyberwarzone
雷峰网
雷峰网
T
Tailwind CSS Blog
V2EX - 技术
V2EX - 技术
T
Threat Research - Cisco Blogs
S
Secure Thoughts
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
O
OpenAI News
C
Cyber Attacks, Cyber Crime and Cyber Security
The Cloudflare Blog
量子位
Apple Machine Learning Research
Apple Machine Learning Research
T
Threatpost
S
SegmentFault 最新的问题
小众软件
小众软件
Google DeepMind News
Google DeepMind News
Help Net Security
Help Net Security

Snorkel AI

Building AI-Native Systems for Federal Infrastructure: A Conversation with Rezaur Rahman Code World Models and AutoHarness for LLM Agents Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory Building FinQA: An Open RL Environment for Financial Reasoning Agents How Tool Discipline Let a 4B Model Outsmart a 235B Giant on Financial Tasks Coding agents don’t need to be perfect, they need to recover Closing the Evaluation Gap in Agentic AI SlopCodeBench: Measuring Code Erosion as Agents Iterate Introducing the Snorkel Agentic Coding Benchmark 2026: The year of environments Part V: Future Direction and Emerging Trends in Rubric-Based AI Evaluation The self-critique paradox: Why AI verification fails where it’s needed most Chat With the Terminal-Bench Team | Snorkel AI Intelligence per watt: A new metric for AI’s future Terminal-Bench 2.0: Raising the bar for AI agent evaluation Snorkeling in RL environments Introducing SnorkelSpatial: A Benchmark for LLM Spatial Reasoning Scaling Trust: Rubrics in Snorkel's Quality Process Evaluating Multi-Agent Systems in Enterprise Tool Use Evaluating Coding Agents with Terminal-Bench 2.0 Parsing isn’t neutral: why evaluation choices matter The science of rubric design The right tool for the job: An A-Z of rubrics Data quality and rubrics: how to build trust in your models Building the benchmark: inside our agentic insurance underwriting dataset Evaluating AI agents for insurance underwriting LLM observability: key practices, tools, and challenges Anthropic Claude + AWS: revolutionizing pharma data analytics with Snorkel AI Data-centric development of an enterprise AI agent with Snorkel Building the data development platform for specialized AI LLM-as-a-judge for enterprises: evaluate model alignment at scale Why GenAI evaluation requires SME-in-the-loop for validation and trust Research spotlight: is long chain-of-thought structure all that matters when it comes to LLM reasoning distillation? Why enterprise GenAI evaluation requires fine-grained metrics to be insightful What is specialized GenAI evaluation, and why is it so critical to enterprise AI? LLM alignment techniques: 4 post-training approaches Research spotlight: Is intent analysis the key to unlocking more accurate LLM question answering? Why enterprises should embrace LLM distillation Retrieval-augmented generation (RAG) failure modes and how to fix them What is large language model (LLM) alignment? Databricks + Snorkel Flow: integrated, streamlined AI development How LLM evaluation drives better models in Snorkel Flow Unlock proprietary data with Snorkel Flow and Amazon SageMaker LLM evaluation in enterprise applications: a new era in ML Snorkel AI joins the AWS ISV Accelerate Program and launches Snorkel Flow Availability in AWS Marketplace AI data development: a guide for data science projects SnorkelCon 2024: Inaugural Snorkel AI user conference gathers leaders from 30+ Fortune 500 companies Snorkel Flow 2024.R3: Supercharge your AI development with enhanced data-centric workflows Explore the new GenAI Evaluation Suite: Snorkel 2024.R3 New NLP features in Snorkel Flow 2024.R3 Enterprise data compliance and security review: Snorkel Flow 2024.R3 How a global financial services company built a specialized AI copilot accurate enough for production Task Me Anything: innovating multimodal model benchmarks Alfred: Data labeling with foundation models and weak supervision RAG: LLM performance boost with retrieval-augmented generation Call center AI for customer experience management: a case study New GenAI features, data annotation: Snorkel Flow 2024.R2 How data slices transform enterprise LLM evaluation Meta’s Llama 3.1 405B is the new Mr. Miyagi, now what? Meta’s new Llama 3.1 models are here! Are you ready for it? Data-centric AI with Snorkel and MinIO Weak supervision for non-categorical applications + superalignment Snorkel AI signs strategic collaboration agreement with AWS to help enterprises cross the demo-to-production chasm AI alignment made simple: innovative solutions for businesses How does the Snorkel Flow label model work? Vision language models: how LLMs boost image classification Long context models in the enterprise: benchmarks and beyond How to build production-grade RAG retrieval with Snorkel Flow How Bonito helps fine-tune specialized LLMs faster than ever Walking safely before building flying saucer seatbelts: introducing Enterprise Alignment Role-based access controls in Snorkel Flow secure enterprise data Accelerating AI development in manufacturing with Snorkel Flow and AWS SageMaker How ROBOSHOT boosts zero-shot foundation model performance Discover what’s new in Snorkel Flow: Flexible data and LLM connectivity, secure data controls, and more! Faster than ever document intelligence with new Snorkel Flow FM-first workflow The art of data development for Enterprise LLMs Crossing the demo-to-production chasm with Snorkel Custom How Snorkel topped the AlpacaEval leaderboard (and why we're not there anymore) CRFM's HELM and enterprise LLM evaluation beyond accuracy How we achieved 89% accuracy on contract question answering Five sessions not to miss at Google Cloud Next 24 Content filtering breakthrough: Snorkel client reaches 96% recall in 3 days Here's how Snorkel Flow + Google AI built an enterprise-ready model in a day Snorkel teams with Microsoft to showcase new AI research at NVIDIA GTC How Skill-it! enables faster, better LLM training Fine-tuned representation models boost LLM systems. Here's how Enterprise GenAI to surge in 2024: survey results Large language model training: how three training phases shape LLMs LoRA: Low-Rank Adaptation for LLMs LLM distillation demystified: a complete guide Enterprises must shift their focus from models to data in AI development Insurance’s GenAI revolution: a business perspective Scaling human preferences in AI: Snorkel's programmatic approach Building better enterprise AI: incorporating expert feedback in system development “Fall in love with your data”—Snorkel AI’s Enterprise LLM Summit Why QBE Ventures invested in Snorkel AI New benchmark results demonstrate value of Snorkel AI approach to LLM alignment Retrieval augmented generation (RAG): a conversation with its creator Snorkel Flow 2023.R4: enhanced UI + PDF and Databricks tools How Snorkel Flow users can register custom models to Databricks
Stanford professor discusses exciting advances in foundation model evaluation
2024-01-03 · via Snorkel AI

During our Enterprise LLM Summit, Snorkel AI co-founder Alex Ratner sat down with Stanford Computer Science Professor Percy Liang for a conversation about what he and his colleagues are doing at the Stanford Center for Research on Foundation Models (CRFM), and about why Liang views this as such an exciting time for evaluation in machine learning and in AI generally. The transcript that follows has been edited for clarity and brevity.

Alex Ratner: I’d like to start with the term “foundation model.” We now see people using a number of different terms (large language model, foundation model, generative AI) synonymously. Sometimes they just use ChatGPT as a moniker for one of these large pre-trained, self-supervised models.

Why did you and the CRFM group pick “foundation model?” Were you at all surprised by the robust debate over the terminology? Lastly, how does that terminology align with your perspective on CRFM’s objective?

Percy Liang: When we founded the center, it involved me going around Stanford University and seeing who was interested in this phenomenon that was happening post-GPT-3. It was clear that this was going to be a big paradigm shift. Maybe we didn’t anticipate it would happen so quickly, but we knew that this was going to go somewhere. And we felt that it was a phenomenon that was deeper than just language models.

Technically, a language model is just a model over language. Often (but not necessarily) it’s associated with auto-regressive language models, where you predict the next word. But of course, there are vision models and other modalities. We felt “language model” undersold the potential of these models. So we coined the term “foundation model.” 

A foundation model is trained on broad data and can be adapted lightly to a wide range of downstream tasks. We thought a lot about the name and the definition. There is a whole section on naming in the report we wrote, but the important part is that it behaves as a foundation. Instead of people building bespoke models and bespoke datasets in a vertical sense, you have this foundation, which gets built once based on a ton of capital. Then this model can be adapted to a wide range of different tasks, including question answering, customer service use cases, information extraction, et cetera. To us, that signified a paradigm shift in how AI systems were built.

 We felt “language model” undersold the potential of these models. So we coined the term “foundation model.”

Percy Liang, Stanford Professor

In short, “foundation models” is a class that contains large language models. But it also contains visual language models, things like CLIP, and so on.

AR: That makes a lot of sense. I like the foundation metaphor. There is often some “house building” required on top. Obviously a lot of what we’ve done at Snorkel, Stanford, the University of Washington, and many other places is about that house building on top. How do you adapt it or fine-tune it or otherwise customize it?

For specific enterprise settings (i.e. a group that has its own data and objectives) how do you adapt and customize the model? That’s the house on top of the foundation.

PL: That’s a really important point. If, for example, you think about ChatGPT or people who are trying to build AGI, it’s a whole stack.

“Foundation model” illustrates that we’re not building the whole stack. We’re building the foundation. You can’t move into a house if you don’t have the rest of the house, but once you have a strong foundation, the house is much easier to build.

People want to build different houses in different ways, and the idea is that you should be able to customize the style of the model that you want. Every enterprise has different data, and customization is a key part of the paradigm.

AR: Many enterprises and other real-world practitioners are finding out now that you can’t figure out how to do customization, or identify where you need it, without some evaluation metric that’s fine-grained enough.

What are some of your intuitions on where that “house building” is most needed? Where do these models not work out of the box?

“Foundation model” illustrates that we’re not building the whole stack. We’re building the foundation. You can’t move into a house if you don’t have the rest of the house, but once you have a strong foundation, the house is much easier to build.

Percy Liang, Stanford Professor

PL: These models really excel in the prototyping phase. Rather than going to a meeting and making a month-long plan to build some prototype, you prompt it and in five minutes you have a working prototype. The power of being able to do that can’t be overstated. It helps you to brainstorm and to think of things that you might not have on your own.

Prototyping obviously is very different from building an actual robust system. As models get better, of course, you start pushing your capabilities out until you can build something that’s not just a prototype, but actually a working system. In the limits of an application that has many users, custom data, and in which you want fine-grained control, I think you have to customize and fine-tune the data. Otherwise, you’re losing a lot of potential.

It is a progression, and in certain cases you can get quite far out-of-the-box. But, if you’re generating user feedback there has to be a way to use it because the system is not going to read your mind. It can’t be perfect.  

you prompt it and in five minutes you have a working prototype. The power of being able to do that can’t be overstated.

Percy Liang, Stanford Professor

AR: It makes a lot of sense. There may be universal grammars in various data modalities out there. But for what you want to do—what your users want, what your enterprise wants—there’s no mind reading.

I really like that metaphor of “first mile” acceleration, and then “last mile” tuning and development. We’re also seeing a trend where you’ll use one of these massive generalist models to do the first mile, and then you’ll not only fine-tune but also distill or shrink the model into smaller, cheaper, lower-latency specialists.

This naturally leads us to evaluation. If you’re going to do your first-mile explorations, you can’t responsibly ship something to production and you can’t find out what tuning you need in order to traverse that last mile unless you have a sense of how the model is performing. A big part of what you’ve been focused on is the HELM [Holistic Evauation of Language Models] project. We’d love to hear a little bit about that.

PL: Yes, there is so much to say. Let’s see what we can do.

We released HELM about a year ago. It stands for holistic evaluation of language models. We’re trying to develop a standardized way of evaluation. We started with language models, but now we’re moving to multimodal models as well. The challenge is: how do you evaluate a language model?  It’s a generalist system, it’s not like you’re evaluating a spam classifier and you’re getting just a notion of accuracy or AUC. You have to imagine and cover the space of possible use cases, which is challenging. On top of that, we focused on not just accuracy but also bias, robustness, calibration, efficiency, and all these additional factors. It’s holistic because we try to look at all of these different dimensions, scenarios, and about 30 separate models and evaluate systematically.

It’s on the website as a community resource—all the predictions, the code, everything. We’ve been updating it over time because a model comes out every week or so and we are trying frantically to keep up.

You should think about HELM as a framework for evaluation, where you can come with a particular evaluation dataset, which could be custom, or you come with your model, which could be a standard model or a custom model that you fine tune. We then do the prompting and the finagling to produce numbers and reports. We’ve seen people and companies using it for their own purposes, and that’s quite exciting. I hope that it will grow into a much more standard platform for the evaluation of foundation models.

We hope to populate HELM with a wide variety of benchmarks that enterprises would care about.

Percy Liang, Stanford Professor

I should mention that we announced today that we’re working with ML Commons to develop safety evaluations on top of HELM, which is really exciting because that’s an area of utmost concern.

We should think about evaluating language models in terms of both upstream and downstream. The language model is upstream, and the evaluations give you a sense of what the models are capable of. This is not necessarily the metric that you would evaluate if you were looking at the product, but I think it’s valuable to evaluate upstream metrics.  

Later, when you get product experience, you look at whatever tasks and product-specific metrics and try to correlate them with the upstream metrics. Then you have (potentially) a good guide for understanding accuracy on MedQA. I know that this is not perfect, but maybe it’s a weak indicator that this model is actually better at medical knowledge, which is relevant for my domain.  

AR: I like that you make it clear that HELM is a framework. There’s a benchmark that people think about, but HELM is really a framework for building and running benchmarks—which is quite powerful given that a lot of the real production and last-mile development is in using custom datasets and custom objectives.

In terms of upstream and downstream, I imagine a scenario in which downstream is user CSAT or feedback scores, and upstream are these generic, public benchmarks. Do you think there’s a middle ground—especially for enterprises that have their own sub-tasks or sub-datasets—that is missing for enterprise evaluation?

PL: What we’re evolving HELM into, besides the framework, is having this notion of “Helmlets” (or “Helmets,” we haven’t decided on the exact name), where we take slices of the use cases. So, for example, coding, or the medical domain, or the legal domain, or the financial domain, or for different languages. We have one where some folks from Hong Kong University have developed a Chinese evaluation, which we’ve integrated into HELM, and Professor of Computer Science Bo Li has developed this decoding trust benchmark, which is being integrated into HELM and which captures certain aspects of trustworthiness.

So we think about these “sub-domains,” if you will, which can capture the specific things that a user might be interested in. I think we’re going to start thinking about the customer service domain as well, working with some companies on that.

We hope to populate HELM with a wide variety of benchmarks that enterprises would care about. Then the task is less about coming up with a new benchmark and evaluating but just curating. A user’s approach would be: “I care about this, I care about that, and I care about that. Now, rank the models.”  

Every good academic project has to have a properly goofy name.

Alex Ratner, CEO of Snorkel AI

AR: A lot of people have likened this ecosystem to a “family tree” of models. You’d have a lot generalist model lineages, but then you’d increasingly see this family tree of specialized models: specialized for the domain, the sub-domain, the specific group, the specific task, et cetera.

I can imagine evaluations would naturally track that. You’d have the generic ones, then the domain-specific ones, then successively more fine-grained ones as you get more precise.

It’s cool to see how that’s going to get supported. And I love the naming! Every good academic project has to have a properly goofy name.

PL: One more comment I want to make is that, in the era of foundation models, it’s an exciting time for evaluation. Before, in machine learning, you had to get a dataset, you had to hire a whole data team and annotate, and then you could divide and train. The ability to do few-shot learning or even zero-shot learning means that you can focus on just evaluation and you can get domain experts who will sit down and tell you what they want. Then you can rely on the general abilities or raw strength of the models to deliver something interesting.

So, I’m optimistic that we’ll see a lot more interesting evaluations coming online as these models get stronger. That’s exciting because that’s what we always wanted in machine learning: evaluation rather than being stuck with these synthetic or semi-synthetic datasets of the past, which everyone complains about.  

AR: Yes. It goes back to your point about lowering the barrier for the first-mile exploration. You always want to be doing test-driven or evaluation-driven development. Start with focusing where the model is performing poorly. And now, you can get to that point much faster. The ability to get to the evaluation sooner and spend more time there, then do your development, your adaptation, your fine tuning, etc., driven by that evaluation, seems very exciting as a better development paradigm.  

PL: Yes. Totally. We agree.  

AR: Percy, thank you so much for spending the time with us and with this group today. You’re doing some awesome work there and the whole community is grateful for it and we really appreciate getting your thoughts on it today.

PL: Thanks for inviting me, Alex.

More Snorkel AI events coming!

Snorkel has more live online events coming. Look at our events page to sign up for research webinars, product overviews, and case studies.

If you're looking for more content immediately, check out our YouTube channel, where we keep recordings of our past webinars and online conferences.