惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

The GitHub Blog
The GitHub Blog
IT之家
IT之家
B
Blog RSS Feed
罗磊的独立博客
GbyAI
GbyAI
博客园 - Franky
Y
Y Combinator Blog
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
Google DeepMind News
Google DeepMind News
博客园 - 聂微东
N
Netflix TechBlog - Medium
博客园 - 三生石上(FineUI控件)
人人都是产品经理
人人都是产品经理
U
Unit 42
博客园 - 叶小钗
Jina AI
Jina AI
MyScale Blog
MyScale Blog
雷峰网
雷峰网
B
Blog
Hugging Face - Blog
Hugging Face - Blog
Blog — PlanetScale
Blog — PlanetScale
Recent Announcements
Recent Announcements
腾讯CDC
酷 壳 – CoolShell
酷 壳 – CoolShell

Snorkel AI

Building AI-Native Systems for Federal Infrastructure: A Conversation with Rezaur Rahman Code World Models and AutoHarness for LLM Agents Benchtalks #1: Alex Shaw (Terminal-Bench, Harbor) – Building the Benchmark Factory Building FinQA: An Open RL Environment for Financial Reasoning Agents How Tool Discipline Let a 4B Model Outsmart a 235B Giant on Financial Tasks Coding agents don’t need to be perfect, they need to recover Closing the Evaluation Gap in Agentic AI SlopCodeBench: Measuring Code Erosion as Agents Iterate Introducing the Snorkel Agentic Coding Benchmark 2026: The year of environments Part V: Future Direction and Emerging Trends in Rubric-Based AI Evaluation The self-critique paradox: Why AI verification fails where it’s needed most Chat With the Terminal-Bench Team | Snorkel AI Intelligence per watt: A new metric for AI’s future Terminal-Bench 2.0: Raising the bar for AI agent evaluation Snorkeling in RL environments Introducing SnorkelSpatial: A Benchmark for LLM Spatial Reasoning Scaling Trust: Rubrics in Snorkel's Quality Process Evaluating Multi-Agent Systems in Enterprise Tool Use Evaluating Coding Agents with Terminal-Bench 2.0 Parsing isn’t neutral: why evaluation choices matter The science of rubric design The right tool for the job: An A-Z of rubrics Data quality and rubrics: how to build trust in your models Building the benchmark: inside our agentic insurance underwriting dataset Evaluating AI agents for insurance underwriting LLM observability: key practices, tools, and challenges Anthropic Claude + AWS: revolutionizing pharma data analytics with Snorkel AI Data-centric development of an enterprise AI agent with Snorkel Building the data development platform for specialized AI
Snorkel teams with Microsoft to showcase new AI research ...
Doug Kelly (Microsoft), Friea Berg (Snorkel AI) · 2024-03-19 · via Snorkel AI

Snorkel AI researchers work on the cutting edge of AI innovation to help expand the boundaries of AI knowledge.

That’s a bold statement, but accurate.

As part of a research-first culture, the Snorkel team has contributed to over 150 academic papers on topics covering LLM data curation, LLM evaluation, model distillation, and more. Snorkel AI founders and researchers present regularly at distinguished AI conferences such as NeurIPS, where the team was recognized with a Best Paper award for “Low-Resource Languages Jailbreak GPT-4. A recent co-presentation on MedAlign, a curated open-source benchmark dataset for the evaluation of LLMs for EHR data retrieval, was awarded Best Findings Paper in GenAI for Health.

Snorkel’s partnership with Microsoft plays a critical role in equipping its research team to experiment with new techniques and approaches. Snorkel is a member of the Microsoft for Startups Pegasus Program, Microsoft’s flagship go-to-market program. Through the Pegasus program, Snorkel has access to premier sales resources and technical assets to accelerate AI workloads including early access to Azure AI services, leading models from OpenAI and Mistral, and accelerated high-performance compute. The ability for Snorkel’s Research team to execute its most demanding projects on Azure AI infrastructure powered by NVIDIA GPUs has been a game changer. 

Snorkel’s recent top tier ranking on the AlpacaEval 2.0 LLM leaderboard would not have been possible without the program’s dedicated startups GPU cluster benefit. Access to state-of-the-art Azure NDm NVIDIA A100 instances via a seamless Azure experience has empowered Snorkel to drive cutting-edge research in programmatic alignment/DPO in a quick & efficient manner. Azure AI Infrastructure VMs, which come pre-configured with InfiniBand and NVLINK for optimized scale-out and scale-up, allow Snorkel researchers to run quick experiments from small projects to large-scale distributed jobs on multiple GPUs reliably and with full monitoring mechanisms.

The value of this research extends far beyond academic papers and benchmark results. The Snorkel Flow data development platform was intentionally designed with flexible abstractions and extensible interfaces that allow for the continual integration of the latest and most remarkable technologies from academic collaborations to create an ever more powerful tool for users. Ultimately, Snorkel’s cutting-edge research plays a key role help enterprises successfully move AI projects from prototype to production.

Today, Snorkel is proud to congratulate Microsoft on the new generally available Azure ​​NC H100 v5 VMs, which are tailored to accelerate large-scale AI model training and batch inference. As Microsoft advances the state-of-the-art with optimized AI GPU VMs leveraging the latest NVIDIA technologies, the Snorkel research team can launch increasingly ambitious and challenging projects.  

To learn more, we invite you to read the Microsoft blog post “Microsoft and NVIDIA partnership continues to deliver on the promise of AI” and visit us at NVIDIA GTC.

Learn more in person at NVIDIA GTC

Visit the Microsoft booth at NVIDIA GTC to learn more about the Snorkel research that resulted in a top tier ranking on the AlpacaEval 2.0 LLM leaderboard. Plus, Snorkel will share how designing projects on Azure AI infrastructure powered by NVIDIA GPUs helps our researchers deliver value for our customers and the OSS community even faster.

Topic: Snorkel AI research leverages Azure AI Infrastructure powered by NVIDIA GPUs for cutting-edge AI/ML.
Location: Microsoft booth #1108
Timing: March 19, 3:40-4pm

This blog post is a collaboration between Microsoft for Startups Senior AI Advisor Doug Kelly and Snorkel Head of Partnerships Friea Berg.