惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com
B
Blog
Jina AI
Jina AI
N
Netflix TechBlog - Medium
Cyber Security Advisories - MS-ISAC
Cyber Security Advisories - MS-ISAC
博客园_首页
Hugging Face - Blog
Hugging Face - Blog
博客园 - 聂微东
美团技术团队
Google DeepMind News
Google DeepMind News
WordPress大学
WordPress大学
阮一峰的网络日志
阮一峰的网络日志
U
Unit 42
The Cloudflare Blog
V
V2EX
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
小众软件
小众软件
罗磊的独立博客
Microsoft Security Blog
Microsoft Security Blog
Apple Machine Learning Research
Apple Machine Learning Research
I
InfoQ
GbyAI
GbyAI
腾讯CDC
MongoDB | Blog
MongoDB | Blog

Vector Institute for Artificial Intelligence

Mohamad Moosavi: Accelerating the search for climate solutions with AI A strategic blueprint for safe health AI implementation: Your 2026 roadmap Vector Institute awards 100 scholarships to Ontario’s top AI graduate students Agentic AI evaluation strategies Hassan Ashtiani: Building trustworthy AI through mathematical foundations Vector researchers advance representation learning and deep learning research at ICLR 2026 Remarkable 2026 Poster Session: 60 research projects shaping AI’s future CRISPNAM-FG: An interpretable Fine-Gray deep survival model for competing risks in health care Demo Day: How the Vector Institute helps Canadian startups turn innovative ideas into commercial reality The New Cartography of the Invisible Vector researchers advance AI frontiers with 80 papers at NeurIPS 2025 New study reveals AI’s $100B economic impact across Canada, with Ontario leading the charge When smart AI gets too smart: Key insights from Vector’s 2025 ML Security & Privacy Workshop Vector Institute names 13 new Faculty Members, expanding core research leadership across Ontario Vector researchers dive into deep learning at ICLR 2025 When AI Meets Human Matters: Evaluating Multimodal Models Through a Human-Centred Lens – Introducing HumaniBench Vector Institute 2024-25 annual report: Where AI research meets real-world impact Vector researchers tackle real-world AI challenges at ICML 2025 Ontario’s AI ecosystem: fueling real economic growth with record number of jobs and private investments Transforming Youth Mental Health Support: FAIIR’s AI-Powered Crisis Response Model Vector Institute awards up to $2.1 million in scholarships to Ontario’s top AI graduate students AI Weather Forecasting Breakthrough: How Canadian Innovation is Transforming Climate Prediction | Aardvark Weather Exploring Intelligence: Vector Faculty Member Kelsey Allen’s Path from Particle Physics to Cognitive Machine Learning Vector Institute Announces the Appointment of Glenda Crisp as President and CEO State of Evaluation Study: Vector Institute Unlocks New Transparency in Benchmarking Global AI Models Real World Multi-Agent Reinforcement Learning – Latest Developments and Applications Principles in Action: Introducing the Vector Institute’s Playbook for Responsible AI Product Development Leveraging Large Language Models for More Efficient Systematic Reviews in Medicine and Beyond Global AI Alliance for Climate Action funding announcement CEO Update
Vector Institute Unveils Comprehensive Evaluation of Lead...
Kylie Williams · 2025-04-10 · via Vector Institute for Artificial Intelligence

At a glance:

  • Canada’s Vector Institute has assessed 11 leading AI models from around the world, using 16 performance benchmarks, including those pioneered by Vector researchers.
  • The State of Evaluation study marks the first time that both open and closed-source models have been evaluated against an expanded suite of benchmarks, revealing leaders and laggards in model performance. 
  • The independent results can help organizations develop, deploy, and apply AI safely and responsibly.
  • In a first for this kind of research, Vector has shared the benchmarks, underlying code, and results in open-source to foster accountability, transparency, and collaboration that builds trust in AI. 

TORONTO, ON, April 10, 2025 — Canada’s Vector Institute has unveiled the results of its independent evaluation of leading large language models (LLMs), offering an objective look at how prominent frontier AI models perform against a comprehensive suite of benchmarks. The study, summarized in a new article on its website, assesses capabilities in increasingly complex tests of general knowledge, coding, cyber-safety, and other critical areas, providing key insights into the strengths and limitations of top AI agents.

AI companies are releasing new and more powerful LLMs at an unprecedented pace, with each new model promising greater capabilities from more human-like text generation to advanced problem-solving and decision-making. Developing widely used and trusted benchmarks advances AI safety; it helps researchers, developers, and users understand how these models perform in terms of accuracy, reliability, and fairness, enabling their responsible deployment.

In its State of Evaluation study, Vector’s AI Engineering team assessed 11 leading LLMs from around the world, including both publicly available (‘open’) models such as DeepSeek-R1 and Cohere’s Command R+, as well as commercial (‘closed’) models such as OpenAI’s GPT-4o and Gemini 1.5 from Google. Each agent was tested against 16 performance benchmarks, making this one of the most comprehensive, independent evaluations conducted to date. 

“Independent, objective evaluation of this kind is vital to understanding how models perform in terms of accuracy, reliability, and fairness,” explains Deval Pandya, Vector’s Vice President of AI Engineering. “Robust benchmarks and accessible evaluations enable researchers, organizations, and policymakers to better understand the strengths, weaknesses, and real-world impact of these rapidly evolving, highly capable AI models and systems, and ultimately to foster trust in AI.”

In a first for this kind of research, Vector has shared the results of the study, the benchmarks,  and the underlying code in an open-sourced, interactive leaderboard to promote transparency and foster advances in AI innovation. “Researchers, developers, regulators, and end-users can independently verify results, compare model performance, and build out their own benchmarks and evaluations to drive improvements and accountability,” says John Willes, Vector’s AI Infrastructure and Research Engineering Manager, who led the project.

The project is a natural extension of Vector’s leadership in developing the benchmarks now used widely across the global AI safety community, including MMLU-Pro, MMMU, and OS-World, which were developed by Vector Institute Faculty Members and Canada CIFAR AI Chairs Wenhu Chen and Victor Zhong. It also builds on recent work by Vector’s AI Engineering team to develop Inspect Evals — an open-source AI safety testing platform created in collaboration with the UK AI Security Institute to standardize global safety evaluations and facilitate collaboration among researchers and developers.

“As organizations seek to unlock the transformative benefits of AI, Vector is in a unique position to provide independent, trusted expertise that enables them to do so safely and responsibly,” explains Pandya, citing the institute’s programs in which its industry partners collaborate with expert researchers at the forefront of AI safety and application. “Whether they’re in financial services, technology innovation, or health or more, our industry partners have access to Vector’s unparalleled sandbox environment where they can experiment and test models and techniques to help address their specific AI-related business challenges.”

  • Read more about Vector Institute’s “State of Evaluation” here.
  • Explore interactive leaderboard here.

About Vector Institute: The Vector Institute is an independent, not-for-profit corporation dedicated to advancing artificial intelligence, excelling in machine learning and deep learning. Our vision is to drive excellence and leadership in Canada’s knowledge, creation, and use of AI to foster economic growth and improve the lives of Canadians. The Vector Institute is funded by the Province of Ontario, the Government of Canada through CIFAR Pan-Canadian AI Strategy, and industry sponsors across Canada.

For further information or media enquiries, please contact: media@vectorinstitute.ai