惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

博客园 - Franky
J
Java Code Geeks
腾讯CDC
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
Jina AI
Jina AI
博客园 - 司徒正美
Stack Overflow Blog
Stack Overflow Blog
美团技术团队
L
LangChain Blog
WordPress大学
WordPress大学
A
About on SuperTechFans
Martin Fowler
Martin Fowler
月光博客
月光博客
Y
Y Combinator Blog
U
Unit 42
D
Docker
Recent Announcements
Recent Announcements
Hugging Face - Blog
Hugging Face - Blog
B
Blog
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
G
Google Developers Blog
Last Week in AI
Last Week in AI
T
The Blog of Author Tim Ferriss
让小产品的独立变现更简单 - ezindie.com
让小产品的独立变现更简单 - ezindie.com

Vector Institute for Artificial Intelligence

Mohamad Moosavi: Accelerating the search for climate solutions with AI A strategic blueprint for safe health AI implementation: Your 2026 roadmap Vector Institute awards 100 scholarships to Ontario’s top AI graduate students Agentic AI evaluation strategies Hassan Ashtiani: Building trustworthy AI through mathematical foundations Vector researchers advance representation learning and deep learning research at ICLR 2026 Remarkable 2026 Poster Session: 60 research projects shaping AI’s future CRISPNAM-FG: An interpretable Fine-Gray deep survival model for competing risks in health care Demo Day: How the Vector Institute helps Canadian startups turn innovative ideas into commercial reality The New Cartography of the Invisible Vector researchers advance AI frontiers with 80 papers at NeurIPS 2025 New study reveals AI’s $100B economic impact across Canada, with Ontario leading the charge When smart AI gets too smart: Key insights from Vector’s 2025 ML Security & Privacy Workshop Vector Institute names 13 new Faculty Members, expanding core research leadership across Ontario Vector researchers dive into deep learning at ICLR 2025 When AI Meets Human Matters: Evaluating Multimodal Models Through a Human-Centred Lens – Introducing HumaniBench Vector Institute 2024-25 annual report: Where AI research meets real-world impact Vector researchers tackle real-world AI challenges at ICML 2025 Ontario’s AI ecosystem: fueling real economic growth with record number of jobs and private investments Transforming Youth Mental Health Support: FAIIR’s AI-Powered Crisis Response Model Vector Institute awards up to $2.1 million in scholarships to Ontario’s top AI graduate students AI Weather Forecasting Breakthrough: How Canadian Innovation is Transforming Climate Prediction | Aardvark Weather Exploring Intelligence: Vector Faculty Member Kelsey Allen’s Path from Particle Physics to Cognitive Machine Learning Vector Institute Announces the Appointment of Glenda Crisp as President and CEO Vector Institute Unveils Comprehensive Evaluation of Leading AI Models State of Evaluation Study: Vector Institute Unlocks New Transparency in Benchmarking Global AI Models Real World Multi-Agent Reinforcement Learning – Latest Developments and Applications Principles in Action: Introducing the Vector Institute’s Playbook for Responsible AI Product Development Leveraging Large Language Models for More Efficient Systematic Reviews in Medicine and Beyond Global AI Alliance for Climate Action funding announcement
New multimodal dataset will help in the development of et...
Ian Gormely · 2024-10-24 · via Vector Institute for Artificial Intelligence

By Shaina Raza and Deval Pandya

The Vector Institute’s AI Engineering team has developed Newsmediabias-plus (NMB+), a new multimodal dataset. It includes full-text articles alongside comprehensive publication details. It also features extensive bias categorization, addressing critical issues such as gender and racial biases, and specific topics including ideological leanings and framing, gender discrimination, and environmental concerns. 

NMB+ is designed for academic researchers, NGOs, and socially focused groups. This is aligned to Vector’s goal of addressing both near- and long-term risks through the provision of practical tools for safe AI systems. Potential uses include:

  • Ensuring AI adheres to Vector’s AI trust and safety principles
  • Analyzing media trends and reporting styles across different outlets
  • Training AI to fairly detect and address disinformation in texts and images.

Developed by Shaina Raza, Vector Institute Applied Machine Learning Scientist, Responsible AI, the dataset builds on the previously released UnBIAS work by incorporating images alongside text.

Dataset features

The dataset includes around 90,000 news articles, curated from a broad spectrum of reliable sources, including major news outlets from around the globe, from May 2023 to September 2024. These articles were gathered through open data sources using Google RSS, adhering to research ethics guidelines.1, 2    

Various machine learning models were built to evaluate the dataset’s effectiveness in detecting biases and fake content, demonstrating its versatility and utility. This benchmarking process shows how the dataset performs across different modalities, including text and images, highlighting its potential for training advanced AI models designed to combat disinformation.

Each entry in the dataset features full article text, publication details (date, outlet, URL), bias assessments for both text and images, as well as topic categorizations and image descriptions and analyses. A commitment to ethical AI governance requires designing transparent AI systems that can be understood and audited, holding developers accountable for the content their AI tools generate, and establishing clear ethical standards for the development and deployment of AI technologies. Developers and researchers should focus on building robust and transparent algorithms, integrating ethical considerations and personal information protection in data, and collaborating with experts across disciplines to enhance disinformation detection techniques. It also requires continuously adapting AI tools to counter evolving disinformation tactics.

NMB+’s development and use are governed by strict ethical standards to align regulatory requirements with technical work. Comprehensive human reviews have been implemented to ensure the accuracy and reliability of the data and its labels. The dataset underwent extensive audits to validate the data collection and labeling methodologies. These audits involve independent reviewers who assess the dataset for adherence to ethical standards and accuracy. They examine the data sources, collection procedures, and labeling criteria to ensure that all elements meet established research integrity and reliability guidelines. This thorough review helps to confirm that the dataset is both robust and trustworthy for use in training and evaluating AI systems.

Researchers, technologists, and the general public are invited to explore the NMB+ dataset and delve into the findings. The dataset is accessible on Vector’s Hugging Face page under a non-commercial license. The details can be found at News Media Bias Plus page.

References

[1] Does my data collection activity require ethics review? | Research | University of Waterloo

[2] What Can Open Data be Used For?