惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

T
The Blog of Author Tim Ferriss
I
InfoQ
H
Hackread – Cybersecurity News, Data Breaches, AI and More
aimingoo的专栏
aimingoo的专栏
小众软件
小众软件
有赞技术团队
有赞技术团队
J
Java Code Geeks
Apple Machine Learning Research
Apple Machine Learning Research
大猫的无限游戏
大猫的无限游戏
Engineering at Meta
Engineering at Meta
B
Blog RSS Feed
博客园_首页
Y
Y Combinator Blog
V
Visual Studio Blog
Google DeepMind News
Google DeepMind News
M
MIT News - Artificial intelligence
雷峰网
雷峰网
博客园 - 司徒正美
OSCHINA 社区最新新闻
OSCHINA 社区最新新闻
H
Help Net Security
P
Proofpoint News Feed
B
Blog
云风的 BLOG
云风的 BLOG
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报

WASP

WASP enters a new phase with long-term funding secured | WASP Christian Berger to lead the WASP Graduate School from 2027 | WASP Two WASP researchers awarded ERC Starting Grants | WASP Three projects awarded funding in first joint WASP and WASP-HS call | WASP WASP strengthens Swedish AI research through recruitment of Julian Togelius | WASP Community building summer school gives new WASP PhD students a first introduction | WASP WASP researchers contribute to ECCV 2026 | WASP WASP research helps Sony AI’s table tennis robot decide in milliseconds | WASP WASP researchers receives the Automatica Best Paper Award | WASP The long game: When someone is paid to think ahead  | WASP Alexandre Bartel receives Nordea's Scientific Prize 2026 | WASP Strong WASP presence at ICML 2026 | WASP WARA Public Safety introduces Vinnova Drone Challenge at Data Collection Week ELLIS adds seven new units – one of them in Sweden WASP-affiliated research accepted to CVPR 2026 Learn more about SE.LLMA WASP researchers receive Best Paper Award for advancing safety in AI-based autonomy The data only industry can provide Jialong Li receives award for his work on open source teleoperation Exploring the intersection of society, life sciences and technology Updates in the WARA Ops portal WASF 2026 explores the foundations of neurosymbolic AI Where are WASP alumni today? Statistics from a recent WASP follow‑up Miriah Meyer: “I’m a fangirl of theory” Martin Monperrus Elevated to IEEE Fellow for Advances in AI-Driven Software Engineering Alexandre Proutiere receives ACM SIGMETRICS Achievement Award
WASP students strengthen ties with Mila’s reinforcement l...
aliro35 · 2026-09-10 · via WASP

WASP cluster members in front of the Mila building in Montreal.

WASP PhD students and postdocs from the cluster Sequential decision making and reinforcement learning visited Mila – Quebec AI Institute to exchange ideas and take part in a conference. The trip turned out to be a success when it offered new research perspectives and opportunities to build international connections. Mila is one of WASP’s partner universities, and the visit contributed to strengthening the long-term collaboration between the two research communities.

The visit was organized in August 2026 to strengthen connections between researchers in the WASP cluster and the reinforcement learning community in Montreal. The program combined research presentations, meetings with faculty and PhD students, informal scientific discussions, and participation in Reinforcement Learning Conference 2026.

“The study trip provided valuable opportunities for scientific exchange and networking. Discussions with researchers in Montreal gave the participants feedback and new perspectives on their research, while presentations from the host groups offered insight into current directions in reinforcement learning and machine learning,” says Stefan Stojanovic, the main organizer and PhD student at KTH Royal Institute of Technology.

Research exchange at Mila

During the visit, the group consisting of 11 WASP PhD students and postdocs, met with researchers and PhD students working on reinforcement learning, machine learning, robotics, and related areas. The first part of the program included a meeting with Professor Alex Hernandez-Garcia, who presented his research on machine learning for scientific discovery and introduced GFlowNets, a method with similarities to reinforcement learning that is designed to sample diverse outcomes in proportion to their reward.

The group also met PhD students supervised by Professor Glen Berseth, who co-directs the Robotics and Embodied AI Lab at Mila, and students from Professor Pierre-Luc Bacon’s research group. Their presentations covered topics such as world models, scaling robotic pretraining, adaptive policy priors, and concentration of cumulative rewards in MDPs.

The breadth of topics aligned well with the WASP cluster, which brings together researchers working across several areas of reinforcement learning and sequential decision making.

On Friday, the group joined a larger reinforcement learning meeting organized by Mila. Six WASP PhD students presented their research and received feedback from researchers at Mila and other visiting researchers. The meeting also featured presentations from researchers visiting from institutions including ETH, DeepMind, and the Max Planck Institute.

Perspectives from Reinforcement Learning Conference 2026

The group also participated in Reinforcement Learning Conference 2026, where they attended talks and workshops, presented their work, and connected with the broader international reinforcement learning community.

For participant Jenni Reuben, Industrial Postdoc at KTH Royal Institute of Technology and Research Scientist at Saab Aeronautics, the conference and study trip highlighted several important questions for safe and trustworthy reinforcement learning.

“One key takeaway was that an agent may perform well under normal conditions but still fail when the environment changes. Many discussions therefore focused on robustness, distribution shifts, and how to identify when an agent moves beyond its area of competence,” she says.

She also noted that uncertainty detection becomes valuable only when it leads to an action, such as slowing down, abstaining, transferring control, or activating a safety filter.

“This connection between detecting uncertainty and deciding how to intervene was especially relevant to my own research,” says Reuben.

Reuben also appreciated the format of the conference, where accepted papers were first presented in short oral sessions and then discussed in poster sessions.

“The format made it easier to identify the papers most relevant to me and then follow up with one-to-one discussions with the authors,” she says.

The study trip offered participants new research perspectives, feedback on their own work, and opportunities to build international connections in reinforcement learning, safe autonomous systems, and related areas.

Jenni Reuben
Jenni Reuben, Industrial Postdoc at KTH Royal Institute of Technology and Research Scientist at Saab Aeronautics.

Interested in organizing your own study trip?

Several options are available for PhD students interested in organizing a study trip. Trips can be arranged through a cluster or organized independently as a self-arranged study trip.

Daniel Lawson, PhD from the REAL group at Mila, presenting his work during meeting with WASP visitors.
Daniel Lawson, PhD from the REAL group at Mila, presenting his work during meeting with WASP visitors.
Raghav Bongole, KTH, presenting work at RL group meeting.
Raghav Bongole, KTH, presenting work at RL group meeting.
Jack Sandberg, Mila cluster trip
Jack Sandberg, Chalmers, presenting work at RL group meeting.
Ahmet Balcioglu, Chalmers
Ahmet Balciouglu, Chalmers, presenting work at RL group meeting.
Gabriele Calzolari, Luleå University
Gabriele Calzolari, Luleå University of Technology, presenting work at RL group meeting.
Mika Persson, Chalmers, presenting work at RL group meeting.
Mika Persson, Chalmers, presenting work at RL group meeting.
David Abel, Deepmind, RLC
David Abel, Deepmind, gave a talk titled “Where is learning?” during the Continual RL workshop at RLC 2026.
All attendees of RLC 2026 had the opportunity to enjoy Echo, a spectacular show by Montreal’s Cirque du Soleil.
All attendees of RLC 2026 had the opportunity to enjoy Echo, a spectacular show by Montreal’s Cirque du Soleil.

Published: September 10th, 2026

[addtoany]