← Events
Speaker Event Past

Tony Wu: Reading and Safeguarding Machine Minds

Exploring technical frontiers in AI safety and interpretability|探索 AI Safety、interpretability 與 Chain-of-Thought monitorability 的技術前沿。

speaker-eventtechnicalinterpretabilitychain-of-thoughtpast

LinkedIn Post: Event announcement

About the Talk|活動簡介

In the second speaker event of the 2026 NTU AI Safety Reading Group, Tung-Yu (Tony) Wu will explore technical frontiers in AI safety and interpretability. The talk will focus on new challenges introduced by increasingly capable AI agents and growing concerns around the faithfulness and monitorability of Chain-of-Thought reasoning.

Tony will also introduce emerging interpretability directions—including LatentQA-style methods and circuit discovery—and discuss recent progress, current limitations, underexplored research questions, and areas where future breakthroughs may emerge.


在 2026 NTU AI Safety Reading Group 的第二場講者活動中,Tung-Yu(Tony)Wu 將分享 AI Safety 與 interpretability 的重要技術前沿,特別聚焦於能力日益增強的 AI agents 帶來的新挑戰,以及 Chain-of-Thought 的 faithfulness 與 monitorability 問題。

Tony 也將介紹 LatentQA 類型方法、circuit discovery 等新興 interpretability 方向,並討論近期進展、現有限制、仍被低估的研究問題,以及未來可能出現突破的領域。

Speaker|講者

Tony Wu is an incoming DPhil student in Engineering Science at the University of Oxford, pursuing research in AI safety and interpretability. He graduated from National Taiwan University with a bachelor’s degree in Electrical Engineering and a double major in Economics. His research collaborations have included teams and researchers at NTU, Oxford, MIT, AI2, UCSD, Mila, MPI, and the University of Toronto, and his work has appeared at venues including ICML and ICLR.

Tony Wu 即將於牛津大學工程科學系攻讀 DPhil,研究方向為 AI Safety 與 Interpretability。他畢業於國立臺灣大學電機工程學系,並雙主修經濟學;曾與 NTU、Oxford、MIT、AI2、UCSD、Mila、MPI 與多倫多大學等機構的研究團隊合作,研究成果亦發表於 ICML、ICLR 等頂尖會議。

Topics|分享內容

  • Safety challenges posed by increasingly capable AI agents|高能力 AI agents 帶來的安全挑戰
  • Chain-of-Thought faithfulness and monitorability|Chain-of-Thought 的 faithfulness 與 monitorability
  • LatentQA-style interpretability methods|LatentQA 類型的 interpretability 方法
  • Circuit discovery|Circuit discovery(神經網路迴路探索)
  • Underexplored and promising research directions|尚未被充分探索、具潛力的研究方向

Event Details|活動資訊

  • Date|日期: June 1, 2026 (Monday)|2026 年 6 月 1 日(一)
  • Time|時間: 7:00 PM (UTC+8)
  • Series|系列: 2026 NTU AI Safety Reading Group speaker event|2026 NTU AI Safety 讀書會講者活動