← Events
Technical Workshop Upcoming

NTU AI Safety Workshop 2026

A two-day hands-on technical workshop covering AI Safety fundamentals, jailbreaks, emergent misalignment, RLHF/RLAIF, and a model attack-and-defense hackathon.

workshoptechnicaljailbreakemergent-misalignmentrlhfrlaifhackathonupcoming

About the Workshop|活動簡介

NTU AI Safety works to raise awareness of AI safety in Taiwan and connect the local community with international networks, research resources, and development opportunities. Following our first workshop, which introduced participants to AI safety and welcomed many of them into the community, we are returning this year with the second NTU AI Safety Workshop.

As AI models advance rapidly, identifying potential risks, reducing unintended behavior, and developing safer training methods have become key priorities for research and industry teams worldwide. This two-day technical workshop takes a hands-on approach, guiding participants from AI safety fundamentals to model attacks, defenses, and safer training methods.


NTU AI Safety 致力於提升台灣社群對 AI Safety 的認識,並連結國內外社群、研究資源與發展機會。繼第一屆工作坊帶領參與者認識 AI Safety 並加入社群後,我們將於今年舉辦第二屆 NTU AI Safety Workshop

隨著 AI 模型快速發展,如何辨識模型的潛在風險、降低非預期行為,並建立更安全的訓練方法,已成為全球研究與產業團隊共同關注的重要課題。本次為期兩天的技術工作坊將從實作出發,帶你從 AI Safety 基礎概念,一路探索模型攻防與安全訓練方法。


What We’ll Cover|工作坊內容

  • AI Safety Fundamentals: Understand the core problems and research directions in technical AI safety, and why they matter for today’s models
  • Potential Model Safety Risks: Explore jailbreaks, emergent misalignment, and other forms of unintended model behavior
  • Training Safer Models: Learn the concepts, applications, and limitations of safety training methods such as RLHF and RLAIF
  • Jailbreak & Defend Hackathon: Analyze model vulnerabilities and design defense strategies through a small-scale attack-and-defense exercise

  • AI Safety 基礎:了解 Technical AI Safety 的核心問題與研究方向,以及這些議題為何與現今模型密切相關
  • 模型潛在安全風險:認識 Jailbreak、Emergent Misalignment,以及模型可能出現的非預期行為
  • 訓練更安全的模型:了解 RLHF、RLAIF 等安全訓練方法的概念、用途與限制
  • Jailbreak & Defend Hackathon:透過小型攻防實作,分析模型弱點並設計防禦策略

Event Details|活動資訊

  • Dates: September 12–13, 2026 (Saturday–Sunday)
  • Time: 10:00–18:00 each day
  • Venue: Room R601, CSIE-DerTian Hall, National Taiwan University
  • Fee: Free of charge, with lunch provided
  • Sponsors: Pathfinder and BlueDot Impact
  • Speakers & TAs: 呂柏頤 (網媒所博五)、陳妍姍 (資工所碩二)、謝子涔(電機所碩畢)、胡皓雍(電信所碩二)、王瑭毅 (電子所碩士)、蔡侑宸(數學五)、蔡孟衡 (資工四)、張嘉泰 (資工四)

  • 日期:2026 年 9 月 12 日(六)至 9 月 13 日(日)
  • 時間:每日 10:00–18:00
  • 地點:國立臺灣大學德田館 R601
  • 費用:全額免費,並提供午餐
  • 贊助:Pathfinder、BlueDot Impact
  • 講者與 TAs:呂柏頤 (網媒所博五)、陳妍姍 (資工所碩二)、謝子涔(電機所碩畢)、胡皓雍(電信所碩二)、王瑭毅 (電子所碩士)、蔡侑宸(數學五)、蔡孟衡 (資工四)、張嘉泰 (資工四)

Who Should Join?|誰適合參加?

This workshop is designed for students and industry professionals with a basic background in ML and Python who want to explore or pursue AI safety research. Before applying, please make sure you:

  • Have experience programming in Python
  • Have a basic understanding of deep learning
  • Can attend both full days of the workshop
  • Are curious about AI safety or interested in pursuing further learning and research in the field

本次工作坊適合具備基礎 ML/Python 背景,並希望深入了解或投入 AI Safety 研究的學生與業界人士。報名前請確認你:

  • 具備 Python 程式設計經驗
  • 對深度學習有基本認識
  • 能完整參與兩天課程
  • 對 AI Safety 感到好奇,或有意進一步投入相關學習與研究

What You’ll Gain|你將獲得什麼?

  • Understand the key questions currently shaping AI safety
  • Build foundational knowledge and hands-on experience in technical AI safety
  • Gain practical experience with model jailbreaks and defense methods
  • Explore whether you want to continue learning, join the community, or pursue long-term research

  • 了解目前 AI Safety 關注的核心問題
  • 建立對 Technical AI Safety 領域的基礎認識與實作經驗
  • 親自體驗模型 Jailbreak 與防禦方法
  • 探索是否要進一步學習、參與社群或投入長期研究

How to Register|報名方式與重要時程

  1. Complete the application form by 11:59 PM on August 25, 2026.
  2. Join the NTU AI Safety Discord for the latest event information and updates.
  3. Check the email address provided in your application. As places are limited, the organizing team will conduct a brief review and send acceptance notifications by August 31, 2026.
  4. If accepted, please reserve both full days and arrive on time for the workshop.

  1. 2026 年 8 月 25 日 23:59前填寫報名表單
  2. 加入 NTU AI Safety Discord,取得最新活動資訊與後續通知。
  3. 留意報名時填寫的 Email;因名額有限,主辦團隊將進行簡單審核,並於 2026 年 8 月 31 日前寄發錄取通知。
  4. 錄取後請預留兩天完整時段,準時參與工作坊。

We look forward to exploring how to make AI systems safer with you at the workshop!
期待在工作坊與你一起探索如何讓 AI 系統變得更安全!