← Events
2026 Fall — Fundamental Reading Group
Explore AI alignment, safer training, evaluations, interpretability, and AI control through a semester of guided discussions.
About|活動介紹
Explore AI alignment, safer training, evaluations, interpretability, and AI control through a semester of guided discussions.
從 AI alignment 出發,循序討論 RLHF、可擴展監督、安全評估、可解釋性與 AI Control,建立對 AI 安全技術與挑戰的整體認識。
Organizing Team|帶領團隊
- Track leads: Jasper, Jonathan
- Track co-host: Teddy
Schedule|活動時程
| Week | Date & Time | Topic / Event | Host | Co-host |
|---|---|---|---|---|
| W1 | 09/28 | AI alignment: Why can’t we build safe AI? | Teddy | Jasper |
| W2 | 10/05 | Training Safer Models Part I: RLHF | Jasper | Teddy |
| W3 | 10/12 | Training Safer Models Part II: Scalable Oversight | Teddy | Jonathan |
| W4 | 10/19 | Connection Dinner / Speaker Event | — | — |
| — | 10/26 | Midterm | — | — |
| W5 | 11/02 | Detecting Danger: Evaluations & Red teaming | Jonathan | Teddy |
| W6 | 11/09 | Understanding AI Part I: Interpretability | Jonathan | Jasper |
| W7 | 11/16 | Understanding AI Part II: Interpretability in practice | Teddy | Jonathan |
| W8 | 11/23 | Minimising Harm Part I: AI Control | Jasper | Teddy |
| W9 | 11/30 | Minimising Harm Part II: AI Control – Chain of Thought | Teddy | Jasper |
| W10 | 12/07 | Speaker Event | — | — |
| W11 | 12/14 | Connection Dinner | — | — |
| — | 12/21 | Final | — | — |
課程參考:BlueDot Impact — Technical AI Safety
Participation Details|參加資訊
- Regular sessions|固定時間: Every Monday / 每週一,19:00–21:00
- Reading group venue|讀書會地點: 臺大資訊工程學系德田館(教室待公布)
Registration|活動報名
- Reading group registration|讀書會報名 — Deadline / 截止日期:2026/09/23
註:讀書會的三個組別採共用報名表單,無論報名一個或多個組別,都只需填寫一次。