← Events
2026 Fall — Technical Paper Reading Group
Read selected AI safety papers on threat models, reward hacking, scheming, evaluations, mechanistic interpretability, and AI control.
About|活動介紹
Read and discuss selected papers on threat models, reward hacking, scheming, safety evaluations, mechanistic interpretability, and AI control.
挑選相關論文進行閱讀與討論,主題涵蓋威脅模型、reward hacking、欺騙與模型謀算、安全評估、機制可解釋性,以及 AI Control。
Organizing Team|帶領團隊
- Track lead: Lily
- Track co-hosts: Zen, Leo, Ted
Schedule|活動時程
| Week | Date & Time | Topic / Event | Host | Co-host |
|---|---|---|---|---|
| W1 | 09/29 | Technical AI Safety Landscape & Threat Models | Zen | Leo |
| W2 | 10/06 | Post-training, Reward Hacking, & Emergent Misalignment | Ted | Leo |
| W3 | 10/13 | Deception, Scheming, & Model Organisms | Leo | Zen |
| W4 | 10/19 | Connection Dinner / Speaker Event | — | — |
| — | 10/27 | Midterm | — | — |
| W5 | 11/03 | Safety Evaluations: Validity, Gaming, & Evaluation Awareness | Leo | Lily |
| W6 | 11/10 | Scalable Oversight | Lily | Zen |
| W7 | 11/17 | Mechanistic Interpretability I: Features, Circuits, & Faithfulness | Zen | Ted |
| W8 | 11/24 | Mechanistic Interpretability II: Probing, Monitoring, & Steering | Lily | Ted |
| W9 | 12/01 | AI Control & Defence in Depth | Ted | Lily |
| W10 | 12/07 | TBD | — | — |
| W11 | 12/14 | Connection Dinner / Speaker Event | — | — |
| — | 12/22 | Final | — | — |
Participation Details|參加資訊
- Regular sessions|固定時間: Every Tuesday / 每週二,19:00–21:00
- Reading group venue|讀書會地點: 臺大資訊工程學系德田館(教室待公布)
Registration|活動報名
- Reading group registration|讀書會報名 — Deadline / 截止日期:2026/09/23
註:讀書會的三個組別採共用報名表單,無論報名一個或多個組別,都只需填寫一次。