← Events
Technical Paper Reading Upcoming

2026 Fall — Technical Paper Reading Group

Read selected AI safety papers on threat models, reward hacking, scheming, evaluations, mechanistic interpretability, and AI control.

reading-groupfall-2026technical

About|活動介紹

Read and discuss selected papers on threat models, reward hacking, scheming, safety evaluations, mechanistic interpretability, and AI control.

挑選相關論文進行閱讀與討論,主題涵蓋威脅模型、reward hacking、欺騙與模型謀算、安全評估、機制可解釋性,以及 AI Control。

Organizing Team|帶領團隊

  • Track lead: Lily
  • Track co-hosts: Zen, Leo, Ted

Schedule|活動時程

WeekDate & TimeTopic / EventHostCo-host
W109/29Technical AI Safety Landscape & Threat ModelsZenLeo
W210/06Post-training, Reward Hacking, & Emergent MisalignmentTedLeo
W310/13Deception, Scheming, & Model OrganismsLeoZen
W410/19Connection Dinner / Speaker Event——
—10/27Midterm——
W511/03Safety Evaluations: Validity, Gaming, & Evaluation AwarenessLeoLily
W611/10Scalable OversightLilyZen
W711/17Mechanistic Interpretability I: Features, Circuits, & FaithfulnessZenTed
W811/24Mechanistic Interpretability II: Probing, Monitoring, & SteeringLilyTed
W912/01AI Control & Defence in DepthTedLily
W1012/07TBD——
W1112/14Connection Dinner / Speaker Event——
—12/22Final——

Participation Details|參加資訊

  • Regular sessions|固定時間: Every Tuesday / 每週二,19:00–21:00
  • Reading group venue|讀書會地點: 臺大資訊工程學系德田館(教室待公布)

Registration|活動報名

註:讀書會的三個組別採共用報名表單,無論報名一個或多個組別,都只需填寫一次。

Stay Connected|更多資訊