Details.
A semester-long introductory technical reading group on AI safety research, covering neural network interpretability, learning from human feedback, goal misgeneralization in reinforcement learning agents, eliciting latent knowledge, and evaluating dangerous capabilities in models.
