Localizing and Editing Knowledge in LLMs with Peter Hase

EPISODE 679

|

APRIL 8, 2024

Watch

Follow

Share

About this Episode

Today we're joined by Peter Hase, a fifth-year PhD student at the University of North Carolina NLP lab. We discuss "scalable oversight", and the importance of developing a deeper understanding of how large neural networks make decisions. We learn how matrices are probed by interpretability researchers, and explore the two schools of thought regarding how LLMs store knowledge. Finally, we discuss the importance of deleting sensitive information from model weights, and how "easy-to-hard generalization" could increase the risk of releasing open-source foundation models.

About the Guest

Peter Hase

University of North Carolina

Connect with Peter

Resources

Related Topics

Cover: TWIML Presents: Large Language Models

Large Language Models

Cover: TWIML Presents: Data-Centric AI

Data-Centric AI