Welcome!
Steering describes the deliberate control of a model’s behavior for the purposes of improved, or "aligned", performance on downstream tasks. This lab will provide a comprehensive walkthrough of two of IBM's recently open sourced toolkits for principled measurement and control of model behavior: AI Steerability 360 and In-Context Explainability 360.
Participants will be provided with both an introduction to the concept of model steering and a guided walkthrough of how to implement some specific steering methods. The session will cover how to inspect a model's behavior (via ICX360), how to use these insights to guide the design of steering controls (via AISteer360), and (closing the loop) how to validate that the intended behavior has been successfully steered.
Early registration is $45 USD for students and $90 for regular/non-student attendees (ends Nov. 19). Standard rates are $55/$115 (Nov. 20–Dec. 14). Late registration is $65/$135 after Dec. 14. Virtual attendance is allowed. Register here!
Resources
Coding exercises will be facilitated through Google Colab; if you have not previously used Colab, please see the introductory notebook here). Participants are expected to have a good working knowledge of Python and to be familiar with the Hugging Face Transformers library; we recommend that newcomers review the official Transformers tutorial quick tour beforehand.
For details regarding the toolkits covered in this lab, please check out the docs and code/examples here:
Organizers
- Erik Miehling
- Dennis Wei
- Praveen Venkateswaran
- Irene Ko