Interpreting and Steering LLM Agents for Social Simulations
We compare prompting, sparse autoencoders, and linear probes for interpreting and steering LLM agents in social simulations. SAEs help inspect internal features, probes offer calibrated control, and stronger prompting is competitive or better on some tasks.