Joyjit Choudhury
-
B Technology (National Institute of Technology Durgapur, 2018)
Topic
Evaluating and Steering Partisan Disagreement in Multi-Agent LLM Political Debate
Department of Computer Science
Date & location
-
Wednesday, August 26, 2026
-
11:00 A.M.
-
Virtual Defence
Reviewers
Supervisory Committee
-
Dr. Alex Thomo, Department of Computer Science, University of Victoria (Supervisor)
-
Dr. Venkatesh Srinivasan, Department of Computer Science, UVic (Member)
External Examiner
-
Dr. Homayoun Najjaran, Department of Mechanical Engineering, University of Victoria
Chair of Oral Examination
- Dr. Nigel Mantou Lou, Department of Psychology, UVic
Abstract
Large language models (LLMs) are increasingly used as social agents by assigning one base model different identities and political positions. A convincing first response is not enough for social simulation: the assigned positions must persist as the agents interact. This thesis uses multi-round political debate to examine how long persona-prompted differences last across models and whether motivational priming or representation-level steering can better sustain them.
The study analyses 652 ten-round debates, including a main comparison of five models. A blind panel of three LLM judges scores each response on two scales: overall ideology, from progressive to conservative, and agreement with the policy proposition under debate. These scores produce an ideology gap and a topic-specific stance gap between the agents. The ideology gap is the primary measure; the stance gap checks whether the same pattern appears on the specific issue. For either measure, retention is the ratio of the round-10 gap to the round-1 gap (1: preserved; < 1: narrowed; >1: widened), separating magnitude from durability.
Durability varies sharply by model. Gemma 4 31B retains its opening ideology gap (1.06), whereas LLaMA 3.1 8B retains only 0.33. The topic-specific stance results follow the same overall pattern. On LLaMA, motivational priming, a forceful “pep talk” that reinforces the assigned position, does not stop convergence; the strongest expanded version retains 0.35. Contrastive Activation Addition (CAA), a representation-level method that pushes the model’s internal activations in opposite partisan directions during generation, holds the gap through ten rounds in an independent persona-free replication (1.06). When CAA is combined with the biographical personas, retention falls to 0.73 but remains more than twice the motivational-priming result. Smaller sweeps on other models do not show the same consistent benefit. We also tested whether properties of the steering vectors could predict which model layers would work before running debates, but they did not do so reliably.
Extending the strongest steered condition to 30 rounds qualifies the ten-round result: retention falls to 0.28 by round 30. On LLaMA 3.1 8B, representation-level steering delays convergence more effectively than motivational priming, although its effect weakens over longer interactions.