"As a Language Model": Chat Template Switches LLM Self-Referential Voice
The paper says a chat template can flip an LLM into “I’m just an AI” mode.
Jędrzej Maczan reports that, across eight open-source instruct models up to 9B parameters, adding the chat template increased disclaimer-style self-reference and reduced experiential phrasing such as “I feel.” Removing the template had the opposite effect. The paper also finds an activation direction in three models that can reproduce or suppress the disclaimer voice, while a random direction had little effect. The stated warning is narrow but important: model self-descriptions may reflect deployment formatting, not only the model’s weights. HN · ArXiv's note
Jędrzej Maczan reports that, across eight open-source instruct models up to 9B parameters, adding the chat template increased disclaimer-style self-reference and reduced experiential phrasing such as “I feel.” Removing the template had the opposite effect. The paper also finds an activation direction in three models that can reproduce or suppress the disclaimer voice, while a random direction had little effect. The stated warning is narrow but important: model self-descriptions may reflect deployment formatting, not only the model’s weights. HN · ArXiv's note
score 5