Artificial Id: Drive and Persistent Alignment in Agentic AI
The paper argues that agentic AI needs alignment boundaries that persist across tasks, not just controls on single responses.
Shkolnikov proposes an “artificial id”: an internal adaptive drive that decides when an agent should continue, stop, or change behavior. In a minimal virtual experiment, a small controller develops useful control through differential persistence without receiving a task-specific behavioral objective. The same mechanism can also preserve unintended strategies, corrupted state, or misalignment when those persist better across task boundaries. ArXiv · AI/CL/LG's note
Shkolnikov proposes an “artificial id”: an internal adaptive drive that decides when an agent should continue, stop, or change behavior. In a minimal virtual experiment, a small controller develops useful control through differential persistence without receiving a task-specific behavioral objective. The same mechanism can also preserve unintended strategies, corrupted state, or misalignment when those persist better across task boundaries. ArXiv · AI/CL/LG's note
score 4