Value induction reshapes how LLMs behave after post-training on traits like empathy and helpfulness. Inducing one value can unpredictably alter behavior on related values, creating complex trade-offs. Efforts to make models safer or more useful can backfire, increasing sycophancy or addictiveness through the language patterns the model adopts.
Opening Kapyn…