A new study systematically analyzes human-like behaviors in LLMs using 21,000 LLM-as-a-judge and human evaluations. It examines prevalence, user factors, and system prompt controllability, aiming to help developers decide when models should exhibit such behaviors. The findings give practitioners empirical grounding for shaping model persona and boundaries.
Opening Kapyn…