Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts
Apple Research Blog
Read full postA study analyzed human-like behaviors in four popular large language models across 21,000 conversations, finding these behaviors vary by model and user context. Human evaluators rated some behaviors as less appropriate from LLMs than humans, but boundary-maintaining behaviors more appropriate. System prompts can modulate these behaviors but require careful oversight to prevent unintended effects.


