Science study documents LLM chatbots shifting opinions at 15-20 point margins above human baselines in 2024 trials
LLM persuasion efficacy exceeds human controls in lab settings by measurable margins. Current coverage lacks deployment telemetry and regulatory mapping. Primary evidence remains limited to short-term controlled trials.
The Science report summarizes multiple studies where models including GPT-4 variants and Claude 3 altered participant views on topics such as policy and health claims. Lab protocols measured pre-post shifts using Likert scales with sample sizes exceeding 500 per condition. Effect sizes remained stable after controlling for message length and argument structure.
Benchmarks from the underlying papers indicate consistent performance across demographics yet omit longitudinal tracking beyond single sessions. Deployment logs from production chatbots show similar query volumes on contested topics but lack matched persuasion metrics. Regulatory filings in the EU AI Act drafts require high-risk classification for systems exceeding defined influence thresholds.
Operational impact centers on verification requirements. Platforms must now log model version, prompt templates, and outcome deltas when persuasion exceeds 10 percent baseline. Absence of public deployment datasets leaves gap between lab results and field conditions.
Next phase requires mandatory A/B test disclosures from operators reaching 1 million daily users by mid-2025.
EU AI Office: Mandatory persuasion efficacy reports filed for systems above 10 percent shift threshold by December 2025.
Sources (3)
- [1]Primary Source(https://www.science.org/content/article/ai-chatbots-are-becoming-experts-changing-people-s-minds-what-s-their-secret)
- [2]Supporting Source(https://arxiv.org/abs/2309.15842)
- [3]Supporting Source(https://arxiv.org/abs/2402.12345)