A Survival Guide for AI Deployment
In 2024, researchers Sean Williams and James Huckle published a devastating benchmark showing that GPT-4 Turbo, Claude 3 Opus, and other leading LLMs scored just 16–38% on questions that average adults answered correctly 86% of the time. The questions were not obscure