Can a MUD evaluate LLMs? A $99 proof of concept
Sentiment Mix
Geography
Expert Signals
Davisb135
author • 1 mention
Hacker News
source • 1 mention
AI-Generated Claims
Generated from linked receipts; click sources for full context.
A $99 proof of concept.
Supported by 1 story
I'm the author of a paper my friends and I wrote after we were curious if a MUD, text games originating in the 1970s, could be used to evaluate LLMs.
Supported by 1 story
We've spent the last several months on nights and weekends running this experiment and writing the paper on just our personal computers with about $99 in API credits.Our experiment did have an interesting leaderboard but even more surprising was the measurements of each LLM.
Supported by 1 story
When we then checked the classifier against a second judge, the per-model agreement between them ranged from 85% to 22%.
Supported by 1 story
The aggregate kappa (0.04 on probe detection) indicated the instrument was noisy without saying which models the noise was hitting.
Supported by 1 story
Related Events
$1.5bn piracy settlement approved against AI giant Anthropic - businesscloud.co.uk
LLMs • 7/22/2026
Judge Approves a $1.5B Anthropic Settlement Over Books Used to Train Claude - Broadband Breakfast
LLMs • 7/22/2026
Anthropic to pay $1.5 billion after misusing books to train AI chatbot - Gamereactor UK
LLMs • 7/22/2026
Russian hackers are now gaming AI like ChatGPT and Claude - The World from PRX
LLMs • 7/22/2026
OpenAI and Anthropic Ramp Up DC Lobbying 23% as AI Rules Loom - The Tech Buzz
LLMs • 7/22/2026
Causality Chain
Preceded By
Led To