SIGNAL GRIDv0.1

Can a MUD evaluate LLMs? A $99 proof of concept

1 sources1 storiesFirst seen 7/22/2026Score34Mixed Progress
Single Source
CoverageRecencyEngagementVelocityBignessConfidenceClipability
Bigness
34
Coverage
13
Recency
93
Engagement
24
Velocity
0
Confidence
49
Clipability
60
Polarization
0
Claims
5
Contradictions
0
Breakthrough
50

Sentiment Mix

Positive0%
Neutral100%
Negative0%

Geography

North America

Expert Signals

Davisb135

author1 mention

Hacker News

source1 mention

AI-Generated Claims

Generated from linked receipts; click sources for full context.

A $99 proof of concept.

Supported by 1 story

I'm the author of a paper my friends and I wrote after we were curious if a MUD, text games originating in the 1970s, could be used to evaluate LLMs.

Supported by 1 story

We've spent the last several months on nights and weekends running this experiment and writing the paper on just our personal computers with about $99 in API credits.Our experiment did have an interesting leaderboard but even more surprising was the measurements of each LLM.

Supported by 1 story

When we then checked the classifier against a second judge, the per-model agreement between them ranged from 85% to 22%.

Supported by 1 story

The aggregate kappa (0.04 on probe detection) indicated the instrument was noisy without saying which models the noise was hitting.

Supported by 1 story

Related Events

Timeline (1 stories)

Jul 22 05:08 PMFirst
Can a MUD evaluate LLMs? A $99 proof of concept
Hacker News165 engagement

Receipts (1)

Bias Snapshot

Center
Left 0%Center 100%Right 0%
Aggcruciblebench.ai7/22/2026