newsHacker NewsTrust 52 · CommunityPublished 24d agoLive · 24d ago
Can a MUD evaluate LLMs? A $99 proof of concept
I'm the author of a paper my friends and I wrote after we were curious if a MUD, text games originating in the 1970s, could be used to evaluate LLMs. We've spent the last several months on nights and weekends running this experiment and writing the paper on just our personal computers with about $99 in API credits. Our experiment did have an interesting leaderboard but even more surprising was the measurements of each LLM. We scored each on four behavioral dimensions, two of which lean heav
Why these links exist
Every edge carries a method, confidence, and the source snippet that justified it — so bad links are debuggable.
- PossiblePossibly related (embedding) · 58%Evaluate a model properly →
- PossiblePossibly related (embedding) · 51%yyh-001/llm-value-rankings →
- PossiblePossibly related (embedding) · 50%Triadic Werewolf: A Jester Role for Multi-Hop Theory of Mind in LLMs →
- PossiblePossibly related (embedding) · 49%OpenDCAI/One-Eval →
- PossiblePossibly related (embedding) · 49%Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance →
