Qwen plays World of Warcraft: a 27B open-weights model quests via MCP — Reddit's verdict
A Redditor wired an open-weights Qwen model into a private World of Warcraft server through an MCP bridge — text-only, no vision — and filmed it questing. r/LocalLLaMA is delighted, and has notes.

On Sunday night, an open-weights language model logged into Azeroth. Not metaphorically — a Qwen 3.8 27B model, running on a single RTX 3090, is playing World of Warcraft on a private server through a Model Context Protocol bridge, and the 84-second shaky-cam video of it questing has r/LocalLLaMA equal parts delighted and unimpressed, in all the right ways.

The setup#
The post, from u/professormunchies on r/LocalLLaMA on Sunday evening, is admirably unglamorous about how it works. The poster runs a private WoW server, forked the old “wowser” browser-based WoW client, and built an MCP (Model Context Protocol) server over websockets so the model can drive the game directly. There is no vision involved at all: the model receives a text-only description of the world state plus a set of tools — find nearby objects, look, fight, use spells — and issues actions from there.
The hardware is the punchline the LocalLLaMA crowd loves: the model is a Qwen 3.8 27B GPTQ W4A16 quantization running on one RTX 3090 on a separate machine, while the WoW server ticks away on an AMD Ryzen 5 5600G. The MCP bridge, by the poster's own admission, was “literally slopped together today.” The project lives at jankcraft.xyz.

No eyes, two seconds a move#
What makes this interesting as an AI demo rather than a party trick is the constraint set. The model never sees the screen — it reasons over structured text describing what's around it, then picks tools. In the video, the character ambles, targets, and casts with the deliberation of a player reading a quest log through a straw. One commenter clocked the pace at roughly two seconds per action and doubted the setup could ever “do the Naxx dance” — the fast, precise movement old raid content demands. Fair. This is not a speedrunner.
The video itself became a minor joke in the thread: the top-voted complaint is that the poster did all this engineering and then filmed the monitor with a shaky phone camera instead of recording the screen. (In his defense, some demos are born in an afternoon.) The clip is 84 seconds of exactly what it says on the tin: a language model, playing World of Warcraft, badly but genuinely.
The community verdict#
r/LocalLLaMA's reaction is the real story here, and it's warm. The thread's running joke — “looking forward to ‘qwen folds my laundry’” — drew the perfect reply: “qwen folds my laundry, but only after a 40 step plan and asking you to approve each sock.” The most substantive praise called it “a better benchmark than just benchmarks”: a persistent world with quests, combat, and navigation is a harder test of long-horizon tasking than most static evals.
The corrections arrived on schedule, too. The most important one, posted within minutes: this can't run on Blizzard's official servers, because it requires reading internal game state that the official client won't expose — private server only. Another commenter raised the botting specter (ten thousand automated farmers), and got the thread's collective “please don't release this publicly.” And when someone asked why you'd automate a human out of a video game at all, the poster's answer was the whole point: it's a benchmark for complex LLM tasking — a deliberate riff on the “Pokémon benchmarks” that have the community training agents on Game Boy classics. (We covered the latest of those, a full Pokémon Red remake built by Opus 5.5, earlier this week.)
One commenter offered the obligatory deflation: AI has been playing WoW for twenty years without any powerful graphics card — true of scripted bots, and exactly the distinction the thread is drawing. A scripted bot executes a plan; this thing has to make one up from a text description of the world, two seconds at a time.
Why games keep becoming the benchmark#
There's a reason this keeps happening — first Pokémon, now Azeroth. Games are the rare evaluation environment that is simultaneously open-ended, legible, and brutally honest: the character either reaches the quest objective or it doesn't, and everyone watching can see which. You can't prompt-engineer your way past a boss fight. As agent benchmarks go, a twenty-year-old MMO is cheaper than a robotics lab and more honest than a multiple-choice test. It's the same impulse behind the community-run physical benchmarks we've covered, like five AIs engineering 3D-printed bridges and letting physics grade them.
The open-weights angle matters here too. This isn't a frontier lab's demo reel; it's a 27-billion-parameter model anyone can download, quantized to fit on a single consumer GPU from 2020, driving game actions through an open protocol. The gap between “frontier capability” and “weekend project” keeps compressing — we've been tracking the local-model side of that story closely, including a full setup guide for this exact model family.
What to watch#
Three things. First, whether the poster turns jankcraft.xyz into a real benchmark — a standardized WoW questing suite for open models would be genuinely useful, and the thread is already asking for it. Second, whether anyone gets past the two-seconds-per-action ceiling into content that demands real-time reactions; that's the difference between a novelty and a capability. Third, the botting question: the same MCP bridge that makes a fun demo makes an industrial farming operation, and the community knows it. The “please don't release this publicly” comments are doing real work in that thread — informal norms are the only guardrail this particular project has.
For now, though: an LLM is out there in Azeroth, reading the world as text, taking two seconds to decide to cast a spell, while a phone camera wobbles. The future arrives in strange aspect ratios.
Sources
- “Qwen plays World of Warcraft” — u/professormunchies, r/LocalLLaMA (Reddit, September 27, 2026)
- jankcraft.xyz — the project's site