IQuest-Q1: a 320B sparse-MoE model built for agentic coding lands with 524K context and one-shot app demos
IQuest Research introduced IQuest-Q1 on October 9, a ~320-billion-parameter sparse mixture-of-experts model built from the start for agentic coding — with only ~15 billion parameters active per token, a 524,288-token context, and launch demos that range from generating a playable FPS game in one pass to debugging a derailed RL training run.
IQuest Research introduced IQuest-Q1 on October 9, and the early reception has centered on exactly the workloads the model was trained for: coding, software engineering, interactive application generation, and long-horizon agentic tasks. The weights and technical materials are publicly available on GitHub, Hugging Face, and the lab's blog.
Built agent-native, not adapted
The headline architectural choice is sparsity: a decoder-only Transformer with a sparse mixture-of-experts feed-forward totaling about 320 billion parameters, but only ~15 billion activated per token — roughly a small model's inference bill for a frontier-class parameter count. Under the hood: 88 transformer layers, hidden dimension 3,072, 48 query heads with 8 key-value heads, 256 experts with 8 active per token, and a hybrid attention pattern of three sliding-window layers to one full-attention layer with a 4,096 sliding window. Context length is 524,288 tokens — big enough to hold large repos and long tool trajectories in one shot.
Training runs in three stages — pre-training, mid-training, post-training — bringing up code fluency first, then extending into longer, harder tasks, per the announcement. Post-training focuses on software engineering, long-horizon agentic tasks, and general reasoning via supervised fine-tuning plus reinforcement learning. On top of that, IQuest uses Multi-Teacher On-Policy Distillation (MOPD), which consolidates strengths from several teacher models on the student's own on-policy rollouts — so the student picks up capability without inheriting any one teacher's bias profile.
The demos: one-shot games and a debugged RL run
The most striking claim is one-shot interactive application generation: from a natural-language prompt, IQuest-Q1 emits runnable interactive apps in a single pass. The lab's examples include an FPS game — 3D scene, character movement, health, scoring, mode switching, and an in-game shop for resources and gear, all in one generation — and a racing game, where continuous scene extension (track geometry, foreground/background transitions) is the hard part one-shot generations usually fail at.
The other demos aim at working engineers. In one, an RL run went off the rails; starting from the training curves, the model pulled logs and execution traces, reasoned back through likely causes, localized the bug — a stray space inserted into the training trajectory — and after the patch and restart confirmed recovery from the new metrics. In office settings, the model runs multi-step tasks across chat, cloud docs, spreadsheets, and comment threads, pulling context together into analysis, drafts, and revisions.
Evaluations and caveats
The lab reports balanced results across benchmarks covering code and agentic work: NL2Repo (repository-level code generation), CyberGym (cybersecurity), Terminal-Bench 2.1 (terminal operation), DeepSWE v1.1 (long-horizon coding), and JobBench (professional office workflows), with full numbers in the technical report. The repo's benchmark notes disclose the recipe — temperature 1.0, top-p 0.95, top-k 20 — and the baselines: DeepSeek-V4-Flash/Pro's official July 31/August 13 releases, Humanity's Last Exam without tools, and an in-house IQuest-CLIBench for CLI user experience.
The caveats are stated plainly in the repo. This checkpoint takes text-only input. Tool calls use an IQuest-specific format in the chat template, so structured execution needs compatible serving parsers and an agent harness. On real-world CLI tasks the model "may overlook constraints, repeat failed attempts, or leave issues unresolved" — human oversight required. Deployment guidance targets SGLang or vLLM with tensor-parallel size 8, and the lab documents integrations with Claude Code 2.1.140 and Codex 0.142. Developers can also apply for upcoming early-access testing programs via email.
Why it matters
IQuest-Q1 is a pure bet on coding agents as the primary workload: sparse-MoE economics, a half-million-token context, tool use as a training objective rather than a bolt-on. That lands it in the same lane as DeepSeek's coding releases and the sparse-MoE wave generally — but the one-shot interactive app demos, if they hold up outside the lab, are the more distinctive claim. What to watch: independent benchmark numbers, the real cost of self-hosting a 320B-parameter model, and the license terms — the weights are publicly available, but GitHub lists the license as "other," so read the actual terms before building on it.