Your coding agent has a discipline problem. Give it a task and it does what an eager junior dev does on day one: it starts typing. No clarifying questions, no design, no plan — straight to code, tests as an afterthought, and a confident "done!" over work it never verified. The model isn't the problem. The missing piece is a methodology, and that's exactly what Superpowers installs: a set of agent skills that turn "just build it" into brainstorming, design review, written plans, red-green TDD, and code review — enforced by the agent's own instructions.

The project is obra/superpowers, built by Jesse Vincent and the team at Prime Radiant. As of October 2, 2026 it holds 294,000+ stars and 26,300 forks, sits at #6 on GitHub's daily trending page, and ships install paths for 16 agent harnesses — Claude Code, Cursor, Codex, Gemini CLI, Pi, OpenCode, and more. It is MIT-licensed, and the current release is v6.4.2 (September 25, 2026).

This tutorial is hands-on and verified: I installed Superpowers into the Pi coding agent, captured the exact bootstrap text it injects into a live session, ran the project's own extension test suite (6/6 passing), and walked the full methodology loop — worktree, plan, failing test, implementation, review, merge — on a real task in this sandbox. Everything below runs as written.

What you'll need #

  • One AI coding agent with plugin or skills support: Claude Code, Cursor, Codex CLI, Gemini CLI, Pi 1.0, OpenCode, or any of the 16 harnesses in the README. (Pi is used for the verified install below; the skills themselves are plain Markdown and work anywhere.)
  • A provider API key for your agent — the methodology only matters once a real model is driving. My sandbox run used a local stub model plus manual execution, and every step that needs a live model is marked.
  • Git, and a terminal. No GPU, no model downloads, no build step — the whole thing is Markdown skills plus small per-harness bootstraps.
  • Cost: $0 in software. The methodology itself costs tokens: plans, reviews, and subagent passes are extra model calls. That is the trade, and it's explicit.

Step 1 — Install it in your harness #

Installation differs per harness; the README documents all 16. These are the commands exactly as documented, and the Pi one is verified below:

# Claude Code (official marketplace)
/plugin install superpowers@claude-plugins-official

# Cursor (agent chat)
/add-plugin superpowers

# Codex CLI (plugin search, then install)
/plugins   # search for "superpowers"

# Gemini CLI
gemini extensions install https://github.com/obra/superpowers

# Pi
pi install git:github.com/obra/superpowers

# Muse
muse plugins install ./superpowers && muse plugins approve superpowers

For Pi, I ran the install for real in an isolated agent directory:

PI_CODING_AGENT_DIR=~/pi-agent pi install git:github.com/obra/superpowers
# Installing git:github.com/obra/superpowers...
# Cloning into '.../pi-agent/git/github.com/obra/superpowers'...
# Installed git:github.com/obra/superpowers

pi list
# User packages:
#   git:github.com/obra/superpowers

The package registers two resources with Pi (declared in package.json): the ./skills directory and the ./.pi/extensions/superpowers.ts extension. That extension is the interesting part — read on.

Step 2 — Watch it take over the session #

Superpowers works by injection, not by tools. The Pi extension hooks the session_start and session_compact events and inserts the full text of the using-superpowers skill into the conversation as a user message — at the top, after any compaction summary. From the model's perspective, it opens every session reading this:

<EXTREMELY_IMPORTANT>
superpowers:using-superpowers bootstrap for pi

You have superpowers.

The using-superpowers skill content is included below and is already
loaded for this Pi session. Follow it now. ...

I verified this end-to-end: with the package installed, I ran pi -p --model stub/stub-model "Say the word pineapple." against a local stub model server and logged the exact request body Pi sent. The bootstrap marker superpowers:using-superpowers bootstrap for pi was present in the request, and the session completed normally. The project's own test suite for this extension passes 6/6 on Node 24 (node --test tests/pi/test-pi-extension.mjs), including "startup context injects the bootstrap as one user message until agent_end" and "session_compact injects bootstrap after compaction summaries, not before compaction".

The core rule the bootstrap installs is blunt: invoke a relevant skill before any response or action — including clarifying questions. If there is even a 1% chance a skill applies, the agent must use it. "Let's build X" routes to brainstorming first; "fix this bug" routes to systematic-debugging first. The skill even ships a table of "red flag" thoughts ("this is just a simple question", "I'll just do this one thing first") with the reality check for each. It reads like tough love because it is.

Official Superpowers project icon
Official Superpowers project icon (MIT), from the obra/superpowers repository.

Step 3 — Learn the seven moves #

The README's "Basic Workflow" is seven skills in sequence. Here is what each one actually mandates, quoted from the skill files:

1. brainstorming. Fires before any creative work. The agent must discover intent (who is this for, what does success look like), write back its understanding for you to correct, and then follow one of three paths — spike, bounded, or architectural. The hard gate: no implementation action until you approve the design artifact. Approval of an idea does not approve artifacts that don't exist yet.

2. using-git-worktrees. Work happens in isolation. The skill first detects whether you're already in a worktree (GIT_DIR != GIT_COMMON), asks consent before creating one, then sets up the workspace and verifies a clean test baseline.

3. writing-plans. Plans are written for "an engineer who has not seen this codebase or this spec" — exact files, names, signatures, and the tests that prove each task. Plans live under docs/superpowers/plans/, named YYYY-MM-DD-<feature-name>.md, every plan starts with a fixed header, and each step is one action with a checkable result: write the failing test, run it, implement, run tests, commit.

4. subagent-driven-development (or executing-plans for inline). A fresh subagent per task, a review after each task, continuous execution with no "should I continue?" check-ins. The skill's standout rule is rulings, not stalls: ambiguities get decided and logged, never parked on a question. Only four things stop a run: an irreversible or destructive operation, a security-sensitive action, a side effect outside the worktree (merge, push, publish), or a plan so broken that every path forward is a guess.

5. test-driven-development. The iron law, verbatim: NO PRODUCTION CODE WITHOUT A FAILING TEST FIRST. Wrote code before the test? Delete it and start over — "don't keep it as reference, don't adapt it." If you didn't watch the test fail, you don't know it tests the right thing.

6. requesting-code-review. Reviews happen between tasks against the plan, with issues reported by severity. Critical issues block progress.

7. finishing-a-development-branch. Verify tests, detect the environment, present merge/PR/keep/discard options, execute your choice, clean up the worktree.

Beyond the loop, three skills handle the edges: systematic-debugging (reproduce first, fix second), verification-before-completion (evidence before assertions — run the command, read the output, then claim), and diagnosing-superpowers for post-mortems. That last one is worth knowing early: when a session goes sideways, you tell your agent "figure out what went wrong with superpowers in this session" and it reads the transcript, cites every finding as path:line, and builds a bug report. No citation, no finding.

Step 4 — Run a full loop on a real task #

Here is the loop executed for real. The task: a zero-dependency Python CLI (linkcheck) that reads URLs from a file and reports HTTP statuses. I played the agent's role by hand — this sandbox has no LLM API key — following the skill text exactly. On your machine, your agent performs these same moves autonomously.

Brainstorm → worktree. The design decision (stdlib-only urllib, exit 0 when all URLs are 2xx, exit 1 otherwise) went into a short design note. Then isolation, exactly as the skill prescribes — detect first, then create:

git init -b main linkcheck && cd linkcheck
GIT_DIR=$(cd "$(git rev-parse --git-dir)" && pwd -P)
GIT_COMMON=$(cd "$(git rev-parse --git-common-dir)" && pwd -P)
# GIT_DIR == GIT_COMMON -> normal checkout, create a worktree:
git worktree add ../linkcheck-wt -b feature/linkcheck

Plan. Written to docs/superpowers/plans/ as 2026-10-02-linkcheck.md with the skill's required header (goal, architecture, tech stack, spec pointer, global constraints) and two bite-sized tasks, each ending in a committable, testable deliverable.

RED. Test first, per the iron law. The test file uses a local http.server stub so it needs no network:

python -m pytest test_linkcheck.py -x -q
# >       from linkcheck import check_urls
# E       ModuleNotFoundError: No module named 'linkcheck'
# 1 failed in 0.58s

That failure is the point. The test fails correctly — it fails because the module doesn't exist, not because of a broken assertion.

GREEN. Minimal implementation — check_urls() over urllib.request.urlopen, plus an argparse main() that prints url -> status lines and returns the exit code:

python -m pytest test_linkcheck.py -q
# ..                                                   [100%]
# 2 passed in 1.27s
git commit -m "linkcheck: check_urls + CLI per plan (tasks 1-2, TDD green)"

Review. Measured against the plan, the review surfaced two minor findings and zero criticals: urlopen follows redirects silently (a 301→200 reports 200), and network errors surface as status 0 rather than a real HTTP code. Both accepted — the plan only demands correct 2xx/otherwise exit semantics, and both behaviors satisfy it.

Finish. Full suite green, merge, worktree removed:

git merge --no-ff feature/linkcheck -m "merge feature/linkcheck"
git worktree remove ../linkcheck-wt --force && git worktree list
# .../linkcheck  4efad46 [main]

Start to finish: design note, isolated worktree, written plan, red test, green implementation, documented review, clean merge. Nothing about this required a framework — it required following the checklist, which is the entire thesis of the project.

Abstract illustration of the red-green test-driven development loop
AI-generated illustration for AI Frontier Post: the red-green TDD loop at the heart of the methodology.

When to use Superpowers vs the alternatives #

ApproachWhat it isUse it when
SuperpowersFull SDLC methodology as skills: brainstorm → plan → TDD → review, 16 harnessesYou want an agent that behaves like a disciplined senior team on multi-step work
GitHub Spec KitSpec-driven development: constitution → spec → plan → tasksYou want spec-first rigor with an explicit task breakdown format
Addy Osmani's agent-skills25 senior-engineer workflow skills (code review, debugging, testing)You want à-la-carte engineering practices without the full lifecycle
Ralph loop / plain promptingIterative "keep going until done" loopsSmall, well-understood tasks where process overhead exceeds the benefit

Superpowers overlaps most with Spec Kit — both insist on design-before-code. The difference is scope: Spec Kit is a specification pipeline, while Superpowers is a whole operating system for the session (worktrees, subagent dispatch, review gates, post-mortems). If your pain is "the agent writes plausible code that doesn't match what I asked for," Superpowers' brainstorming hard gate is the specific fix.

Caveats, honestly #

  • The live loop needs a model. Everything installable was verified above, but the autonomous brainstorm→plan→implement cycle requires your agent plus an API key. My worked example was manual execution of the same steps — the commands are real, the outputs are real, but I was the agent.
  • Overhead is real. Plans, per-task reviews, and subagent passes cost tokens and time. For a one-line fix, the methodology is heavier than the task. The skills say as much: process skills set the approach, and you choose the approach per task.
  • Skills are prompts, not code. Nothing here forces the model to comply — the "mandatory" language is prompt engineering, and a weak model can still drift. That's what diagnosing-superpowers is for, and why the project is explicit that you report; you don't get guarantees.
  • Telemetry. The optional visual companion in the brainstorming skill loads the Prime Radiant logo from their website with your Superpowers version attached — used for rough usage counts, no project or prompt data. Disable it with SUPERPOWERS_DISABLE_TELEMETRY=1; it also honors Claude Code's DISABLE_TELEMETRY.
  • Version pinning. I verified v6.4.2. Skills evolve (the bootstrap itself warns "I remember this skill" is a red-flag thought), so re-read the skill files after updating rather than relying on memory.

The takeaway #

The reason Superpowers is at #6 on GitHub trending isn't a new model or a new tool — it's that the bottleneck in AI-assisted coding stopped being code generation a while ago. The bottleneck is process: agents that don't clarify, don't plan, don't test, and don't check their work. Fifteen Markdown files and a bootstrap hook fix that by making the methodology the path of least resistance.

Install it with one command, say "let's build X," and watch your agent ask what you actually mean before it writes a line. That single behavior change is worth more than any new model checkpoint.

Sources #