Introducing Ultrapowers
Most coding agents do one thing very well: they start writing code the moment you finish typing. That is also their biggest weakness. They skip the questions a senior engineer would ask, they build without a plan, they test after the fact if at all, and nobody reviews the result until it lands on your desk as a pull request you did not expect.
Ultrapowers is my answer to that. It is a plugin for coding agents that turns a set of composable skills into a development method the agent follows on its own, from the first message of a session to the pull request at the end of a ticket. You decide how much of it runs by itself, and nothing merges without you.
This article explains what it is, how it is put together, the three ways teams use it, and what has been proven so far. If you want the step-by-step version, the companion tutorial takes you from install to your first reviewed pull request: Ultrapowers: from install to your first reviewed pull request.
The problem
A ticket arrives. The agent reads the title, guesses the intent, and opens the editor. Twenty minutes later there is code, and three things are usually wrong with it:
- It solves a slightly different problem than the one the ticket meant, because nobody asked.
- It has no tests, or tests written after the code that pass by construction.
- Nothing about the decisions survives: no spec, no plan, no record of what the agent assumed.
Teams that live from tickets cannot work like that. They review specs, they review plans, they review pull requests, and they expect the same bar whether a person or an agent did the work. The question I kept hearing from engineering teams was the same one: give the agent a task and wait for the pull request, without lowering the bar.
What Ultrapowers is
Ultrapowers is a skills library plus a bootstrap that loads at the start of every session, on every supported harness, so the agent uses the skills without being told. It installs as a plugin on fifteen coding agents, among them Claude Code, Codex, Cursor, GitHub Copilot CLI, Gemini CLI, OpenCode, Kimi Code and Antigravity.
What you get, in six lines:
- Workflow skills that trigger on their own. Brainstorm before design, plan before code, write the failing test first, review every task, finish the branch cleanly.
- One setup for every agent.
/ultrapowers:initwrites the project’s instructions, the MCP configuration for each harness, ten knowledge-base folders and the repository hygiene, after showing you the exact file list. It never overwrites a file. - One trail per ticket. The brief, the spec, the plan and the review share the ticket id across
tasks/,specs/,plans/andreviews/. - Autopilot. One command takes a GitHub or GitLab ticket to a pull request, with your approval on the ticket and your merge on the pull request.
- Team memory. Verified lessons saved in git, for every developer and every agent on the project.
- A QA specialist (beta). A tester’s eye on the running app, with one verdict per ticket.
Ultrapowers is a fork of Jesse Vincent’s MIT-licensed skills library, cut from its version 6.4.2. Everything above the original methodology was built from daily work across real projects, and the repository is source-available under a proprietary license.
The architecture in one picture
Five parts, and each one has a single job.
- The bootstrap runs at session start on every harness. It is the reason the skills trigger without a prompt.
- The skills are the method: brainstorming, writing plans, test-driven development, code review, finishing a branch, and the ticket skills that keep one id across the whole trail.
- One trail per ticket lives in the documents repository, the workspace root where init ran. A single-repository project keeps code and documents together; a workspace with nested clones keeps the documents at the root and one branch per code repository.
- The engine is plain code with no AI in it. It posts the review packet, checks your label on the tracker’s own timeline, pushes between stages and opens the pull requests. It has two doors: the command you run in a session, and a watcher that picks up tickets you labelled.
- The guardrail is a hook that runs before every tool call during a stage. It denies the agent a push, a merge, a tracker write, and any edit to the project’s configuration, CI, hooks or settings. It matches patterns; it is not a sandbox.
Three ways to work
Start on the left. Move right when you trust the results.
| Way | You do | The agent does | Where it has run |
|---|---|---|---|
| Manual | Run each skill, answer questions, approve spec and plan, choose merge. | Reads your code, drafts spec and plan, codes test-first, reviews each task. | Every supported harness. |
| One command | Approve spec and plan with one label; review and merge the PR. | Spec, plan, then test-first code and reviews; the engine opens the PR. | Live: Claude Code, GitHub, one repository, gated mode. |
| Watcher (beta) | Run it on a disposable host; then label tickets and merge PRs. | Same stages as One command, started from a ticket label. | Offline tests and vendor documentation only. |
The artifacts are the same at every rung: the brief, the spec with its assumption ledger, the plan, the stage log, the pull request. Moving right changes who starts the run and how many times it stops, never what it produces.
How you stay in control
This is the chart I would show a non-technical lead. Yellow is you, blue is the agent, grey is the engine.
| Moment | What happens | Your move |
|---|---|---|
| The packet | Brief, spec with assumption ledger, plan, repositories in scope; under 25 lines. | Read it on the ticket. |
| The approval | The engine verifies who added the label, their write access, the timing and the commits; a new commit voids it. | Add up:approve. |
| Changes | The run returns to the spec and the plan. | Add up:changes with a comment. |
| Hold | The run pauses; a kill-switch file pauses the watcher. | Add up:hold. |
| The pull request | The engine pushes and opens it, citing the packet, the approver and the stage log. | Review, then merge yourself. |
Two details matter more than they look. The assumption ledger lists every question the agent answered for itself, lowest confidence first, so the lines you most need to check are at the top of the packet. And the approval is a label on the tracker, verified on the tracker’s own timeline, so the record of who approved what, and at which commit, is the tracker’s record and not the agent’s.
Which flow fits your team
- A solo developer on GitHub with Claude Code. Start manual for a ticket or two, then turn on autopilot in gated mode and approve your own tickets. This is the closest to the path that has run live: gated mode, inline execution, the developer’s own account approving.
- A non-technical lead who owns the tickets. You read the packet on the issue, add one label, and later merge the pull request. A developer still runs the command in a session, unless a watcher runs.
- A team on several coding agents. Cursor, Codex and Gemini CLI read the same
AGENTS.md, the same ticket trail and the same team memory from git. One setup, every agent. - A team on GitLab with several repositories. The workspace root holds the documents, each nested clone gets its own branch, and the engine opens one merge request per repository. Documented and covered by the offline suites; not yet run live.
- An agency running a client’s tickets. The client’s tracker is the source, the client’s leads are the approvers, and the client merges. The packet becomes the client-facing record. The watcher (beta) belongs on a disposable machine with read-only stage tokens and a kill switch.
Do you want me to set this up on your project? Contact me.
What is proven and what is beta
| Verified | |
|---|---|
| Live, on the plugin’s own repository | The session door on Claude Code with GitHub: one repository, gated mode, no QA stage. A label added before the packet changes nothing. The agent does not add the approve label for you, even when asked to. |
| Offline suites, every commit | The engine against a fake tracker and a fake harness (54 scenarios), the guardrail’s autopilot profile (120 cases), the in-process hooks of OpenCode, Pi and Hermes, the adapters’ command lines, init’s autopilot mode. |
| Vendor documentation only | The hook contracts of the other harnesses. Devin’s hooks are documented as fail-open, so the watcher refuses it. |
| Not yet run | The watcher on any harness, GitLab, a nested workspace, full mode, the changes loop, the QA stage inside a run. |
I would rather publish that table than a productivity figure. There is no measurement behind one yet, and the people this is built for would ask to see it.
Start here
- The tutorial, from install to your first reviewed pull request: read it here.
- The repository, with the README, the release notes and the install commands for every harness: github.com/raoofaltaher/ultrapowers .
- Questions and ideas go to the repository’s Discussions .
Where this goes next is already written down: a second gate on the pull request, the QA report posted on the ticket, a CI door, and Odoo write-back. Those are the roadmap, not the release.
Work with me
Do you want to automate the development workflow of your project with Ultrapowers? Do you want me to set it up on your project, or teach your team how to use it? Contact me.