New in v1.5.0 / one note per fact, and the summaries write themselves
v1.4.0 put a size limit on the "where we are now" file. This release fixes the reason it grew. Until now, finishing one piece of work meant writing the same news into four or five different files by hand: the to-do list, the decisions file, the test map, the design notes. Nothing said which file owned which fact, so over time they drifted apart and disagreed, and the one you happened to read first was the one you believed. On one real project a single fixed bug had been written up seventeen times across three files.
Forge can now keep each note in its own small file, named after the thing it is about, and generate the summaries from them: the test map, the list of open work, the archive of finished work. You never edit a generated summary, and Forge will not let you: it stops the edit and points you at the note to change instead. Change the note, and every summary that mentions it updates itself.
That also settles which answer is current. When a decision replaces an earlier one, the earlier one is marked as replaced and drops out of every summary, while staying on file if you ever want to look it up. No more reading a stale answer at the top of a long file and a corrected one further down.
What you should notice: finishing a piece of work takes fewer steps and produces fewer contradictions, the summaries are always current because nobody maintains them, and asking "what is left to do" reads a short generated list instead of a very long history. Before anything is pushed, Forge checks that every note still points at something real.
Already have a project on an older Forge? Nothing breaks and nothing is treated as damaged. Your project keeps its current files until you say otherwise. Forge offers the change once, explains what it costs, and takes "not now" or "no" for an answer. Moving an existing project over is done as its own separate piece of work, never in the middle of something else, and it never deletes your old files: it moves them aside. See walking away and coming back.
What this is, and what it saves you from
Software projects fail in boring ways. You describe what you want too vaguely and get the wrong thing. You lose two days of work because nothing was saved properly. You come back after a week and cannot remember what you were doing. You ship something that quietly broke last Tuesday and nobody noticed.
Forge is a set of instructions that runs inside Claude Code. It does not write your software for you so much as it insists on doing it in an order that works, and it stops at the points where a human has to make a real decision.
The professional habits
- Asks until it actually understands your idea
- Sets up your computer with the right tools
- Writes automated checks that catch mistakes
- Saves work constantly so nothing is lost
- Keeps documentation truthful
- Packages finished versions you can point at
The judgment calls
- Answer questions about what you want
- Read the plan and approve it
- Decide who can see the project
- Approve anything that needs admin rights
- Say when a version is ready to release
You do not need to know how to code
You do need to know your own problem well enough to answer questions about it. If you can explain to a colleague what is annoying you and what "fixed" would look like, that is enough to start.
The mental model
Three phases, in order, with a hard stop between each. You cannot skip ahead, and neither can Claude. Think of it the way a building goes up: drawings get approved before anyone pours concrete.
How /forge knows where you are
You never tell it which phase you are in. It looks at your project and works it out. Expand a rung to see what happens in that situation.
01There is no written spec yet
Brand new project. It starts interviewing you about your idea.
02A spec exists but you have not approved it gate
It stops. It summarises what it wrote and waits. It will not approve its own work, no matter how confident it is.
03Spec approved, computer not set up yet
Moves to Phase 2 and starts checking what is installed on your machine.
04Setup started but something did not finish
Picks up at the exact item that failed rather than starting the whole setup again.
05Bench ready, nothing built yet
Writes a build plan: the list of pieces, hardest first, and asks you to approve the order.
06You were halfway through something
This is the one that matters most. It reads its notes, checks them against what is really on disk, and tells you if the two disagree before touching anything.
07Last piece finished, more still on the list
Starts the next piece. A rough edge bad enough to matter is a piece of work like any other, so it gets picked up here rather than waiting for the end.
08Nothing left on the list, but it has not been over the whole thing yet
Before it will propose a release, it does a finish-and-polish pass over every part this milestone touched: runs each one, looks at it properly, and reports what it found.
09The list is done and the polish pass is behind it gate
Stops being autonomous and proposes a release. Releases always need your say-so.
10The list is done but a release check is failing gate
Says exactly which check and what would clear it, and stops there. A requirement with no test proving it, or a rough edge still open on something this release claims to deliver, both land here.
11Everything in the spec is built and released
Reports where the project stands and asks what you want next. More capability means amending the spec, which is a gate of its own.
Which edition to install
Forge comes in two editions. They do the same job and produce the same project. The only difference is which generation of Claude model the instructions inside are written for, and that matters more than it sounds like it should.
Newer models follow instructions closely and check their own work without being told to. Older ones needed to be reminded, repeatedly and in detail. Give a newer model the older style and it over-checks everything and wastes effort. So there is one edition per generation rather than one that tries to suit both.
| If you use | Install | Type this |
|---|---|---|
| A current Claude model Claude 5 era. This is almost certainly you. |
forge-workflowv1.5.2 |
/forge |
| An older Claude model Claude 4 era |
forge4-workflowv0.1.0, frozen |
/forge4 |
If you are not sure which you are on, install the first one. It is the maintained edition, and it is what the rest of this guide describes.
Install one at a time
Both editions run the same background scripts, so having both switched on at once makes those scripts fire twice. Harmless but noisy. Install the one you need, and if you ever want to try the other, switch the first one off first.
The older edition has its own copy of this guide at daileng.github.io/forge4-workflow.
Setting it up, once
Two commands. Claude Code fetches Forge from a public catalog and installs it for you. These install the Claude 5 edition, which is the one you want unless you read section 03 and decided otherwise.
# point Claude Code at the catalog, then install the plugin claude plugin marketplace add DailenG/dailens-claude-toolbelt claude plugin install forge-workflow@dailen
Confirm when it asks, then run /reload-plugins or restart Claude Code. Type
/ and you should see forge in the list along with four others. Those
four are manual overrides you will probably never need.
When a new version ships later, claude plugin marketplace update dailen pulls it down.
One prerequisite worth checking
Forge uses small background scripts that need Node.js installed. If it is missing, everything still works except the automatic memory between sessions, which is a real loss. Forge checks this itself and tells you, so you do not have to remember.
Windows, Mac, or Linux
Describing the idea and building it work the same on every system. The setup phase is the one exception: the commands it uses to take stock of your machine and install what is missing are written for Windows and PowerShell today. On a Mac or on Linux it still runs, but during that phase it translates as it goes and uses your own installer instead, so say which system you are on when it starts and expect it to lean on you slightly more there.
The manual way, if you would rather
Prefer to drop the folder into place by hand? Download the zip from the releases page, close Claude Code, then extract it so the folder lands here:
# Windows C:\Users\<you>\.claude\skills\forge-workflow\ # Mac and Linux ~/.claude/skills/forge-workflow/
On Windows, that .claude folder is hidden. Paste the path into the Explorer address
bar rather than trying to click your way to it. Check it worked:
# should report no problems claude plugin validate $env:USERPROFILE\.claude\skills\forge-workflow --strict # should list forge-workflow@skills-dir claude plugin list
On a Mac or on Linux, make the background scripts runnable afterwards:
chmod +x ~/.claude/skills/forge-workflow/scripts/*.js. On Windows, skip that.
Restart Claude Code and you land in the same place as the two-command install above.
Phase one: describing the idea
Make an empty folder, open Claude Code inside it, and type /forge. Nothing else.
It will start asking questions in small batches, three to seven at a time, with lettered options where there is a sensible shortlist. Most rounds you can answer with a few letters and a phrase.
Round 1. Before anything else I need to know what problem this solves.
1. In a sentence or two, what should this do, and what is unacceptable about how things work today?
2. Who uses it? (a) just you, (b) your team, (c) customers. No recommendation, this is yours to answer.
3. Where is it used from? (a) a command line, (b) a desktop window, (c) a web page, (d) a phone.
2. a
3. a
Confidence: 20 out of 100. Held down by testing, which we have not discussed.
Round 2. Now the part that decides everything else...
The confidence score
After each round it scores how well it understands eleven areas, and reports the lowest one, not the average. That is deliberate. Averaging lets a single unanswered question hide behind ten confident ones, and that question is usually the one that wrecks the project later.
It keeps going until the lowest score reaches 95. In practice that is six to eight rounds and maybe twenty minutes.
One of the eleven areas is what the thing should be like to use: which surfaces it has, the handful of tasks people will do most often, how finished it needs to feel, and any accessibility obligation you are under. A surface nobody described gets designed by accident, so this one holds the score down like any other.
Phase 1 also writes a short design brief, docs/DESIGN.md, and settles what gets
written to a log and whether a failure ever gets reported back to whoever looks after the tool.
Both then become decisions you made, rather than whatever the first piece of building happened
to do.
If the thing has a real interface, it also tells you once what your options are for the visual side, and takes "none of these" as an answer. Claude Design is Anthropic's own design product: you explore looks and build clickable prototypes there, and it can hand the result over to be built. There are also add-ons for Claude Code that make generated screens look less generic, read an existing Figma design system, or drive a real browser so it can take pictures of what it built. Whatever you choose, the decisions get copied into the design brief in your project, so nothing in your depends on an outside account still existing. And nothing about your code or your data goes to any of them without asking you first.
Expect it to push back
It will press hardest on two things people skip. Testing: how you will know the thing actually works. Non-goals: what this deliberately will not do. Vague answers here produce a vague spec, so it does not let them pass.
Choosing the technology
Once it understands the problem, it presents two or three options for what to build it in, with honest downsides for each, and recommends one. You can accept or override. If you override, your reason gets written down alongside the original recommendation, so in a year you can see why.
The approval gate
Phase 1 ends with a written specification and a full stop. Claude will not approve its own spec. This is the most valuable twenty minutes in the project, so actually read it.
Can each requirement be proved?
Every numbered requirement should describe something you could verify. "Fast" is not a requirement. "Returns results in under two minutes for 300,000 files" is.
Is the non-goals list honest?
This is where scope creep gets prevented. If something you secretly want is listed as a non-goal, say so now. Later is much more expensive.
When you approve, the spec becomes a living document. It can still change, but only by adding a dated amendment explaining what changed and why, never by quietly rewriting history.
What to say
"Approved" is enough. Or list what you want changed first. Then type /forge again
and it moves itself to Phase 2.
Phase two: preparing the bench
This phase touches your computer, so it is the one most likely to need you. It assumes nothing is installed and checks everything, but only installs what is genuinely missing.
It inventories your machine
Prints one table: what it needs, what you have, what is missing. Then installs the gaps.
It may ask for administrator rights
Claude cannot grant itself admin. If something needs it, you get instructions for opening an elevated window, one block of commands to paste, and a command to run afterwards to prove it worked. It waits, then checks your result rather than taking your word for it.
It proves the safety nets actually work
This is the part people skip and it is the part that matters. It deliberately writes a broken test to confirm the test system reports failures. A test runner that finds nothing usually says "success", and that one silent lie would make every future green tick meaningless.
It asks one question you must answer
Should the project be public or private? It will never guess this. Private means only you can see it. If you are unsure, choose private; you can open it up later, but you cannot un-publish something.
It also sets up the means to look at the thing later, not only to test it. For anything with a visible interface that means a way to drive it, a way to capture pictures of it, and a checker for contrast and keyboard use. Alongside that it installs the logging the spec asked for, so the first piece of building cannot quietly invent its own.
About that "pre-push check"
Forge installs a guard that runs every time work is saved to the server: it builds the project, runs the tests, and scans for accidentally included passwords. If any fail, the save is refused. It will block you occasionally. That is the point, and Claude is instructed never to force its way past it.
Protecting the main line of work
The main copy of your project must not be deletable, and must not accept a save that throws away history. Forge sets that up in whichever of two ways your hosting account actually allows.
Best case, the host enforces it. The rule lives on the server, so it applies to everything and everyone, including changes made through the website.
Otherwise, your machine enforces it. Some hosting plans reserve that server setting for paid accounts. GitHub free personal accounts, for example, answer "Upgrade to GitHub Pro or make this repository public to enable this feature" on a private project. Forge does not buy anything and never makes a private project public to get around it. It installs a local guard instead, which refuses a history-destroying save from this computer, and tells you once what that does not cover: a copy of the project on another machine that was never set up, a change made through the website, a guard someone deleted, or somebody who has your password. Then it carries on.
Either way, Forge proves the guard works before trusting it, using throwaway practice copies of a project rather than your real one, and writes down which of the two is in force.
Phase three: building in slices
Work happens in . A slice is one complete piece of usable behaviour, not one technical layer. "Scan a folder and print what it found" is a slice. "The database code" is not, because on its own it does nothing you can look at.
Hardest slice goes first. That feels backwards, and it is correct: if something is going to turn out impossible, you want to find out on day one, not in week three.
Why write the test before the thing works
Because a test that has never failed has not been shown to test anything. Plenty of tests pass because they check nothing at all, or check the wrong thing entirely, and you cannot tell the difference by reading them. Watching it fail first, then pass, proves it is connected to reality.
The rule Claude is forbidden to break
It may never make a failing test pass by weakening the test. No deleting it, no skipping it, no loosening what it checks. If a test is genuinely wrong it has to say so and ask you. This is written down explicitly because it is exactly the shortcut anything under pressure reaches for.
Looking at the thing, not just testing it
Passing tests only prove the thing does what it was asked to do. They say nothing about what it is like to use. So once the tests are green, Forge runs the real thing and uses it: walks the tasks people will do most, checks what it shows when there is nothing there yet, while it is working, when something has gone wrong, and when it worked, drives it with the keyboard alone, and reads the actual words on screen.
Every rough edge it finds is either fixed there and then or written down in a list of with how much it hurts: blocks the task, makes it worse, or merely unfinished. Nothing gets waved away as a matter of taste, because an observation nobody wrote down is an observation that is gone.
After each slice you get a short report: what works now, what does not yet, what is next. By default it carries straight on to the next slice. If you would rather approve each one, say so and it will wait.
When it stops you
Forge blocks things on purpose. A block is not a malfunction, it is the system doing the one job you cannot do by staring at the screen. Here is how to read the common ones.
| What you see | What it means and what to do |
|---|---|
| A secret was found | A password or key is about to be saved into the project history, where it would live forever. Remove it from the file and, if it was ever real, change it. Never bypass this one. |
| Tests failed, push refused | Something that worked is now broken. Let Claude fix it. Do not ask it to skip the check. |
| Cannot release, requirement not covered | One of your requirements has no test proving it works. Either it is not actually finished, or it needs a test. Claude will tell you which. |
| Typography violation | A curly quote or long dash got into a file. Harmless in prose, breaks things in code. Claude fixes it itself. |
| Notes disagree with the project | Its record of where things stand does not match reality. It stops rather than guessing. Read what it found and tell it which version is right. |
| Screenshots missing | Documentation needs a picture Claude cannot take, usually of a window on your screen. It tells you exactly which one and what should be visible in it. |
| Cannot release, a rough edge is still open | Something this release claims to deliver has a known problem that blocks a task or makes it noticeably worse, and it is still on the list. Fix it, or say explicitly that it waits for a later version. Leaving it for later is your call, never Claude's. |
| That would change the spec | You asked for something the approved spec does not cover, or a new outside component is needed to do it. It stops, states the amendment, and waits. One deliberate confirmation, then it carries on. |
| That would send your work somewhere else | Something wants to upload your code, a screenshot of real data, or user data to an outside service the spec never named, a hosted design tool included. It always asks first. |
| Something newer is missing from this project | Not a block, and not the project being broken. Forge has gained a habit since this project started. It says what is missing and asks once whether to add it now, add it as its own piece of work, or skip it. |
The one instruction to avoid
Never tell it to skip, bypass, or force past a check. The checks are the whole reason the project stays trustworthy while someone who does not write software is driving it. If a check is genuinely wrong, ask Claude to explain it, then decide.
Walking away and coming back
Real work gets interrupted. A meeting overruns, the day ends, a fortnight passes. This is the thing Forge is best at, and it needs nothing from you.
Claude keeps a file called CONTINUE.md that always describes exactly where things
stand and what the very next action is. It updates it before starting anything risky, not
afterwards, so an interruption cannot catch it out.
When you return, that file is loaded automatically before you say a word. Then, crucially, it checks the notes against the actual project. If the notes claim two tests are failing but three are, it says so and asks you rather than carrying on from a false picture.
Forge itself keeps gaining habits, so a project you started months ago can be missing something newer versions set up from the start, like the design brief, the list of rough edges, or the decisions about what gets logged. That is not the project being broken, and it is not the notes disagreeing with reality. Forge spots the gap, says what is missing and what it would cost to add now, and asks once: add it now, add it as its own piece of work, or skip it for this project. Skipping is offered but not recommended, because the checks that habit feeds go quiet with it. Whatever you answer is written down so you are not asked again, and you can change your mind later by saying so.
All you ever have to type
please continue or /forge. Either works. You never have to remember what you were doing.
Shipping it
A release is a numbered, frozen snapshot you can point people at. Forge cuts the first one as soon as anything works end to end, then again after each meaningful addition, rather than saving them all up for a big finish.
| Version | What it signals |
|---|---|
v0.1.0 | First thing that genuinely works, even if small |
v0.2.0 | A meaningful new capability |
v0.2.1 | A fix, nothing new |
v1.0.0 | Every requirement done, and every one proved by a passing test |
Before any release, Claude switches out of autopilot and asks. It also refuses to cut version 1.0 until every single requirement in your spec has a test proving it works. Not "the code exists", but "here is the check that demonstrates it".
Before a numbered release it goes over every part the milestone touched against a checklist matched to the kind of surface it is: screens, a command line, a library other people write code against, or a service somebody has to run. What it found and what it fixed is written down. Anything it would rather leave for a later version is yours to approve, not its own call.
Each release comes with a plain-language paragraph saying what this version proves the tool can now do. That paragraph is what you will actually read in six months when you are trying to remember where you got to.
The files that appear in your project
Forge creates these. You never have to edit them, but knowing what they are helps.
| File | What it is |
|---|---|
CONTINUE.md | Where things stand right now. The one that makes "please continue" work. |
TODO.md | What is planned, what is being done, what is finished, and the rough edges still owed a fix. |
docs/SRS.md | The specification you approved. The source of truth for what is being built. |
docs/DESIGN.md | What it should be like to use: the surfaces, the tasks that matter, the wording and look rules, and a log of every polish pass. |
docs/DECISIONS.md | Every choice made and why. Read this when you wonder "why on earth did we do it that way". |
docs/ENVIRONMENT.md | What is installed on your machine and what versions. How you would set this up again. |
docs/traceability.md | A grid matching every requirement to the test that proves it. |
docs/images/ | Pictures the documentation refers to, plus a list of which ones still need taking. |
docs/design/ | Only if you use an outside design tool: what it handed over. The decisions are copied into the design brief, so the project never depends on that tool still existing. |
.forge/ | The guard that refuses a history-destroying save, and a note of which kind of protection is in force. Leave it alone. |
CHANGELOG.md | What changed in each release, written automatically. |
README.md | How to install and use the thing you built. Written for someone who has never seen it. |
Plain-English words
The vocabulary you will meet, without the jargon. Terms are underlined throughout the page, so you can check one without leaving your place.
| Word | What it actually means |
|---|---|
| Repository | Your project folder plus its full history. Every version of every file is kept, so nothing is ever truly lost. |
| Commit | A saved checkpoint with a note explaining what changed. Like a save point in a game. |
| Branch | A safe copy to experiment in. If it goes badly you throw the copy away and the real project was never touched. |
| Merge | Folding a finished branch back into the real project. |
| Push | Sending your saved work up to the server so it exists somewhere other than your laptop. |
| GitHub | The server where the project lives. Also your offsite backup. |
| Test | A small program that checks your program. It either passes or fails, no interpretation needed. |
| Test suite | All the tests, run together. "Green" means everything passed. |
| CI | A robot on the server that rebuilds the project and runs every test after each change, on a clean machine, in case something only worked on yours. |
| Hook | An automatic check that fires at a set moment, like just before saving to the server. |
| Coverage | How much of your code the tests actually run. Useful for spotting untested areas, not a measure of quality. |
| Slice | One complete piece of working behaviour, built and tested end to end. |
| Release / tag | A numbered, frozen snapshot. Tags never move once created. |
| Spec / SRS | The written description of what is being built, agreed before building starts. |
| Requirement | One numbered, provable statement of something the tool must do. |
| Surface | Anything a person or another program meets your tool through: a screen, a command line, a published interface, a service. |
| Design brief | The short written statement of what the thing should be like to use. The project's own answer, not a tool's. |
| Design pass | Running the real thing and using it, rather than only testing it. Done inside the piece of work, not saved for the end. |
| UX debt | A rough edge that was found and not fixed on the spot, written down with how much it hurts: blocks the task, makes it worse, or merely unfinished. |
| Polish pass | The sweep over everything a release touches, before the release is even proposed. |
| Gate | A point where Claude stops and a human decides. It cannot pass one on its own, whatever mode it is in. |
| Autopilot | The default: it carries on from one piece to the next and reports as it goes. Say the word and it waits for you each time instead. Near a release it stops asking permission to ask. |
| Backfill | Adding something to an older project that newer versions of Forge set up from the start. |
Quick reference
Commands
/forge | Start, or carry on. The only one you need. |
please continue | Same thing, in plain words. |
/forge-spec | Force the describing phase. |
/forge-env | Force the setup phase. |
/forge-code | Force the building phase. |
/forge-design | Force a design and polish pass. |
Phrases
| "Approved" | Passes a gate. |
| "Explain that" | Any term or decision, in more detail. |
| "Wait for me each slice" | Switches off autopilot. |
| "Where are we?" | Status without doing any work. |
| "Why did we decide that?" | It looks it up in the decisions log. |
| "This feels wrong to use" | Asks for a design pass on that surface, however vague the complaint. |
| "What is still rough?" | Reads back the rough edges on the list and how bad each one is. |
Never say these
"Skip the tests." "Force the push." "Bypass the check." "Just make it pass." Each one removes a safety net that exists precisely because nobody is reviewing the code but the machine that wrote it.
If you only remember four things
- Type
/forge. It figures out the rest. - Read the spec before approving it. Twenty minutes there saves days later.
- A block is the system working. Never ask it to bypass one.
- You can leave at any moment. Come back and say "please continue".