What Is an AI Software Factory?
"AI software factory" is having a moment. It's in pitch decks, in job posts, and on our own homepage, which calls Agentastic.dev "the multi-agent IDE for your AI software factory." If we're going to use the phrase, we owe you a definition.

So here's ours. Then the history, which goes back further than you'd guess, the parts a factory actually needs, and how to start one without reorganizing your team.
The short version#
An AI software factory is a repeatable system in which coding agents do most of the production work, like writing code, running tests, and opening pull requests, while engineers decide what gets built and what ships.
The agents are the workers. The factory is everything around them: how work comes in, where agents run, how their output gets checked, and how it merges.
That second part is the whole point. Anyone can open five terminals and start five agents. If you still have to babysit every step, you don't have a factory. You have a workshop with very fast apprentices.
The term is older than you think#
Software people have been borrowing from manufacturing since the 1960s.
In 1968, Bob Bemer of General Electric argued for a "software factory": standard tools, a common interface, and a database of past projects, so that output depended less on which programmer you happened to get. Hitachi was the first company to actually build one under that name, opening the Hitachi Software Works in 1969. System Development Corporation followed in 1975, and NEC, Toshiba, and Fujitsu set up their own in 1976 and 1977.
The best-documented results came from Japan. MIT's Michael Cusumano studied those shops for his 1991 book Japan's Software Factories. At Toshiba's factory in Fuchu, the share of delivered code reused from earlier projects rose from 13% in 1979 to 48% in 1985, and output per programmer more than doubled between 1976 and 1985.
The term kept coming back. In 2004, two Microsoft architects, Jack Greenfield and Keith Short, published Software Factories, a book about assembling applications from patterns, models, and frameworks instead of writing everything by hand. In 2017 the US Air Force started Kessel Run, its first software factory. By September 2021 the Air Force had 17 of them.
Every version chased the same goal: make software production repeatable enough that output stops depending on heroics. Standard tools, reuse, process, measurement. What none of them could change was the labor. Every station still needed a person, and hiring was the only way to add capacity.
Coding agents change that. The line can now be staffed by software, which is why the phrase is back.
What makes it a factory#
A factory is a line of stations, each with a clear input and output. For software, the line looks like this:
- Intake. Work arrives from where it already lives: the issue tracker, error monitoring, team chat, or a schedule. Nobody should have to copy a ticket into a terminal.
- Plan. A ticket becomes something an agent can execute: a clear prompt, the relevant context, and a definition of done.
- Build. Agents work in parallel, each in its own isolated workspace (a git worktree, a container, or a VM) that starts in a known-good state.
- Review. Every change lands as a diff. Tests and AI review catch the obvious problems, and a person reads what's left and decides.
- Ship. The change goes out as a pull request, CI runs, and it merges.
Two things aren't stations but hold the line together. The first is reproducible environments: an agent dropped into a bare checkout burns half its budget working out how to install dependencies. The second is visibility: with ten runs going, you need one place that shows what's running, what's stuck, and what needs you.
Each missing piece has a recognizable symptom. Without intake, you have a tool you operate by hand. Without isolation, your agents fight over the same files. Without review, you have a faster way to ship bugs.
Review is the bottleneck#
Here's the part the pitch decks skip. Agents can already produce diffs faster than anyone can read them. Once generation is cheap, a factory's throughput is set by review, not by how many agents you run.
A few things actually move that bottleneck:
- Smaller tasks. A 40-line diff gets a real review. A 2,000-line diff gets a skim and a prayer.
- Machine checks first. Tests, type checks, and an AI review pass should reject the obvious failures before a person spends time on them.
- Proof that it runs. A screenshot or a quick look in a browser beats reading a UI diff line by line.
- Standing rules in the repo. Put your conventions in the agent's instructions and your review prompt, so the same mistake doesn't come back every run.
What doesn't work is skipping review. The whole promise of a factory is consistent output, and nothing makes output less consistent than code nobody read.
One vendor's factory, or your own#
There are two ways to get a factory today.
The first is to buy one. Several companies now sell a line where their agent runs in their cloud and you review its work in their dashboard. It's the fastest way to start, and for some teams it's the right call.
The catch is that the agent is the part of the line that changes fastest. The best coding agent this year probably won't be the best one next year. A single-vendor factory ties your intake, compute, review, and merge process to one vendor's agent, and swapping it out means rebuilding the line.
The second way is to build the line yourself and treat agents as interchangeable workers. The stations stay put. When a better agent ships, you add it and route work to it. Your issue tracker, environments, review habits, and CI don't change.
That's the bet we made with Agentastic.dev.
Start small#
You don't need a reorg or a platform team. Build the line one station at a time:
- Pick one boring category of work. Flaky tests, dependency bumps, docs that drifted, small bugs with a clear repro. Well-specified work with an obvious definition of done.
- Make the environment reproducible. Write a setup script that takes a fresh checkout to a working state. It's the single biggest factor in whether agents succeed.
- Run a few agents in parallel. Two or three, each in its own worktree, on separate tasks. Watch where they get stuck.
- Put review on rails. AI review first, then your read of the diff, then CI. Keep tasks small enough that you actually read them.
- Automate intake last. Once the first four are boring, let a schedule or an issue start the runs.
The order matters. Automating intake before review works is how you end up with forty pull requests you don't trust.
How Agentastic.dev fits#
Agentastic.dev is a native macOS app for running coding agents in parallel, and it has something for each station:
- Intake: start runs from Linear or Sentry issues, Slack mentions and DMs, or a schedule.
- Plan: write the task as a prompt, or give a Manager Agent a standing charter and let it coordinate other agents.
- Build: 50+ built-in agents, including Claude Code, Codex, Gemini CLI, and GitHub Copilot, plus any CLI you add. Each run gets its own worktree or container, prepared by your setup script, on your Mac, on a machine you reach over SSH, or in a cloud VM on Modal, Fly.io, or Vercel Sandbox with your own account.
- Review: read every diff, try the result in the built-in browser, and run AI code review with Claude, Codex, or CodeRabbit.
- Ship: open the pull request, watch CI checks, and merge GitHub pull requests without leaving the app.
Two honest limits. Local runs need your Mac awake with the app open, so work that should keep going after you close the lid belongs on an SSH host or a cloud VM. And issues don't turn into pull requests on their own: you start the run, or a Manager Agent or scheduled task picks the work up while the app is running. We think a person should still be the one who merges.
For the whole line on one page, see the AI software factory in Agentastic.dev. If you mostly want to run several agents side by side, start with the multi-agent IDE.
The job moves to the floor#
The factory metaphor gets one thing right that "AI pair programmer" never did. The interesting work moves from the line to the floor: deciding what gets built, designing the stations, and signing off on what ships.
That's still engineering, just a different part of it. The teams that get the most out of agents won't be the ones running the most of them. They'll be the ones with the cleanest line: clear tasks, reproducible environments, and review that someone actually does.
Agentastic.dev is free for macOS. Bring the agents you already use and build the line one station at a time.