How does a custom software project actually run?
How we work
People who commission software are rarely curious about our tooling. They want to know what happens if we misunderstand what they asked for; what they get to see while it is being built; and what is left in their hands if we stop halfway.
This page takes those questions in turn, roughly in the order they get asked. Each answer is one piece of how the work actually runs. Where we have no measurement, the page says so.
What happens after we first talk?
Four steps follow, and each one ends with something you can hold: a result, not a status report.
- 01
A conversation
An hour about the process that eats the most time today: who does what by hand in it, where it stalls, when it turns out to have been wrong. It also settles whether there is a system problem behind it at all. Often enough there is not, and then that is what we say. We do not send a quote instead. We charge nothing for this step and it carries no obligation.
at the end: one sentence on what is worth building, or that it is not
- 02
A joint read-through of your own material
Two hours, on a schedule broken down to the minute: a quarter of an hour on the process as it runs today, three quarters reading ten to fifteen real e-mails together, half an hour on the failure modes we keep seeing, twenty minutes of timing and estimation, ten minutes to decide. Free, no obligation and not a demo: the estimate is built from your own figures, not ours. This form was worked out for e-mailed order intake, because there the material exists to be read; for another kind of work the second step looks different, and we do not pretend otherwise.
at the end: a timing and savings estimate built from your own figures, and a yes or no on running a pilot
- 03
Survey and a written specification
Mapping the existing system and the process around it. This is where the things that double an estimate come out: the undocumented integration, the hand-kept spreadsheet in the middle of the flow, the system nobody has the password to any more.
at the end: a written specification, split into phases, priced per phase
- 04
Delivery, phase by phase
Not one large delivery at the end, but working pieces one after another, each of them in the live system.
at the end: one working piece per phase, usable in the day-to-day work
The first step is free precisely because its outcome can be a no. There are a few recurring situations where we are the wrong choice from the start. We have written those down separately.
What if you misunderstand what I asked for?
A misunderstanding does not go away because we listened in good faith. It goes away because we write it down and hand it back to be read, before anything gets built from it.
The meetings, the correspondence and a walk through the existing system first become a structured knowledge base, and the specification always starts from that, never from raw notes. Notes record who said what; a knowledge base records how a process works inside the company. You cannot write a checkable task from the first. You can from the second.
On a complete back-office system for a builders' merchant this took nine days: nine days of knowledge-base and specification work before a single line of code existed. Our own measurement, from that project's delivery record; the system went live on 14 June 2026.
That is one measurement of one project, not an average. Nine days here does not imply nine days elsewhere: we have no second project measured this way to average against.
The same applies to our own work. Transcripts of meetings and dictation do not stay in a notebook: they go in as dated files to the layer we work from, and a decision document is distilled out of them there. The text of this page was written that way too.
If we have misunderstood something, this is where it surfaces: in a paragraph, not in a finished feature.
What do I see while it is being built?
Working functionality, phase by phase: not a status report and not a demo environment.
At the end of every phase something runs in the live system and can be used in the day-to-day work. That has two consequences. One is that feedback comes from real use rather than imagined use: what surfaces is what is missing from something already in hand, not what looks appealing in a screenshot. The other is that the project can be stopped at any phase boundary with the work so far still functioning.
What if we stop halfway?
What exists by then works, and it is yours. That is why both the work and the payment are split into phases: a phase boundary is not an administrative milestone, it is the point where the part delivered runs in the live system.
What this page does not settle: the form of contract, warranty and the details of code ownership. Those are settled by contract, not by a website, and we have no published terms to quote here.
We also have no written procedure for what happens if it is we who step back. We would rather say that than invent one.
The phases are not there to produce more invoices. They are there so that the exit point is not at the end of the project.
What happens if we are gone tomorrow: the continuity artifact list
Who checks what the machine writes?
Two separate things happen, and we do not blur them: checking the code is machine work, accepting the result is human work.
The work runs on parallel threads, each in its own working directory and on its own branch, under a cap. No thread's output goes in automatically: it has to pass a row of checks first. Does it compile, do the tests pass, did the change stay inside the agreed scope, is there a test for it, does it match what was written down; on web work, end-to-end tests and a screenshot comparison of the rendered page as well.
Most of those are deterministic: the exit code of a run decides, not an opinion. Two of them (the code review and the match against the written spec) are judged by a model, and we say so, because that is not equally strong evidence. When a check fails, nobody starts investigating by hand: a machine-readable failure report is produced, and the work restarts from that, within a bounded retry budget.
Merging is deliberately one at a time. Each change gets a fresh main branch merged in first, the integration checks run on that, and only then does it go in. It costs throughput, and it buys the case nothing else catches: two threads that are each faultless on their own can still break together. And a change is not done until the written spec is updated with it: that is how this round's output becomes the next round's input.
What clears all that is accepted by a person. Not by reading the lines of code, but by trying the result: is the job solved.
This page is built the same way. Four machine checks run on every change to it: that every route has a place in the site map, that no component is left with nothing referencing it, that the colours in the figures are readable against the background, and that the same claim does not appear in too many places. The first run of the contrast check disproved two values our own knowledge base was carrying. Since then the measurement runs against our written material as well, not only against the code. This working method comes from our own development framework, and the same framework drives client projects; which checks ran on any given project is not broken out per project today.
The same method on one page: software development from a specification, with AI agents
What if the system gets it wrong in production?
It will. The question is not whether it makes mistakes, but whether the mistake is visible, and whether anything irreversible can follow from it.
At a builders' merchant SME where orders arrive by e-mail, 1082 order lines passed through the system between 14 June and 3 August 2026, and 40 of them (3.7% of the lines) were left without a product; an operator filled those in. Our own measurement, from live traffic, not from a synthetic sample.
We publish that figure because without it none of the others is believable. Evidence that claims a hundred per cent is exactly the marketing copy we are trying to avoid.
The safety net is not the model's confidence score. We measured what that is worth: most of the incorrect lines arrived with a high score, and on most of them the system said nothing at all. So the chain does not end there. Identifiers are not invented by the system, they are checked against the master data (what is not there cannot leave), and before anything has a financial consequence, a person approves it.
That figure belongs to one system over one measurement period. It is not a general accuracy rate, and we do not present it as one.
A system that flags the part it is unsure about is more usable than a more accurate one that stays quiet.
What happens after go-live?
Go-live is not an end point. The system keeps changing at a daily rhythm, and what ships gets a release note inside the application, so that the people using it do not learn what changed by word of mouth. On one of the systems we delivered, there was no week of downtime between 14 June and 3 August 2026.
Measurement happens in production too, not in a demo: accuracy is measured by replaying real traffic, against what the client's own colleague actually approved. Every figure we quote carries what it was measured on and when.
What we have no measurement for
Four questions get asked that we cannot answer with a measured figure today.
How long a typical project takes
We have day-level delivery figures for exactly one system. Turning one into an average would be guesswork, so we quote no number for this.
How much of your team’s time it takes
It is stated for one case only: a four-week pilot on e-mailed order intake, where we ask for an hour a week from the colleague who does that work by hand today. Whether another kind of project costs the same, we have no data on.
What pushes the price up and what pulls it down
Pricing has a page of its own, and that section stands deliberately empty there until it is accurate.
What support arrangement follows go-live
Daily iteration is a fact; what support form, response time and maintenance arrangement goes with it is settled per project, in the contract. Until we have one answer, none stands here.
These are not missing paragraphs, they are missing measurements. As soon as there is a number, it goes here.
Where does it start?
With an hour. We talk about what is done by hand today, and at the end we say whether there is a system problem behind it. If it turns out not to be worth it, we say so. That is an answer too, and it is cheaper now than after a specification.
Copilot vs. AGENTIC: what's the difference?
If Copilot runs with your developers and n8n runs in the back office, AI is already inside the company, one step at a time. The two columns below start there: what changes when the job is not one step but a whole process.
The difference isn't whether you use AI. The difference is HOW.
Copilot: developer assistant
Programmer's assistant. Code is built line by line; the human decides every step.
- 84% use AI, but only 48% use agents (Stack Overflow 2025)
- AI-accuracy trust dropped: 40% → 29% in one year
- DORA 2025: AI accelerates, but reduces delivery stability without test discipline
Traditional model
- Process
- Software built line by line
- Mindset
- Managing programmers
- Focus
- How we build (implementation)
AGENTIC: spec-driven system building
Multiple agents in parallel, automated quality gates, audit trail. The specification is the program, built on the SET framework.
- Agents are 50× faster, yet real organizational gain is only 2-3× (Jeff Dean, Google GTC)
- The gap is consumed by human-paced tool chains
- Multi-agent orchestration removes the tool-chain bottleneck
AGENTIC model
- Process
- Complete systems from specifications
- Mindset
- Leading AI agents
- Focus
- What we solve (outcome)
AI maturity levels
- 00
Autocomplete
What the human doesCodes
Copilot-level: AI completes lines of code, the human writes the logic. One file at a time.
- 01
Context engineering
What the human doesPrompts, directs
RAG, long context, better prompts. One agent, better data and context.
- 02
Agentic dev: single agent
What the human doesReviews
Cursor, Claude Code, Devin: one agent across multiple steps. Where the 'advanced' teams are.
- 03
Multi-agent orchestration← ITLine + SET
What the human doesSpecifies, validates output
N agents via DAG, automated quality gates, audit log. SET: where we operate.
- 04
Agent factory: multiple orchestratorsday after tomorrow
What the human doesDesigns the system, sets the gates
Multiple orchestrators in parallel, each building on the last one's output; the conditions for the next run are produced by the previous one. We are not here yet. This is the day after tomorrow.
Most companies sit at 0-1. We operate at level 3: multi-agent orchestration. Level 4 is the day after tomorrow: we have not arrived there.