Project

Rewriting 5,000 Apps

How a five-person team at Phorest re-platformed every whitelabel salon app — by turning a porting project into a generation pipeline. Bookings through the new apps up ~10%.

The problem

Phorest gives every salon its own branded consumer app as part of the package — their name, their look, their bookings, live in the app stores. That’s a wonderful promise and a brutal engineering liability: thousands of salons meant thousands of apps on an aging mobile stack, where every improvement, fix and OS update carried a per-app cost. The product was slowly being crushed by its own distribution model.

The obvious plan — rebuild them on a modern stack — looked impossible on paper. The core team was five people: an architect, a junior engineer, a designer, a QA engineer, and me. Against ~5,000 apps.

The bet: generate, don’t port

We stopped treating it as a porting project and treated it as a generation problem. One modern React Native codebase, with every salon’s app produced from it by an automated pipeline — branding, theming, configuration, assets, store packaging. The unit of work stopped being “an app” and became “the machinery that makes apps.”

That reframing changed what the five of us actually spent time on. The job was never to build 5,000 apps; it was to make the generator’s output trustworthy enough that a human only touched the exceptions.

The real product problem: trusting generated output

Once generation works at all, the whole project hinges on one question: is this generated app good enough to ship without a person reworking it? Answering that at scale meant watching where the pipeline’s output failed, tightening it against each failure class, and driving manual intervention down from the norm to the exception. Every manual fix was treated as a bug in the machinery, not a task to staff.

In 2026 vocabulary, we were managing a touchless rate — the share of generated output needing zero human edits — years before AI products made that a standard metric. The discipline is identical: define quality for machine-generated output, measure it, and let nothing ship on vibes.

Research on the shop floor

None of this was done from behind a dashboard. I was in salons weekly with prototypes in hand — watching owners and staff hit the booking flow on their own counters, between clients. That loop set the quality bar for what “good enough to ship” meant, because the people who’d live with a mediocre generated app were never abstract.

Results

  • ~5,000 apps rewritten onto the modern stack — by five people
  • Bookings up ~10% through the new apps
  • A maintainable system where improvements ship to every salon at once, instead of 5,000 times

It remains the piece of work I’m proudest of: a small team overachieving not by working more, but by refusing to do the same work five thousand times.

Why this story matters more now

The defining product problem of 2026 — agentic and generative systems creating output at a scale humans can’t review piece-by-piece — is this story’s shape exactly. When I build evaluation harnesses for generative pipelines or design systems where my team’s work compounds, I’m applying the same principle this project taught me: automation earns trust through measurement, and the metric that matters is how rarely a human has to step in.

The team wasn’t five people building apps. It was five people teaching a machine what good looked like.

Want the longer version — how the pipeline was structured, what we got wrong first, what I’d do differently with today’s AI tooling? Get in touch.