An AI systems firm: rescue what is broken, build what is traceable. Illustrative.
AI system rescue & maintenance
For a system that already exists and is no longer behaving.
The signals that bring someone here
- Output quality degraded quietlysignal
Nobody noticed until customers did.
- The original builder is gonesignal
A freelancer or agency can disappear once paid.
- Nobody can explain the pipelinesignal
Not even the team that owns it.
- Costs crept up with no clear causesignal
- Prompt or model changes break something elsesignal
- No monitoringsignal
So failures surface as complaints rather than alerts.
How a rescue runs
- 1Diagnose
- 2Stabilize
- 3Write it down
- 4Fix
- 5Monitor
- 6Maintain or hand back
Diagnosis happens before any code changes, and it is documented rather than performed. Stabilize first — a risky rewrite is a last resort, not a first move. Ownership transfers cleanly if you want it, with no dependency on the firm to keep it running.
What a diagnosis looks like
Sample findings. The written diagnosis belongs to the client whether or not the engagement continues.
Sample findings
- No evaluation setfinding
Nothing told the team whether a change helped or hurt, so every release was a guess with a deploy attached.
- Silent fallbacksfinding
Failures were swallowed rather than recorded, so the system looked healthy while it degraded.
- No cost guardfinding
Spend could only be discovered at the end of the month.
- Prompts untrackedfinding
The behaviour in production could not be tied to any version of anything.
A sample engagement
- week 1Diagnosis deliveredEuclid
written, with the evidence behind each finding
- week 2Evaluation set builtEuclid
so later changes can be judged rather than argued about
- week 3Tracing addedEuclid
failures become visible before customers report them
- week 4+Handover or retainerclient
the client's choice, not a default
AI-native system builds
For when there is no system yet, or the existing one was never architected for production.
The signals
- AI wired in as an afterthoughtsignal
On top of an existing product rather than through it.
- Needs to be reliable, not demoablesignal
An agentic or RAG system that holds up under real usage.
- No in-house AI engineering capacity yetsignal
- Previous attempts fell apart under loadsignal
Worked in a demo, failed in production.
- Needs ownership through launchsignal
Not a prototype and a goodbye.
The engagement
- 1Scope
- 2Architect
- 3Build
- 4Deploy
- 5Maintain
Scoped before it is architected — the problem and the success criteria come first. Evaluation and tracing are designed in from day one rather than bolted on after launch, and every component has to earn its place rather than arrive with a template.
What the firm builds for itself
The products exist because the engagements kept running into the same missing pieces.
| Product | What it does | State |
|---|---|---|
| Cascaid | predicts cascading failures in AI pipelines | on PyPI |
| LocalForge | on-device code review before a commit lands | released |
| Bernn | cloud and AI spend for solo developers | building |
Most observability tools trace a failure after it happens. Cascaid was built to flag it before it cascades — which is the same argument the rescue work makes, turned into a package.
The case the firm makes
The alternatives
- A freelancer or agencycan disappear once paid
- An in-house hiretakes months
And leaves with the knowledge.
- Most toolingtraces after the fact
What is promised instead
- Written downalways
The system stops being a black box to the team that owns it.
- Traceableby design
- Transferableon request
Continuity is an option, not a lock-in.
Rescue and build are two shapes of the same engagement: find out what the system actually does, write it down, and make the next change safe to make.