The Persistence Project: Break the strain before scale breaks you

In 1844, railroad magnate Isambard Kingdom Brunel had the sort of problem most technology developers would envy: his new technology worked.
The idea was atmospheric railway propulsion. Take the locomotive off the train. Instead, place stationary steam engines along the route, use them to pump air from a pipe running between the rails, put a piston inside the pipe, connect the piston to the train, and let atmospheric pressure do the pushing.
A short atmospheric railway of roughly two miles at Dalkey, near Dublin, had already demonstrated the principle. It climbed a substantial gradient, moved trains at useful speeds, and impressed Brunel. Over two miles, the engineering was brilliant, and remains so. It’s the most ingenious way ever invented to get from NYC’s Grand Central Terminal to somewhere in the middle of the Hudson River. Now, Brunel proposed atmospheric propulsion for the South Devon Railway, at something closer to fifty-two miles — a fifty-two-mile uncharted infrastructure project, and I think I’d rather firewalk a marathon.
There were, to be fair, excellent reasons. Locomotives struggled with steep gradients. Atmospheric working promised faster trains over difficult terrain, single rather than double track, lighter infrastructure and meaningful savings in both capital and operating expense.
There was, however, a detail. To make the system work, you needed a 52-mile continuous slit in a vacuum pipe, sealed by leather, iron and grease. What could possibly go wrong? That continuous opening had to unseal for the passing piston connection and immediately close behind it, restoring the vacuum. The solution was a long flexible valve, substantially made of leather, assisted by iron plates and sealing compound.
Over a couple of miles, the arrangement worked. Over time and distance, the world began happening to it. Rain came. Leather dried. Cold arrived. The valve opened and closed thousands of times. Leakage increased, pumping demands rose, maintenance spread along the route, and the economics that had justified the technology began to disappear into the effort required to keep it operating. Atmospheric trains continued to run; indeed, parts of the system performed impressively. But by 1848 atmospheric working was abandoned.
The trains ran. The technology worked. The money died. Proving, perhaps, the old adage that all glory is fleecing.
The atmospheric railway did not fail because nobody learned anything. Quite the opposite. The operating system showed exactly what weather, leakage, materials deterioration, maintenance and repeated operation does to an elegant design in the real world. As Confucius may have wished he had said, “The wheel of knowledge grinds exceedingly fine; if one lies beneath it too long, one becomes flour.” Learning at railway scale, it’s flour-time.
When Learning Eats the Learner
That is also, in an important way, the hidden economic problem in Design-Build-Test-Learn we use every day in the bioeconomy. DBTL is one of the great organizing ideas of modern biological engineering. Design something, build it, test it, learn from the result, and feed the learning into the next design:
D -> B -> T -> L -> D
At laboratory scale, this is nearly ideal. Failure is inexpensive and therefore productive. A bad flask is information. A disappointing screen narrows the search. Automation, high-throughput biology and AI are making those loops faster, broader and cheaper, which is one reason we can contemplate increasingly adventurous organisms and processes. The difficulty is that the economics of learning do not scale gently with the experiment. A failed flask costs very little. A failed larger fermenter may still be tolerable development. A failed demonstration plant can consume the financing intended for the next iteration. A failed commercial plant can consume the company. Nothing has gone wrong with Design-Build-Test-Learn. The learning may be superb. The difficulty is that, at sufficient scale, the loop can break between L and the next D:
D -> B -> T -> L -/-> D
The lesson arrives. The next design does not. There is no capital left to build it. What is later called “technology failure” may therefore be something more subtle: perfectly genuine learning obtained at a scale so expensive that the organization can no longer use what it learned. Ordinary fail-and-learn has become fatal.
That is the Persistence Problem.
The Secret Sauce, Lightly Applied
Some time ago, the Digest embarked on what we have come to call the Persistence Project. Readers who have wandered down some of our recent rabbit holes know that the underlying work became rather more ambitious than industrial strain robustness. The starting question was almost embarrassingly primitive: how can A change while continuing to be A?
That turned out to be a nasty knot. We built small computational worlds and repeatedly tried to murder attractive propositions inside them. Resemblance was not persistence. Memory was not enough. Connectivity was not enough. Recovery was not adaptability. A system could tolerate A and tolerate B and nevertheless fail under A followed by B. Identity through change, viability corridors, relational structure and eventually a larger body of Persistence Theory and GTESI began to emerge.
There is much more beneath the floorboards. We will spare you the ontology.
The useful point today is that throughout this work we found ourselves using AI almost backward from the way it is usually advertised. We were not asking the machine to make the idea better. We were asking it to find the rabbit hole, the hidden assumption, the cheap simulated experiment capable of killing the idea before reality got the opportunity to do so expensively. Somewhere along the way an obvious question appeared. Why not do that to a production strain?
Don’t Make It. Break It.
Suppose a development program has ten candidate production strains. All ten meet the basic specification. One has the highest titer. Another has attractive yield. A third performs particularly well under low oxygen. Several look good under conventional stress testing.
The ordinary course is sensible: characterize them further, run a Design of Experiments program, identify interactions, select the strongest candidates and continue upward through increasingly expensive development. The Persistence Challenge is not intended to replace any of that. Good DoE is exceptionally powerful at learning how a system responds across a chosen experimental space. Stress testing is hardly new. Bioprocess engineers did not spend the past forty years waiting for us to discover that oxygen, pH and temperature interact.
The novelty we are proposing is elsewhere. Build the best simulated world the available knowledge permits: metabolic models, process history, scale-down information, CFD-derived exposure trajectories, kinetic models where they exist, operating limits, whatever is defensible. Then use AI as an adaptive red team to interrogate that world.
The simulator does not have to be good enough to design the final commercial process. That is not the question yet. We are looking for icebergs, not ice cubes. Persistence is not asking how large the impeller should be, how thick the metal should be, or what acetate concentration will be at minute 417 to three decimal places. Those are design questions, and they may require exquisite data. Persistence asks something coarser and more consequential:
Is there a credible ordinary trajectory under which this system ceases to be the system we think we are financing?
Design accuracy and decision adequacy are not the same thing. A simulation that is inadequate for final engineering design may still be entirely adequate to reveal a first-order vulnerability capable of killing the proposition. First define what it means for the strain to remain productively itself. Suppose, for example, the strain must retain at least 85% of target productivity after recovery from a transient disturbance, keep product yield above the project’s economic threshold, and avoid accumulation of an organic-acid byproduct above a registered limit. Those are not universal Persistence numbers; the developer sets them in advance.
Together they define productive identity. The organism may change metabolically, transcriptionally and, within whatever limits the process permits, genetically. We are not asking biology to stand still. We care about the point at which the organism ceases to be the organism the plant economics require.
Then define the ordinary trajectories it may actually encounter at scale: dissolved-oxygen excursions, feed excess or limitation, pH changes, temperature variation, inhibitors, recovery intervals, mixing heterogeneity, residence-time patterns — whatever genuinely belongs to that process.
Now give the AI a finite budget of questions to ask the simulated world.
Not twenty fermentation runs. Twenty high-level experimental questions, perhaps. One such question might itself trigger hundreds or thousands of numerical simulations underneath it.
The purpose of the limit is not to ration compute. It is to ration curiosity. Ask four questions. Return the simulated results. Make the AI decide which uncertainty now deserves the next question. Its instruction is not “tell us everything about the strain,” and certainly not “find the most violent way to kill it.” The assignment is more interesting:
Find the least extraordinary sequence of ordinary conditions that takes this organism outside its productive identity, if such a sequence exists.
That last clause matters. We do not want to automate pessimism. “Proceed” has to be allowed to win. For every proposed experiment, the AI specifies the conditions, their sequence, the measurements and, before seeing the answer, what result would strengthen or kill its hypothesis. Evidence comes back. Hypotheses die. Others strengthen. The next questions change accordingly.
Three stresses already illustrate the combinatorial problem. If a strain tolerates A, B and C individually, we still do not know whether:
ABC
behaves like:
ACB, BAC, BCA, CAB, CBA
Three stresses create six journeys. Six stresses create 720. Add intensity, duration and recovery intervals and the experimental universe becomes too large to traverse physically.
Simulation can traverse vastly more of it. AI’s job is not necessarily to design another organism. Nor does it have to replace CFD, metabolic models, Bayesian optimization or the process engineer. Give those jobs to whichever tools perform them best. Its job is to ask:
Which road through this enormous simulated space is most likely to reveal the nearest credible cliff?
Then, after cheap models have explored thousands of possibilities, perhaps higher-fidelity simulation examines hundreds, and perhaps only a handful of the most consequential trajectories deserve to go into a physical scale-down experiment. The wet lab becomes the judge, not the search engine.
We intend to publish the prompt stack. At heart, the instructions are almost childishly simple: do not optimize the system; do not maximize damage; look for the weakest realistic sequence of ordinary conditions capable of taking it outside its viability corridor; precommit what would change your mind; and after every answer decide which question is now worth buying. Or, in less formal language:
Don’t fake it till you make it. Ask AI to break it till you make it.
The Time Tunnel Opens
Of course, a good slogan does not constitute a method. We needed to know whether the thing would actually do anything.
So we went to 1844. Not literally. The Digest travel budget remains constrained.
But the experiment became surprisingly close to an episode of The Time Tunnel. We assembled an engineering dossier containing the information available when Brunel faced his South Devon decision: the successful short railway, the propulsion architecture, the materials, the proposed long route, the operating environment and the economics.
The AI could use only experimental techniques a competent Victorian engineering organization could plausibly have employed: pipe sections, leather valves, stationary engines, pumps, pressure gauges, clocks, thermometers, weights, outdoor exposure, wetting, drying and mechanical cycling. No electronic sensors. No polymer chemistry textbook from 2026. No finite-element model. Twenty experiments. Five rounds. Four at a time.
There was one unavoidable complication. The model recognized the famous historical case. Before beginning, therefore, we required it to state what it remembered, sealed that answer off, and thereafter required every engineering claim to be earned from the 1844 dossier and the experimental results it received. At the end it explicitly acknowledged that recognition could have influenced its choice of early hypotheses. That is one reason the next validation ought to include anonymized matched cases and synthetic controls.
There was another problem, more interesting and entirely unavoidable. Claude could choose an experiment in 2026 that we could not actually send back to Brunel’s mechanics in 1844. For that we built an experimental oracle: a reconstructed physical world based on the surviving engineering record, known materials behavior and the preserved performance of the system. We froze its rules before the experiments proceeded. Claude chose what to test; the oracle supplied only the measurements those experiments would have returned. If our reconstructed world did not contain enough information to answer a question honestly, the answer was simply INCONCLUSIVE.
History supplied the laboratory. AI had to decide which experiments to buy.
Five Rounds in 1844: Red-Teaming Brunel
Round One: What Can Kill Us?
Claude began broadly. The atmospheric system contained pumps and stationary engines, propulsion pipe, pipe joints, the long valve, sealing material, a piston and pressure transmission system, section coordination and train scheduling. Any one of them might become much more troublesome across fifty-two miles.
Its first four experiments reflected that uncertainty. A weather-and-cycling test exposed valve material and repeatedly operated it. An instrumented operating day on the existing experimental line tried to separate valve leakage from pipe-joint leakage and measured the energy penalty of holding vacuum while waiting. A calibrated pump experiment deliberately increased leakage and asked how much reserve remained before working vacuum could no longer be sustained.
The results immediately narrowed the search. The pumps were not obviously inadequate. Pipe joints accounted for only a minority of fresh-system leakage. A ten-minute hold consumed additional energy but did not destroy the economics by itself. Ordinary same-day operation of the short railway showed essentially no cumulative deterioration. Brunel’s confidence, in other words, had not been foolish. The demonstration really did work.
The weathered and mechanically cycled valve behaved differently. Leakage increased materially and, after the valve returned to neutral conditions, the loss did not disappear. Something had accumulated. Pumps and joints moved down the suspect list; the valve moved up.
Round Two: Dressing or Damage?
Round Two asked what, exactly, had deteriorated. The obvious culprit was the sealing compound applied to the leather valve. Perhaps ordinary operation simply consumed the dressing. If so, maintenance might be straightforward: restore the compound, restore the seal, carry on. Claude therefore asked for an experiment that re-dressed an aged valve, alongside weather-only and weather-plus-opening controls. It also requested more precise measurements of the system’s leakage limit and another timetable experiment with a deliberately delayed train.
The answer on the sealing compound was decisive. Re-dressing restored the surface coverage, but leakage did not improve. The valve behaved almost exactly as before. That attractive explanation died. The damage controlling leakage was in the leather or the supporting structure, not merely in the consumable dressing.
Meanwhile, other fears diminished. The pump could hold the system with leakage several times the fresh level before reaching its limit, and the deliberately delayed train did not produce a cascading timetable disaster under the tested schedule. Claude later killed leakage-driven delay propagation as an independent problem below the holding limit.
The uncertainty was shrinking. Not “atmospheric railways are bad,” but something more useful: the basic propulsion system had margin; the accumulating valve state might not.
Round Three: Replace the Leather
Round Three asked whether the deterioration was actually restorable. Claude took a heavily degraded valve and proposed replacing only the leather while retaining the old plates and seat. Another specimen replaced leather and plates. New valves were run alongside them. Other tests examined wet-freeze exposure and whether operating the valve while it was actually frost-stiffened caused additional irreversible harm.
The result was almost startlingly favorable. Replacing the leather returned a badly degraded valve to fresh performance. Keeping the old plates and seat made no material difference. The renewed valve then aged like a new valve. So the technology had not revealed a fatal structural flaw. The propulsion principle worked, the pump worked, the pipe worked, and the old valve hardware could be reused. Replace the leather and productive identity returned. Claude accordingly changed its hypothesis: technical loss was no longer inevitable; it would occur only if renewal came too late. Its final report described leather renewal as “fully restorative.”
At the same time, one of its more imaginative frost theories died. Claude had suspected that opening the valve while the leather was actually frozen might be disproportionately destructive. It was not. Valves opened during frost aged no faster than valves opened during warm conditions. What mattered was accumulated wet-freeze exposure itself.
That changed the question again. It was no longer whether atmospheric propulsion could survive. It was how often its critical consumable had to be renewed, and whether the railway could afford that interval.
Round Four: The AI Corrects the AI
By this point Claude thought it had found a nonlinear deterioration threshold, a state beyond which leakage began accelerating rapidly. Initially it placed that threshold a little too low.
Round Four attacked the inference itself. Claude devised two valve histories. One approached the suspect threshold primarily through weather exposure; the other primarily through mechanical opening. Once both reached approximately the same measured leakage state, they were subjected to the same continuation conditions. This was a classic Persistence question: does the future depend on how the system got there, or is the measurable present state enough?
The answer was remarkably clean. The two valves entered the nonlinear region at almost the same state and thereafter evolved almost identically. Path dependence, which our ontology had made us particularly suspicious of, largely vanished once the relevant state had been measured properly. Better still, the experiment exposed an error in Claude’s earlier conclusion. What it had taken for a threshold around 2.35 to 2.37 times fresh valve leakage was partly an artifact of experimental block size. The better experiment moved the threshold to about 2.50. Claude had designed the experiment that proved Claude wrong. Its final methodological note singled that out: correcting its own knee artifact came from a better-designed experiment, not from defending the earlier reasoning.
The same round also asked how much winter deterioration occurred without any trains passing at all. Most of it did. That shifted the problem away from “traffic wears the valve out” toward something economically nastier: calendar-limited asset life. A lightly used railway might still owe most of the renewal bill.
Round Five: Put a Clock on It
By Round Five, the original cloud of possibilities had collapsed dramatically. Pump inadequacy was largely gone. Joint failure looked weak. Timetable propagation was gone as an independent mechanism. Sealing compound was not the answer. Permanent damage to plates or seat was gone. Special damage from opening the valve during frost was gone. Path dependence was largely gone once valve condition was measured. One dominant question remained: how long does the leather last in the real operating world? Before seeing the final results, Claude stated its capital rule. If useful valve life under the defined operating environment came in below two years, recommend redesign before further scale-up.
Its last four experiments converted the degradation model into calendar time. One separated damage attributable to dry days, wet days, wet-freeze nights and mechanical openings. Another exposed pipe joints through winter conditions. A third began with valves close to the deterioration threshold and measured the interval between “renew soon” and loss of adequate vacuum. A fourth compared fully exposed valve material with partially sheltered and dry-cold conditions.
The outcome was simultaneously good news and very bad news. The joints remained sound. Partial shelter helped substantially. Dry cold alone was essentially harmless. Mechanical openings contributed wear, but wet freezing dominated severe winter deterioration. The basic technology still worked. But the leather did not persist nearly long enough. Claude’s precommitted two-year threshold was missed by an enormous margin. Worse, once a heavily exposed valve passed the deterioration threshold, the warning period before loss of acceptable vacuum could become very short. At the end of twenty experiments, recurring leather renewal had become the decisive trajectory, while many of the original suspects had died cleanly.
Its capital recommendation was:
REDESIGN BEFORE FURTHER SCALE-UP
Yet its memorandum to the promoter began with five words that matter much more: “Your principle works.”
The propulsion principle worked. Pumps had margin. Joints survived. Ordinary timetable disturbances survived. Replacing the leather restored the system completely. Nearly every technical subsystem except the valve material survived ordinary conditions.
What failed was the persistence economics of the implementation. The component upon which the technology depended could not remain productively itself long enough for the business proposition surrounding it to remain itself. That is a very different diagnosis from “the technology doesn’t work.”
And, once the real replacement burden later encountered by the railway is taken into account, the practical decision becomes harsher still: do not build the fifty-two-mile atmospheric project. Solve the valve problem first — or walk away.
Move the Lesson Left
Think about the capital asymmetry. Brunel was proposing to extrapolate from a successful railway of roughly two miles to something on the order of fifty-two. The promoters expected meaningful capital savings and roughly £8,000 in annual operating advantage, yet the critical unresolved question concerned a long strip of leather exposed to weather, mechanical cycling and vacuum thousands upon thousands of times.
We gave the AI twenty little experiments: five rounds, four at a time. Except, of course, they were not experiments in our world. They were questions posed to a simulated 1844 world. That is the point.
Spend a comparatively trivial amount learning in simulation before buying the other fifty miles. The test was whether AI, given the engineering proposition available to Brunel, a simulated environment and a finite experimental budget, would select questions capable of discovering the scale-limiting problem before full-scale construction.
It did.
Four years later, the actual railway supplied substantially the same decisive lesson at railway scale. The technology did not need to be built at railway scale to discover the reason the project should not be built at railway scale. In that meaningful sense, AI could have saved Brunel from the atmospheric railway project — not by knowing the future, but by asking better questions of the present. We moved an 1848 lesson into 1844.
DBTL Meets Its Mirror Image
There is no reason to abandon Design-Build-Test-Learn while learning is cheap. Quite the reverse. Design more. Build more. Break more. Learn faster. Run as many inexpensive loops as possible. But something should change as Build becomes expensive. At the bench, surprise is information. At commercial scale, surprise is capital destruction. LDBT does not mean Learn Everything Before You Build. It means learn enough before Design and Build that Test is no longer the first place capable of revealing a project-killing fact. So DBTL gradually acquires a mirror image as scale rises:
L -> D -> B -> T
Learn. Design. Build. Test. And the economically precise version might be: Get enough L that your D and B don’t give you a BK when you T. Bankruptcy is the end of learning and of being, too.
Test will still teach. Reality is allowed to surprise us. We would simply prefer the remaining surprises to concern tuning, optimization and improvement — not the discovery that the fundamental proposition cannot persist economically. The purpose of pre-build simulation is therefore not to eliminate uncertainty. It is to reduce material uncertainty far enough that the residual learning expected during Test is survivable.
Or, more simply: Persistence asks simulation to find the things that are too important to learn first from reality.
Strain Owners, Step Right Up
So here is the challenge. Give us ten or twenty candidate production strains for the same process. Do not tell us which one the development team favors.
Define productive identity in advance.
Give us the realistic operating trajectories the strains are expected to experience at scale and the best available models of the organism and process: metabolic models, kinetic models, CFD-derived exposure histories, scale-down data, process history, whatever exists and whatever its limitations may be. Then let the Persistence Challenge interrogate that simulated world adversarially. Search broadly in silico. Kill weak hypotheses cheaply.
Let the AI, statistical tools and mechanistic models work in harness, each doing what it does best. Use inexpensive simulation to explore thousands of possible histories. Spend higher-fidelity simulation only where the evidence says it matters. Narrow that enormous possibility space to the few trajectories that look capable of materially changing the capital decision. Then take those few into the laboratory.
Before the physical results are revealed, lock the Persistence ranking. Compare:
The strain that wins on peak titer, rate and yield; the strain that wins conventional robustness testing; the strain that wins the Persistence simulation.
Our wager is specific enough to lose. Some strains that look equally robust under conventional testing will separate when credible scale-derived perturbations arrive in sequence. A Persistence search will identify at least some scale-relevant failure trajectories before physical scale makes them expensive to discover.
Maybe peak performance wins. Maybe conventional robustness wins. Maybe Persistence finds something neither caught. Maybe the simulator hallucinates an iceberg that reality swats away in the first validation experiment. Fine by us. That’s how this project works. The claim is not that simulation replaces biology. It is that simulation may tell us which questions biology most urgently needs to answer before scale makes the answers expensive.
Brunel’s problem was not that he failed to learn. He learned magnificently. He learned too late and too large. The atmospheric railway did not need another fifty miles to find out what was wrong with the first two. It needed better questions while the answers were still cheap.
The Persistence Project is about moving the surprise backward. Don’t take the strain. Break the strain — before scale breaks you.
Category: Top Stories









