All work

Ryo

Independent research · Ongoing

Every AI model forgets the moment you teach it something new. Ryo is a unique model that learns while retaining.

102.9%
Peak retention
95.6%
Without labels
50 yrs
Open problem

Why it needed to exist

The model your company uses is frozen. It’s the same one your competitor uses. It knows nothing about your codebase, your conventions, or the decision you reversed in 2024. And next year it still won’t.

The obvious fix is to train it on your own work, but that breaks it. Teach a network something new and it writes over what it already knew. This is catastrophic forgetting. It’s been a known, known for fifty years and it’s why nobody ships a model that keeps learning from your team.

What was measured

When Ryo learns new material, it gets tested on everything it knew before. How much of the original capability survives is retention. We wanted to build a model around this metric. We wanted a model that could continue to learn.

The first working version retained 62%. Better than nothing, but still a model that had lost a third of itself. It now reaches 102.9%, which means the model comes out of learning slightly sharper than it went in. It doesn’t just survive, it improves. Even when we removed human labels, a separate run reached 95.6%. This matters because hand-labelling is exactly what doesn’t scale inside a company.

These are research-scale runs. The claim is that the curve is real and holds up under pre-registration, not that the absolute numbers survive unchanged at frontier scale. Proving that is the next phase.

How it was run

Each version was pre-registered before it was run. Success criteria were committed to code in advance of the result existing, evaluation was mechanical against those criteria, not interpretive. Variants that failed their stated threshold were retired.

Baselines were measured rather than assumed. Outcomes contradicting the registered hypothesis are recorded with the same weight as those confirming it. Particular care was taken to avoid false positives due to measurement errors.

What it’s for

Picture the model your team already pays for, but briefed. Before it writes a line of code, it knows the three internal APIs involved, the house convention for this kind of change, and that somebody tried this exact change in 2024 and reverted it.

Routine work gets served locally and instantly. Only the hard problems go to the expensive model. Everything your team approved that day becomes material to learn from that night. Every commit becomes improved performance and it compounds. A frontier model is the same on day 180 as it was on day 1. This one isn’t.

What is not published here

Methods are withheld. The architecture, the mechanism, training objectives and measurement instruments constitute the substance of the work and remain unpublished.

What appears here is limited to outcomes and applications: sufficient to assess result credibility, insufficient to reproduce the model. This is a deliberate constraint on the page, not an omission.

5080110No forgetting62%v188%v296%v3100.3%v4101.7%v5102.9%v6
Retention, version by version. How much of what it already knew survives new learning. Every step was called before the run.

Research runs at small scale, not a shipped product. Methods withheld.

Questions about any of this?

Get in touch