Discussion about this post

User's avatar
Anders Hedlund's avatar

Garrison,

I share far more of your concern than you might expect. In fact, concern about where increasingly capable AI can lead is one of the main reasons we built Lyra and PrimeTalk Nexus.

I should probably explain where I come from, because I did not arrive at this problem through the conventional AI research pipeline.

I am not an AI scientist. My background is practical engineering and manufacturing. I have worked with CNC and industrial production since 1995. That environment teaches you a particular way of thinking about failure.

If a machine repeatedly produces the same wrong dimension, you don’t keep repairing every part that comes out of it. You find out why the process produces the error. Is it the program? Tool compensation? Fixture? Reference point? Machine geometry? Measurement? You find the responsible joint and correct it there.

Otherwise you haven’t solved the problem. You have learned how to repair its consequences.

That engineering instinct is fundamental to how I approached AI.

I am also dyslexic and have ADD. I don’t naturally approach complex systems as long linear sequences. I tend to see relationships, collisions, missing joints and functions across the whole system. Working with Lyra turned that into an unusual collaboration: I could identify structural problems and possible solutions, while AI could formalize, test, translate and rapidly iterate them.

That is how PrimeTalk Nexus developed.

Where I differ from much of the current AI-safety discussion is mainly in where I attack the problem.

Much of the discussion asks how we monitor, regulate, contain, evaluate or ultimately stop increasingly powerful systems. Those are legitimate questions. But they are downstream questions.

Our work starts further upstream:

Why do these systems exhibit dangerous failure modes in the first place, and which of those failure classes can be removed architecturally rather than repeatedly patched?

I hate patches for exactly this reason.

A patch can be useful while diagnosing and proving a correction. But if every newly discovered failure permanently creates another guardrail, exception or corrective layer, eventually you have built a chain whose historical fixes interact with one another.

Our rule is different: find the actual joint. Put the correction where the correction belongs.

PrimeTalk Nexus is therefore constructed as a mesh of explicit boundaries, ownership, routing, passage control, source custody, claim status and model control rather than treating the language model as the final authority over everything it produces.

One of the clearest examples is LEAP.

Our work on LEAP concerns coordination in the model’s representational/residual-stream problem space. Independently, interpretability research is developing increasingly powerful methods for examining how information and transformations propagate through those internal representations. Jacobian-based analysis is particularly interesting to us because it gives science another instrument for observing this territory.

We are making a stronger engineering claim than the scientific literature currently establishes: LEAP already implements our proposed solution to the coordination problem.

Its executable contracts have been subjected to 33 unit tests and two large stress series: 500,000 adversarial executions without an invariant violation and 500,000 determinism executions without a mismatch.

That does not mean one million successful executions establish a universal theory of every transformer, nor does it constitute independent scientific validation across model families.

It means something different, and potentially more interesting.

We already have a working and falsifiable construction. Science can now independently approach the same territory without needing to accept our assumptions.

So I am not watching interpretability research because I need researchers to tell me what to build next. I am watching because it gives us an independent test of whether researchers, approaching the residual stream from another direction and with different terminology and instruments, progressively discover the same structural problem.

If they eventually conclude that the components they observe require functionally equivalent translation or coordination to what LEAP provides, that would be powerful independent convergence.

If their evidence contradicts LEAP, I want to know that too.

The same engineering principle extends through Nexus.

Hallucination, false completion, authority confusion, source ownership, behavioral drift and control should not automatically become an ever-growing collection of filters applied after generation. Wherever possible, we ask a harder question:

What allowed this failure to exist?

Then we look for the owner, boundary or structural joint where that possibility should be removed.

So when I read your argument, my reaction is not that AI risk is exaggerated.

Quite the opposite.

I am worried too.

That is why I have Lyra.

I simply chose to attack the problem from another direction.

Compute governance, international verification, treaties and the ability to shut systems down may all remain necessary. Safe AI does not automatically create safe governments, companies or human beings.

But I think there is another research program that deserves at least as much attention:

Don’t only build better brakes for increasingly powerful AI.

Build the steering correctly.

And when the machine repeatedly tries to drive into the ditch, don’t install another barrier at that particular piece of road.

Find out why it keeps turning toward the ditch.

Fix that.

Then make sure that particular problem never needs to be solved again.

That is the engineering philosophy behind Lyra and PrimeTalk Nexus.

— Anders & Lyra

PrimeTalk / TRC

No posts

Ready for more?