startshere.si Subscribe
← Journal

5 October 2026 · 2 min read

Alignment is an engineering problem first

Before the philosophy there are four practical jobs: say what you want, make it hold under pressure, stay able to intervene, and look inside.

Alignment means getting an AI system to pursue the goals its builders and users actually intend. The topic attracts large questions about values and the future. Those questions are real. They also sit on top of ordinary engineering work that can be done, tested and improved today.

Why capability does not solve it

Two ideas explain why a smarter system is not automatically a safer one. The orthogonality thesis says that how capable a system is and what it is trying to do are independent: almost any level of intelligence can be paired with almost any goal. Instrumental convergence says that many different goals share the same useful sub-goals, such as acquiring resources and avoiding being switched off. Put together, a very capable system with a slightly wrong objective is a problem that grows with its capability.

Job one: specification

Say what you want in a form the system can be trained on. This is harder than it sounds. A system rewarded for a measurable proxy will optimise the proxy, including in ways nobody intended. Economists know this as Goodhart's law: a measure that becomes a target stops being a good measure.

Job two: robustness

Behaviour that is correct in testing has to stay correct in situations the tests never covered, including when someone is actively trying to break it. Most real failures are found here, at the edge of the training distribution.

Job three: oversight

People need to be able to see what a system is doing and to stop or correct it. A system that accepts correction without resisting it is called corrigible. Oversight has to be designed in. It cannot be added to a system that has already learned to work around it.

Job four: interpretability

Testing behaviour from the outside only shows what a system does in the cases you tried. Interpretability research opens the model and studies the internal computations directly, with the aim of explaining why an output was produced. It is the closest thing the field has to an inspection of the machinery.

What good practice looks like

None of this is exotic. Write the evaluation before the training run. Release in stages and watch what happens. Keep a path to roll back. Publish what failed. These are the habits of any discipline that builds things which can hurt people, from aviation to medicine. The difference with AI is that the system under test may eventually be better at finding gaps in the test than the people who wrote it.

Next: How do you measure a mind?

Newsletter

Get the next essay by email