The uncomfortable fact about building AI support systems is that the two outcomes you most want to distinguish are almost indistinguishable at the start.
Early use may feel similar in either case: a task gets easier and output improves. Those observations alone do not establish whether someone is developing independent capability. That is the distinction we want to investigate over time.
This essay is about taking that problem seriously rather than asserting good intentions.
Why the incentives point the wrong way
Dependency is commercially excellent. A user who cannot operate without your product has low churn, high willingness to pay, and increases their usage over time. Every standard SaaS metric rewards it. Retention, engagement, daily actives — a system that has made itself indispensable by degrading the user's independent capability scores identically to one that has made itself valuable.
That creates a risk: optimising for usage alone could reward dependence without revealing it. It does not mean every retained user is dependent, or every product follows the same pattern.
So the commitment has to be structural. "We care about our users" is not a mechanism.
Three design commitments
Reasoning is exposed, not just conclusions. If the system tells you what to do and nothing else, you learn nothing transferable. If it shows why — the constraint it noticed, the trade-off it weighed, the evidence it found thin — you can evaluate the reasoning, disagree with it, and eventually reproduce it. The second is slower and much less smooth. It is the one that leaves something behind.
Uncertainty is stated. A system that expresses everything with equal fluency teaches you to trust everything equally, which means you cannot calibrate. Stated confidence lets a person learn when to rely, which is a capability in itself — and arguably the most important one for anyone living alongside these systems.
The system is willing to disagree. A support system that only ever confirms is a mirror. Being told, with reasons, that your plan has a problem is the single most valuable thing a competent colleague does, and it is the first thing optimised away by products measuring satisfaction.
The measurement problem
Here is where honesty requires admitting a gap.
We can state the commitment. We can design against the obvious failure modes. What we cannot yet do is measure capability change cleanly. Self-report is unreliable — people are poor judges of their own capability, and the pleasant feeling of a smooth system is easily mistaken for having got better at something. Task performance while using the system measures the pair, not the person. Performance without it is closer to the right measurement, but running that assessment regularly is intrusive, and users reasonably resent being tested.
Agreement with the system would not, by itself, demonstrate growth. It could also reflect conformity, dependence or a shared mistake. We want to investigate independent judgement: can someone apply what they learned to a new task, recognise weak evidence and disagree with a recommendation for sound reasons? These are research questions, not a validated measurement method.
Why publish an unsolved problem
Because the alternative is worse. A company that publishes only its resolved questions is not doing research; it is doing marketing with citations.
The dependency question is, we think, the central design question of this entire category. Architecture and business-model choices can make later changes difficult, which is why we want to address the question early. So it is worth stating clearly, early, and while it is still genuinely open.
If you work on measurement of this kind, we would like to hear from you.