The first attempt produces an outcome. The research question is whether it also produces a better learner.
A working definition
We use self-learning to describe improvement that comes from a system’s own experience: acting in an environment, observing the result, updating a working understanding, and using that change in a later encounter.
This is intentionally a demanding definition. A system has not necessarily learned because it completed a task, retrieved an old answer, or generated a reflection. Something about its future behavior must change—and the change must remain useful when the surface details do.
Learning under uncertainty
Unfamiliar environments do not provide clean lessons. Their rules may be hidden, their feedback delayed, and their apparent patterns temporary. The learner must act before it fully understands.
That makes uncertainty part of the mechanism, not an inconvenience to hide. A self-learning system should be able to hold competing explanations, design bounded tests, and recognize when the world has contradicted what it believed.
The loop around the model
Self-learning may not live inside a single model call. It may emerge from the wider system around the model: the interface that exposes consequences, the record that preserves evidence, the process that turns observations into hypotheses, and the evaluation that tests whether those hypotheses help.
We are exploring the shape of that loop:
- Act where the result can reveal something.
- Separate direct observation from possible explanation.
- Test the explanation against another encounter.
- Carry forward what survives, with its uncertainty intact.
- Revise or forget it when the evidence changes.
What should persist?
Not every event deserves to become knowledge. Some details belong to one episode. Some strategies recur. Some beliefs appear stable until the environment changes.
The problem is not simply persistence, but selective persistence: choosing what to retain, how to represent it, when to retrieve it, and how to keep it corrigible. Memory is useful only when it serves that learning process.
How would we know?
Self-learning cannot be established by a compelling transcript. It must be visible across attempts. Does the system reach a useful hypothesis sooner? Does it recover after a belief fails? Does learning transfer to a related environment without overwhelming new evidence? Does it know when the old lesson no longer applies?
Evaluations should make these changes observable. The interesting unit is not one score in isolation, but the trajectory of behavior as experience accumulates and the environment pushes back.
Questions we are pursuing
- How can interaction produce evidence that an agent can actually use?
- How should working beliefs be represented, tested, and revised?
- What enables learning to transfer without turning into overconfidence?
- Which evaluations distinguish self-learning from retrieval, repetition, or chance?
Where the work begins
The goal is not a machine that starts with every answer. It is a system that can meet the unknown, learn something defensible from the encounter, and return changed in ways that can be tested.
Self-learning begins when experience becomes more than memory. Everything after that sentence is still an open question.

