The Didaflow white paper
Machine in the Loop
Returning knowledge and decisions to the people who own them
Where this comes from
Didaflow builds decision-support systems for education that put human expertise — not algorithms — at the centre of knowledge and decisions.
The question behind our work is older than AI. How do people build knowledge together — through practices, tools, and institutions? Anthropology asks it about cultures. Philosophy of science asks it about research communities. We ask it about schools and universities: the institutions whose entire purpose is the building of knowledge, and which today are being reshaped by technologies they did not design, cannot inspect, and struggle to govern.
Artificial intelligence was born as a tool for thinking about thinking — a way to understand intelligence by modelling it. Applied to education, that promise was extraordinary: computational models that help us understand how people learn, and use that understanding to teach better. The early history of the field delivered on it. Systems were built to test theories of cognition; when they worked, both the science and the classroom gained.
That link has been broken. Our work is about repairing it.
The Educational AI Divide
Look at educational AI today and you find a field split in two.
On one side, the epistemic cycle: the slow, careful work of understanding learning — theories, models, mechanisms, the questions educators actually ask. On the other, the pragmatic cycle: systems optimised to predict and personalise — dashboards, risk scores, recommendation engines — powerful, scalable, and increasingly disconnected from the understanding they were supposed to serve.
The divide is not an accident of engineering. It is the product of a chain: a theoretical retreat (the belief that learning is too complex to model), a technical turn (from models you can read to models you can only measure), and a market pull (prediction sells; understanding doesn't). None of these alone would have split the field. Together, they did — and they hold the split open.
The result: institutions that can see more than ever and understand less than ever. They know who is at risk with remarkable accuracy. They do not know what to change. We have built excellent thermometers, and very little medicine.
Three ways it fails
The divide produces three uncertainties that feed one another.
I. Technological uncertainty — How does it work?
Modern educational AI is opaque by construction. Predictions arrive without mechanisms; explanations, when offered, are computed after the fact by systems no more inspectable than the ones they explain. An educator cannot contest what they cannot see. A recommendation you cannot interrogate is not advice — it is an instruction.
II. Socio-technical uncertainty — Cui bono?
Behind every educational technology stands a network of actors — vendors, platforms, incentive structures — whose interests are built into the tool and travel with it into the classroom, invisibly. Opacity shields those interests; those interests, in turn, select for opaque systems that scale. The loop closes on itself. Black boxes are not only hard to read. They are convenient not to open.
III. Educational uncertainty — At what cost?
We measure what these systems do for grades and completion rates. We barely ask what they do to learners — to their capacity for judgement, their ownership of their own learning, their autonomy. When an apparently neutral tool quietly takes over the work of evaluating knowledge, the student gains an output and loses a faculty. And because our metrics only see the output, the benchmark improves while the learner erodes.
These three uncertainties are not separate problems awaiting separate fixes. Opacity hides interests; interests define what counts as success; eroded judgement loses the very capacity to contest either. The system does not self-correct. It deepens. That is why educational AI keeps disappointing the institutions that adopt it — and why the answer cannot be a better dashboard.
The response: Machine in the Loop
For a decade, the field's answer to these worries has been to keep a human in the loop — a person supervising an automated pipeline, approving what the machine proposes. We think this gets the geometry backwards.
Machine in the Loop (MITL) puts the machine inside the human loop. Educators, coordinators, and institutional leaders remain the authors of hypotheses, the owners of goals, the makers of decisions. The machine enters their process to do what machines do well: formalise the hypotheses experts state, simulate their consequences before anyone acts, confront them with evidence, and keep every assumption visible and revisable.
Three principles govern every system we build:
Epistemic precedence. Hypotheses come from people who understand the domain — teachers, programme coordinators, the research literature — before any data is touched. Data refine expert knowledge; they never replace it. The machine proposes nothing it was not asked; it never invents a number and never decides a causal claim.
Causal transparency. Every assumption is written down. Every effect estimate carries its source — a cited study, an institution's own historical data, a declared expert judgement — and its uncertainty. Anyone affected by a decision can trace the reasoning behind it to its origins. Where our own systems use opaque components, they are caged: constrained by explicit protocols, barred from making judgements, and named as limitations rather than hidden as features.
Intervention orientation. Education does not ask "who will fail?" — it asks "what should we change?" Our systems answer the second question: they estimate what happens if the institution acts — if it introduces tutoring, redesigns a course, reorganises support — so that decisions can be compared, simulated, and evaluated before and after they are made. Prediction ranks people. Intervention changes conditions. We build for the latter.
Each principle cuts one strand of the loop that holds the divide open: transparency removes the shield, epistemic precedence restores the authority to contest, intervention orientation takes back the definition of what counts as success. This is why it is one methodology and not three features: the failures are a system, and only a system-level response holds.
Where it was forged
MITL was not designed on a whiteboard. It grew out of doctoral research at the National PhD in Artificial Intelligence and years of work inside the University of Bologna — one of Europe's largest universities, and the oldest in the world.
The turning point was failure. Our group had spent years building dropout prediction models that worked: accurate, early, state of the art. And they changed almost nothing. Knowing who was likely to leave gave the institution no lever on why — and targeting individuals with risk scores raised harms of its own: labels that stick, prophecies that fulfil themselves, responsibility quietly shifted onto the very students the system was meant to help.
So we inverted the question. Students do not simply abandon universities; universities, through a thousand structural conditions, abandon students first. The object of analysis moved from the student to the system — from predicting the person to understanding and changing the conditions. That inversion demanded new machinery: explicit causal models authored with the institution's own experts, effect estimates grounded in evidence and local history, simulation before action, evaluation after it — and a discipline in which every step of that reasoning stays open to challenge.
That machinery is now software. Our platform walks educators and decision-makers through a structured scientific protocol — from a vague question to an explicit hypothesis, from hypothesis to simulated intervention, from simulation to evidence and back. It refuses shortcuts by design: it will not skip the step where a human thinks, it will not produce a number it cannot trace, and when local evidence disagrees with the literature, it does not silently pick a side — it puts the choice in front of the expert, where it belongs. Developed with and tested inside a real university on real institutional questions, it is the working proof that the methodology is not a manifesto. It runs.
What we believe
Educational decisions belong to educators. Not to models, not to vendors, not to us. Our systems are instruments in expert hands — they extend the reach of human inference and never absorb it. The measure of our success is not how much our tools decide, but how much better the people using them decide.
Values must be declared, not smuggled. Every educational technology carries a worldview: assumptions about what education is for, what counts as success, who is responsible for what. When these stay buried in design choices and training data, they act on students anyway — silently, unchallengeably. We believe institutions should have the courage to state their educational ideals out loud, precisely so they can be discussed, contested, and improved. Our tools are built to make that declaration unavoidable: goals written down, assumptions on the table, criteria in the open.
Responsibility must run upward, not downward. A field's technologies can quietly shift the burden of learning onto learners while presenting themselves as neutral services. We build in the opposite direction: our systems concentrate accountability on institutions and educators — the actors with the power to change conditions — and give them the instruments to carry that responsibility knowingly.
Trust is earned by honesty, not certainty. No model of anything human is certain. Trustworthy systems are not the ones that hide this; they are the ones that quantify their uncertainty, declare their assumptions, keep a public record of when they were right and wrong, and make revision cheap. We would rather tell an institution "the data are insufficient" than hand it a confident fiction.
Contestability is the feature. Any system used long enough acquires unearned authority; its assumptions harden into defaults. We design against our own ossification: provenance on every number, revision as a first-class workflow, and the standing invitation — to educators, to institutions, and ultimately to the communities they serve — to challenge the models that affect them. A tool that cannot be argued with does not belong in education.
And the same standards apply to us. Our platform includes components we did not build and cannot fully inspect. We name them, constrain them, and treat them as risks to manage rather than magic to market. The scrutiny we advocate is scrutiny we accept.
Towards learning institutions
Mature sciences earned public trust without certainty. Weather forecasting is trusted not because it is always right, but because its models are open, its uncertainty is quantified, and its track record is public. Climate science handles an even harder problem — a future that depends on human choices — by declaring those choices as explicit, revisable scenarios rather than pretending to predict them. Medicine built trial registries, evidence hierarchies, and revision procedures for a domain with no exact laws at all.
Education has never had this infrastructure. We are building it.
The horizon we work toward is the learning institution: a university or school system that treats every decision as an opportunity to understand itself better — where interventions are designed from explicit hypotheses, evaluated against their declared goals, and feed what they teach back into the institution's shared knowledge; where different experts' models of the same problem can be compared instead of fought over; where the question "how do we know this works?" always has an answerable form. Institutions that do not merely use intelligence, but practise it — collectively, transparently, and on the record.
This is a long project, and it is not ours alone. It belongs to every educator who refuses to outsource their judgement, every institution that would rather understand its problems than dashboard them, and every builder who believes technology should make experts more powerful rather than more optional.
Work with us
If you lead an educational institution wrestling with decisions that matter — student success, programme design, resource allocation — and you want understanding, not just monitoring: talk to us. Our platform is deployed with partner institutions and designed to learn each institution's own context.
If you are a researcher in education, learning analytics, causal inference, or human-centred AI: our methodology was born in research and stays accountable to it. We collaborate, we publish, and we welcome scrutiny.
If you are an educator or instructional designer asking what AI is doing to your students and your practice: those questions are the beginning of everything we build. We are developing open instruments to help you ask them rigorously.
If you build technology for education: the divide we described is a choice, and it can be unchosen. We are glad to compare notes with anyone building on the same side of it.
Didaflow — Bologna, Italy.