A useful result because it separates two things
The usual education argument asks whether students should use AI or learn to think for themselves. A new randomized trial at Bocconi University gives a more practical answer: those interventions may change different parts of the work.
The trial involved 1,053 first-year students in economics, finance and management. They worked on a real marketing case for the university’s merchandise store. By class period, they were assigned to one of four groups: access to ChatGPT Edu with GPT-4o, a short exercise in causal reasoning, both, or neither.
The ChatGPT group produced stronger-looking answers on the study’s main rubric. The reasoning exercise produced a broader range of ideas. The combined group showed both patterns. That makes this a more interesting finding than a simple claim that AI helps or harms education.
What the study measured
The researchers used trained human graders and a five-point rubric that assessed how well recommendations addressed two marketing goals: increasing awareness and use of the store. They also used text analysis to examine the number and variety of ideas, signs of causal reasoning and similarity to three expert submissions.
According to the researchers and OpenAI, students with access to ChatGPT scored almost a full point higher on that five-point rubric. Their submissions contained more ideas, followed clearer logic and were more similar to recommendations written by experts. Students still had to decide what to ask, assess the responses and choose what went into their final work.
The causal-reasoning exercise taught a specific kind of critical thinking: linking cause and effect, and asking why an idea may work or fail. It did not raise the main rubric score. But it did lead to more varied ideas that were less conventional when compared with peers’ submissions.
Polish and originality are not the same measure
A neat answer can be valuable. It can also be easier to recognise and reward than a novel one. This trial shows how a single grade can miss that difference. The rubric captured stronger structure and more expert-like recommendations. It did not capture the wider idea variety associated with the reasoning exercise.
That does not make the rubric wrong. It was designed for a defined marketing task. It does mean that schools and employers need to be clear about the capability they want to observe. If an AI tool makes a final response more coherent, the final response on its own may say less about how a person explored the problem.
For a classroom, that could mean keeping some process evidence, asking students to explain rejected options or grading the assumptions behind a recommendation. The right approach depends on the course. The study offers a reason to make the question explicit.
The limits are as important as the result
This was a well-defined business assignment at one institution, with first-year students and one particular model. It measured submitted work, not long-term learning, memory, motivation or performance across disciplines. The results should not be treated as a universal answer to how AI changes education.
The study also came from a collaboration between Bocconi and OpenAI Economic Research. Its randomized design is a strength, but the results still need the usual wider discussion, replication and evidence from settings where the task is less suited to a language model.
The practical takeaway is modest. Access to a capable tool and practice in reasoning do not have to compete. If the goal is good work and good judgement, the training around the tool may be at least as important as the tool itself.
What is confirmed, what the researchers say, and what is open
Confirmed: Bocconi University and OpenAI published accounts of the trial on 27 August 2026. Both describe a randomized four-group experiment with 1,053 first-year students, ChatGPT Edu access with GPT-4o and a causal-reasoning exercise.
The researchers’ and OpenAI’s findings: ChatGPT access improved quality and coherence on the five-point rubric, while the causal-reasoning intervention produced broader and more distinctive ideas. The combined group showed gains on the measures attributed to each intervention.
Open questions: whether the pattern holds in other subjects, institutions and age groups; whether it improves durable learning; how results change with different models; and how assessment should distinguish polished output from understanding, originality and responsible use.
Sources
- Bocconi University — Better evaluations with ChatGPT. Critical thinking broadens ideasPrimary university account, published 27 August 2026. Source for the 1,053-student trial, four groups, task, interventions and the study authors’ reported findings.
- OpenAI — Better answers, broader thinking: What students gain from ChatGPT and critical-thinking trainingPrimary research partner account, published 27 August 2026. Source for the rubric, text-analysis measures, reported effects and interpretation of the randomized trial.



