We’re seeing a lot of students, teachers, and scientists produce more, and worse, outputs because of improper AI use. It’s possible to use it better: Have it imitate a critic. Learning and creativity come from getting criticism and experiencing friction.
We found that when knowledge workers intentionally use AI to challenge their ideas – to generate friction – they can significantly improve their performance. Previous research on human-AI collaboration for loan evaluations had similar results.
A study I was not involved with found that creative writers produce better copy when they use AI as a sounding board rather than to ghostwrite.



GenAI is an averager (as all empirical models are), so it tends towards predicting the mean of whatever it’s trained on (in a given context). But it also has noise added, so some variance comes back, but there’s no guaranteed that that mean+noise produces something meaningful/true/valuable.
But, GenAI is very good at producing syntactically correct language. This is a problem because it lulls the reader into a sense that the author knows what it’s talking about, when it doesn’t. When a junior scientist produces text, the awkward language alerts the reviewer to the poor thinking (same with a junior coder producing weak code). With an LLM, you don’t get that - it produces the impression of knowledge without any actual understanding.
This, combined with the lure of efficiency and the feeling of effectiveness, make it a honey pot for quick but sloppy thinking. I think this is what connects 2 to 1, especially in domains where it’s hard to get external validation from other people who can understand what you’re trying to mean, and not just say.
As for your point 3: I agree. People don’t spot flaws already, and that’s part of why we saw the replication crisis in behavioural science. Adding LLMs to the mix will just make things like that more likely.