We’re seeing a lot of students, teachers, and scientists produce more, and worse, outputs because of improper AI use. It’s possible to use it better: Have it imitate a critic. Learning and creativity come from getting criticism and experiencing friction.
We found that when knowledge workers intentionally use AI to challenge their ideas – to generate friction – they can significantly improve their performance. Previous research on human-AI collaboration for loan evaluations had similar results.
A study I was not involved with found that creative writers produce better copy when they use AI as a sounding board rather than to ghostwrite.



If you’re doing knowledge work, you are an expert. You’re very likely to catch biases and untruths (and those you don’t catch, humanity couldn’t catch. These concerns are just as applicable to a human by themselves.)
Sorry, but that’s just not true. Researchers might be experts in teir field generally, but day-to-day work is very rarely focused on well-known parts of their domain, because the whole point is to obtain new knowledge.
Source: I’m a climate scientist, have been in and around academia for nearly 20 years.
Could you flesh this out for me, I’m not sure I understand? I think you’re saying:
I don’t see why (2) implies (1), but I agree with (2). As I intended it, (1) is my main claim. I follow it up with
I don’t see how (2) helps with (3) either.
GenAI is an averager (as all empirical models are), so it tends towards predicting the mean of whatever it’s trained on (in a given context). But it also has noise added, so some variance comes back, but there’s no guaranteed that that mean+noise produces something meaningful/true/valuable.
But, GenAI is very good at producing syntactically correct language. This is a problem because it lulls the reader into a sense that the author knows what it’s talking about, when it doesn’t. When a junior scientist produces text, the awkward language alerts the reviewer to the poor thinking (same with a junior coder producing weak code). With an LLM, you don’t get that - it produces the impression of knowledge without any actual understanding.
This, combined with the lure of efficiency and the feeling of effectiveness, make it a honey pot for quick but sloppy thinking. I think this is what connects 2 to 1, especially in domains where it’s hard to get external validation from other people who can understand what you’re trying to mean, and not just say.
As for your point 3: I agree. People don’t spot flaws already, and that’s part of why we saw the replication crisis in behavioural science. Adding LLMs to the mix will just make things like that more likely.