We’re seeing a lot of students, teachers, and scientists produce more, and worse, outputs because of improper AI use. It’s possible to use it better: Have it imitate a critic. Learning and creativity come from getting criticism and experiencing friction.
We found that when knowledge workers intentionally use AI to challenge their ideas – to generate friction – they can significantly improve their performance. Previous research on human-AI collaboration for loan evaluations had similar results.
A study I was not involved with found that creative writers produce better copy when they use AI as a sounding board rather than to ghostwrite.



(There is a separate claim about genAI producing the average/typical behavior from its inputs; this isn’t true for RLHF and boosting reasons, have a gander at boosting and wisdom of the crowds for how you can take a cheap source of mediocrity and create something much better. It is true that genAI outputs feel unoriginal and omnipresent, as so very many folks are throwing around genAI outputs everywhere. But I think that’s not a property of the statistics, rather of the economics/behavioral science.)
I don’t really see how those avoid the problem. I can see how they produce much better training results, but ultimately a trained LLM is a neural net that has fixed set of parameters that represent a function over n-dimensional space (I.e. the information in the context window), plus a noise term. If you turn the temperature down to zero, they are perfectly deterministic, no? That’s the mean that they are producing values around…
Perhaps this sentence is not what you wanted then? The mean of whatever it’s trained on (that is, the training data) need not be the mean of the trained LLM function. This is because we use complicated loss functions to update our LLMs, stochastic processes in training, and that we change those functions after training on input data by using RLHF and other models (eg, via using a mixture of itself to get better outputs with the boosting algorithm).
You are right that once we fix an LLM’s parameters, at temp 0 they are deterministic, though they are recursive functions so it’s weird to talk about a long string of outputs at temp 0 as the ‘mean’ of the LLM; high temp outputs won’t ‘regress’ to this temp 0 output over time, for example. There also may be hallucinations or mistakes that are common at temp 0, but extremely rare at all other temperatures (because these mistakes only get made when it looks at its own temp 0 context). So I’m not sure the deterministic output string(s) are particularly useful.
(An analogy that I suspect isn’t very good: the butterfly effect, adding a single small air perturbation can lead to drastically different deterministic weather simulation results. This is very much the norm for LLM outputs)
(They are somewhat more than a neural network; the ‘attention’ layers add something pretty strange. And output temperature isn’t implemented as a noise term, but this is a fine way to think of it.)