

(There is a separate claim about genAI producing the average/typical behavior from its inputs; this isn’t true for RLHF and boosting reasons, have a gander at boosting and wisdom of the crowds for how you can take a cheap source of mediocrity and create something much better. It is true that genAI outputs feel unoriginal and omnipresent, as so very many folks are throwing around genAI outputs everywhere. But I think that’s not a property of the statistics, rather of the economics/behavioral science.)

Perhaps this sentence is not what you wanted then? The mean of whatever it’s trained on (that is, the training data) need not be the mean of the trained LLM function. This is because we use complicated loss functions to update our LLMs, stochastic processes in training, and that we change those functions after training on input data by using RLHF and other models (eg, via using a mixture of itself to get better outputs with the boosting algorithm).
You are right that once we fix an LLM’s parameters, at temp 0 they are deterministic, though they are recursive functions so it’s weird to talk about a long string of outputs at temp 0 as the ‘mean’ of the LLM; high temp outputs won’t ‘regress’ to this temp 0 output over time, for example. There also may be hallucinations or mistakes that are common at temp 0, but extremely rare at all other temperatures (because these mistakes only get made when it looks at its own temp 0 context). So I’m not sure the deterministic output string(s) are particularly useful.
(An analogy that I suspect isn’t very good: the butterfly effect, adding a single small air perturbation can lead to drastically different deterministic weather simulation results. This is very much the norm for LLM outputs)
(They are somewhat more than a neural network; the ‘attention’ layers add something pretty strange. And output temperature isn’t implemented as a noise term, but this is a fine way to think of it.)