We’re seeing a lot of students, teachers, and scientists produce more, and worse, outputs because of improper AI use. It’s possible to use it better: Have it imitate a critic. Learning and creativity come from getting criticism and experiencing friction.
We found that when knowledge workers intentionally use AI to challenge their ideas – to generate friction – they can significantly improve their performance. Previous research on human-AI collaboration for loan evaluations had similar results.
A study I was not involved with found that creative writers produce better copy when they use AI as a sounding board rather than to ghostwrite.
Or don’t use AI for any work that matters.
Don’t use it for conclusions at all.
Use it to expand your vocabulary on the topic and as a search engine supplement to find sources
Or don’t use AI for knowledge work when being correct is important.
Sometimes knowledge work is about learning. That’s also a nice time to skip AI. But I can see from the data that student’s aren’t doing that, so we harm-reduce.
Using it for learning stuff that’s already well known (e.g. beginner level coding) is fine.
Using it for “knowledge work” is a bad idea idea, because AI pushes toward the mean, while research by definition is focused on new edges of knowledge. This is is exactly the space where AI will introduce biases and untruths that will be very hard to spot. Using it also reduces critical thinking ability which is critical in this space.
If you’re doing knowledge work, you are an expert. You’re very likely to catch biases and untruths (and those you don’t catch, humanity couldn’t catch. These concerns are just as applicable to a human by themselves.)
Sorry, but that’s just not true. Researchers might be experts in teir field generally, but day-to-day work is very rarely focused on well-known parts of their domain, because the whole point is to obtain new knowledge.
Source: I’m a climate scientist, have been in and around academia for nearly 20 years.
Could you flesh this out for me, I’m not sure I understand? I think you’re saying:
- experts are not well posed to catch biases and untruths generated by genAI in their (research work)
- because as an academic climate scientist, your day-to-day work is spent on new climate science (and not established climate science).
I don’t see why (2) implies (1), but I agree with (2). As I intended it, (1) is my main claim. I follow it up with
- if an expert cannot spot a bias or flaw in genAI output, then they wouldn’t catch it from a peer either.
I don’t see how (2) helps with (3) either.
GenAI is an averager (as all empirical models are), so it tends towards predicting the mean of whatever it’s trained on (in a given context). But it also has noise added, so some variance comes back, but there’s no guaranteed that that mean+noise produces something meaningful/true/valuable.
But, GenAI is very good at producing syntactically correct language. This is a problem because it lulls the reader into a sense that the author knows what it’s talking about, when it doesn’t. When a junior scientist produces text, the awkward language alerts the reviewer to the poor thinking (same with a junior coder producing weak code). With an LLM, you don’t get that - it produces the impression of knowledge without any actual understanding.
This, combined with the lure of efficiency and the feeling of effectiveness, make it a honey pot for quick but sloppy thinking. I think this is what connects 2 to 1, especially in domains where it’s hard to get external validation from other people who can understand what you’re trying to mean, and not just say.
As for your point 3: I agree. People don’t spot flaws already, and that’s part of why we saw the replication crisis in behavioural science. Adding LLMs to the mix will just make things like that more likely.
(There is a separate claim about genAI producing the average/typical behavior from its inputs; this isn’t true for RLHF and boosting reasons, have a gander at boosting and wisdom of the crowds for how you can take a cheap source of mediocrity and create something much better. It is true that genAI outputs feel unoriginal and omnipresent, as so very many folks are throwing around genAI outputs everywhere. But I think that’s not a property of the statistics, rather of the economics/behavioral science.)
So I think you’ve landed on the problem that I have with student use, and that the article has with scientist default use. It is very bad for folks to use AI to try to do the knowledge work directly, yet people are making this mistake constantly. YSK: it’s much better when you make AI increase the friction of knowledge work. (note neither are claiming that this use is particularly good; we’re doing damage control.)
I definitely notice sloppy thinking when I see it in someone else’s review of my papers. You certainly pick it up when you read climate-change-denialists. Experts are great at noticing sloppy thinking when it disagrees with them. That is the relationship you want with genAI if you are using it for knowledge work. Make it disagree with you, so you notice where it is very sloppy (and sometimes where you’ve been sloppy because the bad thinker is kinda right).
Or better yet, don’t! Think for yourself while you still can.
yeah but you’re very biased about what you think
it actually is really refreshing to get a LLM to break down your arguments and ideas
cognitive friction really gets the best out of your brain even if the AI is only averagely good at disproving your ideas
thinking for yourself and using AI isn’t mutually exclusive
This is what peer review, teacher feedback, and group workshops are for. In college I didn’t need an llm to critique my work. That was the job of my classmates, and mine for their writing. And for those outside of college, you can still just seek out writing circles, either irl or online. It takes some work, but you’ll get better feedback from real humans who can actually feel the emotion in your writing. You can also seek out similar communities for other hobbies. The only reason for someone to pick an ai critique over a human one is convenience and the cost is way too high to justify that
Idk about you, but my mentors and peers are all busy stressing over the polycrisis.
I’m sure the subjects of the research mentioned in the article do all that.
I can say AI is being used for a lot of stupid wasteful things. I don’t know what it costs for AI to seemingly improve our knowledge workers, but to me that is worth something.
Knowledge work is generally pretty expensive; it’s done in dense places, uses a lot of air travel, has pretty big long-term consequences. Even the high-end estimates on AI costs are probably worth modest gains for knowledge workers.
mmm yes, using the thing famous for being fundamentally programmed to be a sycophant to challenge your ideas…
Did you have the AI challenge that idea?
don’t antromorphize the computer sycopathy is just how they programmed the thing all the information to heavily challenge your ideas is in there
try it
tell it you think using an AI to challenge your ideas is great (the opposite of what you think)
AI is going to tank academia. It’s already going down hill after decades of corporate-style management.
Someone else I know said this in a conference and I disagree. You should instruct it to not do any extraneous language and give back data in the most impartial way possible. You should use it for things you know well so that you can see how it fails and double check any results you are going to use. Having it purpsefully be adversarial is not going to improve what you get out of it and will just waste time.
Thoughts on the study from the article? link
oh I won’t go to links off the fediverse.
I’ve done this. Hilariously frustrating when it starts doubling down on hallucinations.
As I’ve been messing around with it, I am becoming a fan of ‘hallucination -> new context’.
prompt like this:
"what did Nietzsche think was the influence on western society of the victory of jewish christian moral on the ancient greek values
and why was he so stupid about it"
I’m not walking around touting it, cause this is a time of battling cults. But I have been using an AI as a tutor to teach me how to write C code, with an eye towards kernel contributions and other things that require that speed. Obviously it’s a more marketable dev skill, even in this time, than the Python I know. It is very good at teaching, actually. Too bad about us all perishing by asphyxiation.
I wouldn’t trust a professor who lies to me as often as LLMs do.
it’s not even lies it’s just not being able to tell what is true because you’re not actually conscious
(Why is consciousness required to tell true/false statements? I think I’d even say ‘that sign is dishonest/wrong/a liar’, regardless of the fact that the sign is very much dead.)
Programming is a very non-fuzzy topic; it’s no surprise that as the great delusion shakes out, generated art is too despised to be a significant revenue stream, and so far as I can tell the generated human text is still amateurishly bad.
But it’s a fact that these models have been trained on just about every bit of publicly-available code, which by and large, is good, working code. Not every bit of it, but enough that I have yet to see the thing spit out bad syntax… though it catches my bad syntax every time.
I have not and will never try to use it to write code for me, but it’s a smashing tutor, and I could certainly see an experienced dev who knows exactly what they want, handing off the building of tedious functions and stuff, piece by piece, and getting back good, compact code. I believe that the problems we’re having are to do with someone handing off the building of an entire website, or an entire app, that’s where you end up with runaway dependencies, bloated and badly-written functions, etc.
Now, it still goes idiot all the time. I point out that my push function is failing because the ringbuffer_is_full check is returning true every time (I’m very dumb when I’m first starting out on anything), it “sees” the problem and tells me correctly what to do to fix it, but it talks about the function returning false every time. Correct answer, but the dressing on the salad is wrong. This only happens, though, after I’ve been in the same chat for a while (context window overflow).
In other words, it’s a tool that humans can use, while keeping an eye for that kinda thing, but it certainly will not replace anyone, though we’re going to absorb a lot of damage from stupid rich people trying.
In other news, I had a dream last night where I was hanging out with an old ex, and when she showed up, she (somehow) told me that not only had she gone blind since we last talked, she had then gotten a brain tumour and had her head amputated. So I was walking around with my headless ex all day in this dream. I don’t dislike this ex, we were just walking around (how?) and talking to people (HOW?).
I wonder if this is a product of me conversing with something that has no consciousness so much this last week. Or maybe it’s the Ed Gein series we’re watching, I dunno.
I’ve actually been able to shape GPT this way via a bunch of global settings I’ve invoked over the years. Don’t get me wrong-- it’s a PITA, and pretty much a never-ending process, but I think I have two dozen settings now it follows as best it can in order to provide something like a ‘pros and con’ list for any opinion it spits out. This cuts way down on its classic tendency to hallucinate bullshit and spit out half-baked / incorrect replies.
(ready for the downvotes now assuming that I love AI as a general concept 😅)







