We’re seeing a lot of students, teachers, and scientists produce more, and worse, outputs because of improper AI use. It’s possible to use it better: Have it imitate a critic. Learning and creativity come from getting criticism and experiencing friction.

We found that when knowledge workers intentionally use AI to challenge their ideas – to generate friction – they can significantly improve their performance. Previous research on human-AI collaboration for loan evaluations had similar results.

A study I was not involved with found that creative writers produce better copy when they use AI as a sounding board rather than to ghostwrite.

    • Artisian@lemmy.worldOP
      link
      fedilink
      English
      arrow-up
      2
      arrow-down
      3
      ·
      3 days ago

      Sometimes knowledge work is about learning. That’s also a nice time to skip AI. But I can see from the data that student’s aren’t doing that, so we harm-reduce.

      • naught101@lemmy.world
        link
        fedilink
        English
        arrow-up
        3
        arrow-down
        1
        ·
        2 days ago

        Using it for learning stuff that’s already well known (e.g. beginner level coding) is fine.

        Using it for “knowledge work” is a bad idea idea, because AI pushes toward the mean, while research by definition is focused on new edges of knowledge. This is is exactly the space where AI will introduce biases and untruths that will be very hard to spot. Using it also reduces critical thinking ability which is critical in this space.

        • Artisian@lemmy.worldOP
          link
          fedilink
          English
          arrow-up
          0
          arrow-down
          2
          ·
          2 days ago

          If you’re doing knowledge work, you are an expert. You’re very likely to catch biases and untruths (and those you don’t catch, humanity couldn’t catch. These concerns are just as applicable to a human by themselves.)

          • naught101@lemmy.world
            link
            fedilink
            English
            arrow-up
            3
            ·
            2 days ago

            Sorry, but that’s just not true. Researchers might be experts in teir field generally, but day-to-day work is very rarely focused on well-known parts of their domain, because the whole point is to obtain new knowledge.

            Source: I’m a climate scientist, have been in and around academia for nearly 20 years.

            • Artisian@lemmy.worldOP
              link
              fedilink
              English
              arrow-up
              0
              arrow-down
              2
              ·
              2 days ago

              Could you flesh this out for me, I’m not sure I understand? I think you’re saying:

              1. experts are not well posed to catch biases and untruths generated by genAI in their (research work)
              2. because as an academic climate scientist, your day-to-day work is spent on new climate science (and not established climate science).

              I don’t see why (2) implies (1), but I agree with (2). As I intended it, (1) is my main claim. I follow it up with

              1. if an expert cannot spot a bias or flaw in genAI output, then they wouldn’t catch it from a peer either.

              I don’t see how (2) helps with (3) either.

              • naught101@lemmy.world
                link
                fedilink
                English
                arrow-up
                4
                ·
                22 hours ago

                GenAI is an averager (as all empirical models are), so it tends towards predicting the mean of whatever it’s trained on (in a given context). But it also has noise added, so some variance comes back, but there’s no guaranteed that that mean+noise produces something meaningful/true/valuable.

                But, GenAI is very good at producing syntactically correct language. This is a problem because it lulls the reader into a sense that the author knows what it’s talking about, when it doesn’t. When a junior scientist produces text, the awkward language alerts the reviewer to the poor thinking (same with a junior coder producing weak code). With an LLM, you don’t get that - it produces the impression of knowledge without any actual understanding.

                This, combined with the lure of efficiency and the feeling of effectiveness, make it a honey pot for quick but sloppy thinking. I think this is what connects 2 to 1, especially in domains where it’s hard to get external validation from other people who can understand what you’re trying to mean, and not just say.

                As for your point 3: I agree. People don’t spot flaws already, and that’s part of why we saw the replication crisis in behavioural science. Adding LLMs to the mix will just make things like that more likely.

                • Artisian@lemmy.worldOP
                  link
                  fedilink
                  English
                  arrow-up
                  1
                  ·
                  9 hours ago

                  (There is a separate claim about genAI producing the average/typical behavior from its inputs; this isn’t true for RLHF and boosting reasons, have a gander at boosting and wisdom of the crowds for how you can take a cheap source of mediocrity and create something much better. It is true that genAI outputs feel unoriginal and omnipresent, as so very many folks are throwing around genAI outputs everywhere. But I think that’s not a property of the statistics, rather of the economics/behavioral science.)

                  • naught101@lemmy.world
                    link
                    fedilink
                    English
                    arrow-up
                    2
                    ·
                    6 hours ago

                    RLHF and boosting

                    I don’t really see how those avoid the problem. I can see how they produce much better training results, but ultimately a trained LLM is a neural net that has fixed set of parameters that represent a function over n-dimensional space (I.e. the information in the context window), plus a noise term. If you turn the temperature down to zero, they are perfectly deterministic, no? That’s the mean that they are producing values around…

                • Artisian@lemmy.worldOP
                  link
                  fedilink
                  English
                  arrow-up
                  2
                  arrow-down
                  1
                  ·
                  9 hours ago

                  So I think you’ve landed on the problem that I have with student use, and that the article has with scientist default use. It is very bad for folks to use AI to try to do the knowledge work directly, yet people are making this mistake constantly. YSK: it’s much better when you make AI increase the friction of knowledge work. (note neither are claiming that this use is particularly good; we’re doing damage control.)

                  I definitely notice sloppy thinking when I see it in someone else’s review of my papers. You certainly pick it up when you read climate-change-denialists. Experts are great at noticing sloppy thinking when it disagrees with them. That is the relationship you want with genAI if you are using it for knowledge work. Make it disagree with you, so you notice where it is very sloppy (and sometimes where you’ve been sloppy because the bad thinker is kinda right).

                  • naught101@lemmy.world
                    link
                    fedilink
                    English
                    arrow-up
                    2
                    ·
                    6 hours ago

                    Yeah. Agree in theory. I’m a bit sus on the adversarial AI approach in practice, because most LLMs seem to produce a lot of sycophancy, and I think that’s a far too easy trap to fall in to.

                    I do think there are some practical uses, such as reducing availability bias by asking it “is there anything I’ve forgotten to consider in this thing I wrote?” AFTER you’ve done the thinking and writing. But I think the temptation to use it for generative purposes is pretty huge, and likely to suck a lot of people in.