I’ve tried giving screenshots of phishing emails to a local Qwen instance and so far it always correctly detected it as scam, even points out the exact elements that it based its judgement on. Sending screenshots to it ad-hoc isn’t too scalable for family and friends. I’d like to be able to either forward emails for screening, or perhaps have it screen everything from a mailbox.

Has anyone done anything like this? Is there anything self-hostable that does this?

  • Shadow@lemmy.ca
    link
    fedilink
    English
    arrow-up
    0
    ·
    6 days ago

    I’d be pretty concerned about prompt injection risks with feeding a LLM unsanitized data. You definitely need a good harness around it…

    • Avid Amoeba@lemmy.caOP
      link
      fedilink
      English
      arrow-up
      0
      ·
      6 days ago

      Good point. It’ll have to have no access to the internet or anything local outside of its container. Just text in, text out.

      • Dran@lemmy.world
        link
        fedilink
        English
        arrow-up
        0
        ·
        5 days ago

        You could (and probably should) use a system-one style inference system for spam classification. Much cheaper and the structured output means it’s impossible to go rogue and curl some malware or whatever. It can absolutely misclassify but its output is programmatically structured and just ranks a pre-selected set of output tokens.

        In your case that’s

        Spam

        Not_spam

        • tigerhawkvok@startrek.website
          link
          fedilink
          English
          arrow-up
          0
          ·
          1 day ago

          I only read about that today, so I hadn’t even considered that angle. They’re doing some cool stuff in that space. jeff is the self-hostable one that does best as far as I know.

          • Dran@lemmy.world
            link
            fedilink
            English
            arrow-up
            2
            ·
            10 hours ago

            You actually have a lot of options in this space. You can take the prefill engine of just about any LLM and turn it’s transformer into a classifier by lobotomizing out the decoder. You can also use a diffusion model on single pass to surprisingly competent result.

            If Jeff looks easy enough to deploy by all means start there, but don’t discount the idea if Jeff sucks; you have a lot of options.