Since the explosion in generative AI, there has been a rash of “decompilations” of video games and various other software (such as this example from earlier today disputed – see below) that have been published to Github and advertised as “Open Source.” That claim is a lie.

The phrase “Open Source” has a specific meaning (and similarly for “Free Software,” by the way1), and it isn’t merely that the source code is there for you to look at. It means that the copyright holder is explicitly giving you permission to read that source code, modify it, redistribute it, etc. Without that element of permission, the code cannot be “Open Source” even if you can physically read it. At best, it might be “Fair Use” depending on the circumstances, but it’s most likely just a fancy means of copyright infringement.

Remember, copyright is a legal construct, not a technical one. It depends much more on the intent of the human doing the copying than it does on the technical details of what they actually did. If the thing they have is obviously intended to be a copy of something else, it is a Derivative Work no matter what technical means were used to create it. That means the original copyright still attaches to it and the person who made the copy doesn’t get to choose a new license for it, “Open Source” or otherwise.

Why YSK:

You don’t have to like the way copyright law works – I sure don’t! – but you do have to understand it because there’s a lot of misinformation going around right now with people claiming things are “Open Source” when they aren’t and a lot of people are going to get in trouble for it. It also dilutes the public understanding of what actual legitimate Open Source software is, which is a problem in and of itself.

Conflating real Open Source software with proprietary software that’s been ‘pirated with extra steps’ is harmful both for developers of the former, who have their reputations damaged by association, and for users/sharers of the latter, who might be misled into not taking the same precautions that they would if they understood that they were dealing with warez. Just because you might think Big Tech can get away with laundering copyright through LLMs – and even that remains to be seen – doesn’t mean the little guys can.

TL;DR: Proprietary software cannot become Open Source software by any means except (a) the express consent of the copyright holder or (b) the copyright expiring and the work becoming Public Domain. Whatever technological end-run you think you have around this legal fact, no you don’t.


EDIT: dispute over example

In giving that example I was relying on the claim in the linked thread, which comes from this guy on BlueSky. Seems like a lot of people think he’s wrong, so maybe that’s not a good example after all.

However, there are also things like this, and those are examples I feel very confident in citing because (a) they explicitly call them “decompliations,” (b) at least one of them has a LICENSE file that says it’s MIT, and (c) there’s zero chance Nintendo or Rare or anyone else legitimately gave them permission for it.


footnote

1 “Free Software” has essentially the same denotation as “Open Source” – close enough that every “Free Software” license is also “Open Source” and vice-versa – but a different connotation. The term “Free Software” tends to get used by people who wish to emphasize the rights of the end-user, while the term “Open Source” tends to get used by people who wish to emphasize that the software is available to be modified.

  • bus_factor@lemmy.world
    link
    fedilink
    English
    arrow-up
    19
    arrow-down
    1
    ·
    22 hours ago

    Whether it qualifies as a clean room implementation is going to spark some debate (and some lawsuits).

    • floofloof@lemmy.ca
      link
      fedilink
      English
      arrow-up
      23
      arrow-down
      1
      ·
      edit-2
      22 hours ago

      You could argue that “vibe coded” and “clean room reimplementation” are by nature incompatible with one another, since you can’t prove the AI didn’t ingest the original code. But you could also question whether the AI could possibly be trained on Adobe’s code when Adobe keeps that under wraps. It could be trained on leaked code, but that would be hard to prove.

      If they used decompilation as a guide for the AI, that might be easier to establish regardless of what the LLM was trained on. There could be telltale implementation details that Adobe could cite.

      In any case, Adobe has a huge number of patents on its software, so I imagine there could be grounds for legal action just because these applications faithfully copy the look and feel and functionality of Adobe products. That might be easier for them to prove in court than anything about the code’s provenance.

      • bus_factor@lemmy.world
        link
        fedilink
        English
        arrow-up
        3
        ·
        22 hours ago

        There’s certainly going to be lawsuits claiming it’s effectively a decompilation as well. We’ll see how it goes.