In the U.S., copyright protections don’t extend to functionality, though. So copying even copyrighted code in a manner to replicate functionality is fair use when there’s no other way to accomplish the same thing.
For example, if you have a piece of equipment that checks the software on a cartridge for a string of text that says “Produced by or under license from Sega Enterprises Ltd.” before running that software, then it’s fair use to copy that exact text so that your software can run on that equipment. And it’s fair use to reverse engineer and decompile licensed cartridges to see what the bare minimum necessary to make it work.
One way to prove that you didn’t copy the software any more than is strictly necessary for functionality is to fully document the functionality, and then have a skilled programmer take the documentation and write new software from scratch, without ever having seen the original software whose function is being copied. That’s a “cleanroom implementation.” Compaq and other IBM clones built their own BIOS software to implement the exact same functionality of copyrighted IBM code, and created an entire industry of IBM compatible PCs that weren’t actually licensed from IBM. Similarly, Google moved Android off of Sun-licensed Java using a cleanroom implementation (when Oracle bought Sun and Google wanted to get away from Larry Ellison’s abusive pricing practices).
Ok, so if it’s permissible to reverse engineer the code to create documentation of how it works, and then have someone else take that documentation and implement the functionality using new code, how do AI/LLMs fit into this? Can it be said that it’s truly a “cleanroom” when the reverse engineering and decompilation functions are done by the same software that converts the decompiled code into documentation in human language, and then is the same software that converts the documentation into newly implemented code? Doesn’t quite hit the same way, and I’m not sure the courts would see it the same way.
All of this is a gray area, and people shouldn’t confidently predict what the courts will decide in specific nuanced examples. A lot will depend on the specific details, so there isn’t going to be much room for sweeping generalizations.
In the U.S., copyright protections don’t extend to functionality, though. So copying even copyrighted code in a manner to replicate functionality is fair use when there’s no other way to accomplish the same thing.
For example, if you have a piece of equipment that checks the software on a cartridge for a string of text that says “Produced by or under license from Sega Enterprises Ltd.” before running that software, then it’s fair use to copy that exact text so that your software can run on that equipment. And it’s fair use to reverse engineer and decompile licensed cartridges to see what the bare minimum necessary to make it work.
One way to prove that you didn’t copy the software any more than is strictly necessary for functionality is to fully document the functionality, and then have a skilled programmer take the documentation and write new software from scratch, without ever having seen the original software whose function is being copied. That’s a “cleanroom implementation.” Compaq and other IBM clones built their own BIOS software to implement the exact same functionality of copyrighted IBM code, and created an entire industry of IBM compatible PCs that weren’t actually licensed from IBM. Similarly, Google moved Android off of Sun-licensed Java using a cleanroom implementation (when Oracle bought Sun and Google wanted to get away from Larry Ellison’s abusive pricing practices).
Ok, so if it’s permissible to reverse engineer the code to create documentation of how it works, and then have someone else take that documentation and implement the functionality using new code, how do AI/LLMs fit into this? Can it be said that it’s truly a “cleanroom” when the reverse engineering and decompilation functions are done by the same software that converts the decompiled code into documentation in human language, and then is the same software that converts the documentation into newly implemented code? Doesn’t quite hit the same way, and I’m not sure the courts would see it the same way.
All of this is a gray area, and people shouldn’t confidently predict what the courts will decide in specific nuanced examples. A lot will depend on the specific details, so there isn’t going to be much room for sweeping generalizations.