Would be a terrible shame if lots of people opted out.

If you are a EU citizen you might also want to write a complaint to [email protected] because they are collecting your personally identifiable information in machine-readable form which they are distributing to third parties.

  • katy ✨@piefed.blahaj.zone
    link
    fedilink
    English
    arrow-up
    40
    arrow-down
    1
    ·
    21 days ago

    We want to give developers agency over their source code by letting them decide whether or not it should be used to develop and evaluate machine learning models.

    fuck them seriously; if you want to do that then don’t steal the repositories in the first place.

    • kibiz0r@midwest.social
      link
      fedilink
      English
      arrow-up
      30
      ·
      21 days ago

      Meanwhile at work we just had a training course that specifically said doing “opt out” instead of “opt in” violates the principle of informed consent.

      • atomicbocks@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        9
        ·
        20 days ago

        Notice how a shit load of these people keep turning up to be rapists and pedophiles… They have no shits to give about informed consent. In their mind you don’t even have the right to the same agency they do.

    • lps2@lemmy.ml
      link
      fedilink
      arrow-up
      12
      ·
      21 days ago

      Kinda hope it uses my code. It’s so terrible there’s no doubt it will make the resulting code from the model worse even if the impact is miniscule

      • kboy101222@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        5
        arrow-down
        1
        ·
        20 days ago

        I just opted out of all my repos except the God awful ones from middle and highschool. Those are basically prompt poison so fuckem

    • JakenVeina@midwest.social
      link
      fedilink
      arrow-up
      3
      ·
      21 days ago

      Unless I’m mistaken, this wasn’t written by the folks that scaped GitHub in the first place, someone just wrote a small tool to semi-automate the process of searching the scraped dats, and submitting a GitHub issue to have it removed.

      • PlexSheep@infosec.pub
        link
        fedilink
        arrow-up
        6
        ·
        20 days ago

        It’s not, but it may be violation of licenses. And also, if it has personal information on it, that’s probably illegal under the GDPR.

        • FizzyOrange@programming.dev
          link
          fedilink
          arrow-up
          3
          arrow-down
          3
          ·
          20 days ago

          but it may be violation of licenses

          They excluded code with non-permissive licenses apparently:

          Each file is labelled permissive (at least one permissive license detected, no conflicting non-permissive license), no_license (no licenses detected, or only non-license legal texts such as CLAs), or non_permissive. The permissive allowlist follows the Blue Oak Council list plus licenses categorized as Permissive or Public Domain by ScanCode. Files classified as non_permissive are excluded from both released datasets.

          also, if it has personal information on it, that’s probably illegal under the GDPR.

          It’s all public so I would be extremely surprised if that were the case.

    • fruitcantfly@programming.dev
      link
      fedilink
      arrow-up
      7
      ·
      edit-2
      20 days ago

      Other git hosts are also getting scraped, and have had to implement counters because of it. For example, this is the kind of thing Codeberg shows crawlers. I’ve even seen people who self-host complaining about getting overloaded because of bots scraping their forge

      • PlexSheep@infosec.pub
        link
        fedilink
        arrow-up
        1
        ·
        20 days ago

        I’ve put Anubis before most of my website, including my forgejo instance. For the projects hosted there, which is not all, I can only hope that that’s enough.

        I like to have the visibility and CI of GitHub. But this sucks ass.

    • Eager Eagle@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      edit-2
      21 days ago

      note that former users would have needed to remove their GitHub data before August 2025 to not be in this dataset

  • gedfromgont@piefed.ca
    link
    fedilink
    English
    arrow-up
    17
    ·
    20 days ago

    I never really needed any justification, but ever since AI companies just take stuff illegally, and that openly, it has become my justification to just pirate the shit out of everything.

    (Excluding indie games of developers I like).

    • Melvin_Ferd@lemmy.world
      link
      fedilink
      arrow-up
      1
      arrow-down
      6
      ·
      edit-2
      20 days ago

      It should have been anyways. Like what the fuck were you all doing on the internet. Trying to sell your stupid art. Helping corporations steal data and build pay walls so you can earn $500/month selling your kitchy Garfield key chains. The internet should have always been about data replication and sharing. It’s too late to put that cat back in the bag. It was suppose to be stopped 20 years ago. Now we have all this stuff ruining the world and bringing back nazis. But at least the lefty capitalist artist sold some key chains

        • Melvin_Ferd@lemmy.world
          link
          fedilink
          arrow-up
          2
          arrow-down
          1
          ·
          edit-2
          20 days ago

          Keychain just a stand in for anything. More like how someone making videos of making key chains used patent trolls to build tools to scan all the videos they can in order to find users they can take down. Just one example. The little lefty entrepreneur was used to beat the old school leftists into the dirty on behalf of the corporate interests. Now everything is owned by corporate and we lost the internet to racist nazis. There’s a moment 20 years ago things could have been different that we’ll never get back. Frontier was new and we needed to rally around the wagon to defend the frontier. Instead we sold out.

          And now all the lefties are in the smallest corner of the internet and the rest of the internet and social media is owned by the capitalist and nazis. It all started by forgetting the roots and in favor of selling our key chains and trying to build a market rather than a collaborative space.

          • Ricaz@lemmy.dbzer0.com
            link
            fedilink
            English
            arrow-up
            2
            ·
            20 days ago

            What’s with the obsession with “lefties” vs nazis?? Sounds like somebody stayed on the propaganda carousel a bit too long.

            You choose what parts of the internet you want to use, so who cares if 99% is profit oriented garbage, honestly.

            People choose to publicly share their stuff for free, so of course it’s gonna be taken advantage of. That’s literally capitalism.

            • Melvin_Ferd@lemmy.world
              link
              fedilink
              arrow-up
              1
              arrow-down
              2
              ·
              edit-2
              20 days ago

              Yea it’s so crazy. Lol

              Everything is fine. Nazis are marching in streets again primarily driven by online recruitment. Billionaires are investing in removing leftist from digital spaces to facilitate this. But because I care, I’m obsessed.

              I’ll just choose to hang out here with the people who say they oppose this stuff but really, they’re just what, posting about beans and moths while nazis take over the government, March in the street and feed children their content through multiple pipelines that are growing because there’s fuck all resistance. After all the “resistance” is busy trying to decide which hole they want to hide in.

      • vanillama@programming.dev
        link
        fedilink
        arrow-up
        3
        arrow-down
        1
        ·
        20 days ago

        But at least the lefty capitalist artist sold some key chains

        Selling your shit and getting paid doesn’t make you into a capitalist. At most you’re an artisan, if you own your tools and you get paid for your final product rather than your labour hours.

        To be a capitalist you necessarily need to own capital. In the example you picked you’d need to own a factory where you underpay people to and keep the surplus from all the keychains they produce.

        And regardless of what you meant, you’re absolutely barking at the wrong tree here, market consolidation and monopolies predate the internet, and they would have always had the final say in how we shape our infrastructure and share shit online under capitalism.

        • Melvin_Ferd@lemmy.world
          link
          fedilink
          arrow-up
          1
          arrow-down
          1
          ·
          20 days ago

          You would be a supporter of capitalist. The capitalist needed to take the internet away from the left. It was hostile to making capital. Early days people were much more aware of defending it against capitalist which like you said, people were tracking how capitalist ruined all other frontiers and left them dried out husks that stifled originality, creativity and actual innovation. Look at the leftovers from the early days. Websites from them were focused on free use and sharing information freely, gnu, Winamp, VLC, sites like reddit, digg, Wikipedia wouldn’t exist with today’s internet if they were created today. Capitalist took over.

        • Melvin_Ferd@lemmy.world
          link
          fedilink
          arrow-up
          2
          arrow-down
          1
          ·
          edit-2
          20 days ago

          Which goes to my point. You fought wrong. What i see is that the left is very likely manipulated by right wing interest groups way more than people want to admit. I think for decades now, they have targeted leftist online for issues we traditionally dominated and redirected people into areas that were very ineffective. Meanwhile they socially engineered the right side and MAGA to actually effective social engagement. After decades the left is just a dead in the water idea with people who no longer know how to use the internet correctly for social engineering and engagement. The only time the left rallies is when the effort is the most high effort, high energy, low reward movement people could possibly imagine.

          The left in the past were intelligent intellectual building and innovating things. That is no longer the case. We’re resting on the laurels of the best and looking around at lemmy should make it very obvious how true this is.

          The sad reality i think people need to understand its the MAGA isn’t winning just because of money. They’re wining because they’re putting effort in areas that the left don’t even realize are important. They are playing chess and the left aren’t even playing because competition makes the left feel anxious.

          I have multiple people in my life on the right who spend all their time online posting and networking and planning. They will shit post all day. Between the 3 of them, they probably create tons of content. Let alone sharing and promoting others content.

          People on the left I know will read an article. Can’t be bothered to share or even upvote. Maybe they’ll write a comment like “This sucks, Trump is stupid” and then they go look for star trek memes all day. Give this behavior 20 years for the kids today to grow up. The global cultural image of the left right now is a weak, gutless screaming loser who nobody wants to admit to being outside of small circles. Kids today hate the left. It’s going to get so much worse.

  • SpaceCowboy@lemmy.ca
    link
    fedilink
    arrow-up
    14
    ·
    20 days ago

    A decade ago, if someone asked someone working on an Open Source project if they’d be ok with an AI reading their code and learning from it, they’d most likely say “yeah that sounds really cool!”

    Somehow the tech-bros have fucked up AI so much that something that should be really cool seems creepy, lame, and nefarious all at once.

    • FishFace@piefed.social
      link
      fedilink
      English
      arrow-up
      1
      ·
      20 days ago

      Making arguments from popular perception is never very strong. AI is cool if you actually think about it - the capability is incredible. Anything that can produce working code was going to have this ambivalent result where execs pushed it way too hard.

        • NewNewAugustEast@lemmy.zip
          link
          fedilink
          arrow-up
          1
          ·
          edit-2
          20 days ago

          Simultaneously they have created the part I wanted from Star Trek, while making the worst possible anti star trek a reality.

          E.g.

          I want to be able to ask a computer about history, art, codeing, well anything. And be able to clarify and question and put together new ideas.

          But not by anyone owning that ability or profiting on the labor of others or causing environmental harm.

          We got the cool computer but haven’t achieved the post-scarcity part.

          I don’t think you can have one without the other.

          • SpaceCowboy@lemmy.ca
            link
            fedilink
            arrow-up
            0
            arrow-down
            1
            ·
            19 days ago

            Post-scarcity isn’t actually possible. Not even in Star Trek this is true. Picard’s family owns a vineyard in France filled with artifacts and antiques. Not everything is fungible. People will desire these non-fungible things. Not everyone that desires these things will be able to have them, because they aren’t fungible. You can’t have a billion people all owning vineyards in France. Some people won’t get everything they want. There will always be scarcity.

            Star Trek was made in the 1960’s at the height of the cold war. They didn’t want the show to be about how capitalism was superior to communism, or vice versa. So they side stepped the issue by saying in the future there’s no scarcity, no money, and they live under some ideal future economic model. The show writers don’t know what that ideal economic model is, so it’s deliberately vague and inconsistent. And that’s fine because the show isn’t about economics.

            It was wise of the writers of Star Trek to avoid trying to make predictions of that nature. 300 years ago, Wealth of Nations wasn’t yet written, the field of economics didn’t really exist. They believed shiny rocks had intrinsic value. They had some vibes about things like currency devaluation and inflation, but economics was still mostly about acquiring shiney rocks 300 years in the past. So what will economics be like 300 years in the future? None of us know, and certainly writers of a TV show don’t know.

            The writers of a TV show saying there will be no scarcity in the future just means they didn’t want to discuss economics. It’s not any kind of prediction about the future. There will always be scarcity.

            • NeilBrü@programming.dev
              link
              fedilink
              English
              arrow-up
              1
              ·
              edit-2
              17 days ago

              I tepidly agree that there will always be scarcity, but this:

              300 years ago, Wealth of Nations wasn’t yet written, the field of economics didn’t really exist. They believed shiny rocks had intrinsic value. They had some vibes about things like currency devaluation and inflation, but economics was still mostly about acquiring shiney rocks 300 years in the past. Taylor it for Reddit, this is a particularly leftist space where purported or self-styled leftists tend to jam their half-baked theories about Marxism and neo-Marxism in juxtaposition to late-stage capitalism.

              Tl;dr: Your claims about history reveal that you’re full of shit.

              The idea that pre-Adam Smith economics was just “cavemen liking shiny rocks” is a complete caricature. 300 years ago (1720s), thinkers weren’t staring at gold because it was pretty; they were managing global trade empires, state finance, and imperial expansion.

              ​A few actual facts:

              ​"Shiny rocks" was about material power, not shiny aesthetics.

              Mercantilists hoarded silver and gold because bullion was the only globally liquid medium of exchange that could buy naval fleets, pay mercenary armies, and settle trade balances with empires like China that refused European paper money. It was cold geopolitical pragmatism.

              ​They understood inflation long before 300 years ago.

              Copernicus literally wrote a treatise in 1526 explaining how debasing currency drives up prices, laying the foundation for the Quantity Theory of Money. In the late 1500s, Jean Bodin explicitly tracked how the influx of stolen American silver was causing massive inflation across Europe.

              ​The theory of political economy was thriving.

              The School of Salamanca in the 1500s had already analyzed value theory, supply/demand, and the ethics of trade. By the early 1700s, economists like Richard Cantillon were mapping out how money creation trickles through different social classes (the Cantillon Effect).

              ​Adam Smith didn’t invent the field of economics out of thin air in 1776—he was responding to centuries of detailed, material economic critique and state policy.

              Dismissing centuries of economic history as “vibes and shiny rocks” is just lazy.

    • FizzyOrange@programming.dev
      link
      fedilink
      arrow-up
      1
      arrow-down
      3
      ·
      20 days ago

      I mean they’d have been ok with it because tech bros were the ones automating other people out of jobs and never thought it would come for theirs.

      The level of AI we have now was impossible science fiction a decade ago.

    • qaz@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      20 days ago

      They also scraped some of my assembly code of which I’m pretty sure the latest version has a major bug. Let’s hope they don’t use these models to program my new PC’s bios.

  • Prior_Industry@lemmy.world
    link
    fedilink
    English
    arrow-up
    6
    ·
    20 days ago

    When I hear this company’s name I can’t get the idea out of my head that it’s related to the face huggers from Alien.

  • purplemonkeymad@programming.dev
    link
    fedilink
    arrow-up
    6
    arrow-down
    1
    ·
    20 days ago

    Am I alone in not wanting to put my username into that field? If they don’t have it will they then just decide that it’s now a good time to scrape it? Or are they going to record that it was searched?

  • chicken@lemmy.dbzer0.com
    link
    fedilink
    arrow-up
    5
    ·
    21 days ago

    A little bit infuriating since huggingface itself requires login to access a large portion of the content on their site

    • cecilkorik@lemmy.ca
      link
      fedilink
      English
      arrow-up
      2
      ·
      edit-2
      20 days ago

      As long as the datasets are open, it is our best hope. I know it doesn’t compensate the people whose work’s copyright and licenses have been violated, but I think it’s the only realistic hope we’ve got of getting out of this informational dystopia with a reasonably intact library of humanity’s knowledge that hasn’t been locked down and/or monetized. The AI scrapers and generators are in the process of burning down the great library of Alexandria that the Internet had become, and we are already starting to feel its loss. We cannot stop the wave of toxic pollution that is spreading through all our digital content now, but the archives from before this apocalypse started will become the most valuable thing humanity has ever produced. This is information war, and we are losing.

      • hexagonwin@lemmy.today
        link
        fedilink
        arrow-up
        2
        ·
        20 days ago

        IA is a nonprofit and archives to preserve human history. shitty AI startups do this to monetize the data, and their end goal is to “replace” the people who made that data in the first place.

        Thanks to these AI mfs the IA now prevents access to many items because they can be used as training material which fucking sucks

      • trem@lemmy.blahaj.zone
        link
        fedilink
        arrow-up
        1
        ·
        20 days ago

        As a developer, you hold the copyright to your code. When you make it open-source, you grant a license to use the code and the resulting program under certain terms.

        This is a contract. If you copy my code without following these terms, then that’s theft.

        The Internet Archive’s use complies with these terms for all open-source licenses. These AI companies do not. In particular, here’s a quote from the MIT license, which you will find in a similar wording in all open-source licenses:

        The above copyright notice and this permission notice shall be included in all copies or substantial portions of the Software.

        https://mit-license.org/

        In effect, what this means, is that when you copy my code, I demand that you also copy the license text along with it, so that anyone else looking at this code knows the permissions I grant and the terms I require.

        And now guess what these AI companies are doing. They copy my code and reproduce substantial portions upon a user asking, yet they do not include my license terms. They violate the contract under which they obtained my source code.

        I suspect you don’t realize how shit that is, because source code is so abstract.
        It’s like spending hundreds of hours painting a great artwork and then deciding that everyone should be able to give a copy to everyone they know, under the simple condition that they inform those people that they have this right as well.
        And then comes along a company and sells my artwork for money, without informing their customers that they can pass it on for free. That’s, plain and simple, a criminal operation.

          • fruitcantfly@programming.dev
            link
            fedilink
            arrow-up
            1
            ·
            edit-2
            20 days ago

            If nobody owns the code, then then nobody can enforce the terms of the license it was released under, and free software under the FSF definition becomes impossible. All you have is public domain.

            For example, a company could take the Linux kernel, modify it and distribute it with their gadgets. And they could simply not release the modifications they’ve made, as is required by the GNU Public License. But nobody would be able to do anything about it. Currently, copyright laws allow the people who wrote the Linux kernel to sue the company for breaking the license and violating the authors’ copyrights

  • qaz@lemmy.world
    link
    fedilink
    English
    arrow-up
    3
    ·
    edit-2
    20 days ago

    I just followed the opt out link and they’re making you open a public issue with a Markdown list with all your repositories you want removed.

    Surely there has to be a better way to do this (probably the point).

    Also why are they storing 4.71 TB in a Git repo? How are they going to deal with removal requests?

    EDIT: It seems like they’re manually responding to the issues, wtf? There’s a perfectly fine GitHub auth system that they could use to verify everyone’s GitHub account / repository ownership

    EDIT 2: This is apperently a collaboration between Hugging Face and ServiceNow, why is my companies IT ticketing system scraping all of GitHub?

    • Kissaki@programming.dev
      link
      fedilink
      English
      arrow-up
      2
      ·
      20 days ago

      with a Markdown list with all your repositories you want removed.

      The repo readme linked FAQ says

      You can choose to request either (1) all repos, or (2) you can specify select repos that you own to be removed.

      so “all of them” should be acceptable

  • Kissaki@programming.dev
    link
    fedilink
    English
    arrow-up
    3
    ·
    edit-2
    20 days ago

    We want to give developers agency over their source code by letting them decide whether or not it should be used to develop and evaluate machine learning models.

    crawled directly from GitHub and built to pre-train code LLMs with full-repository context

    Repositories that opted out are removed from the dataset before each patch release.

    “agency”

    Which AI company will not use v1 which has all of the data but will use later patch releases instead which have less data?

    • qaz@lemmy.world
      link
      fedilink
      English
      arrow-up
      3
      ·
      edit-2
      20 days ago

      No, it seems to only be a subset of public repo’s.
      I have like 65 repo’s and only 13 were scraped. I don’t get why they specifically scraped those though. They don’t have the most stars, they aren’t the oldest or newest, not the ones with the most forks, nor do I see a pattern based on programming language.

  • Kissaki@programming.dev
    link
    fedilink
    English
    arrow-up
    2
    ·
    edit-2
    20 days ago

    Noteworthy: They crawled only the default branch HEAD and inlined all source content.

    • The file contents are included inline. The decoded UTF-8 source text is embedded directly in the dataset, so it is fully self-contained — you can start training the moment the download finishes.
    • It reflects the state of GitHub in August 2025. The corpus is a direct crawl of GitHub repositories at their default-branch HEAD, capturing roughly two additional years of open-source code compared to The Stack v2.
  • G_M0N3Y_2503@lemmy.zip
    link
    fedilink
    arrow-up
    2
    ·
    20 days ago

    Pretty sure all my repos are MIT licensed for the betterment of everyone, but I’m not on the list! So I guess I’m not good enough, or they are failing to follow the attribution clause of it.

      • Alex@lemmy.ml
        link
        fedilink
        arrow-up
        2
        ·
        20 days ago

        Mine are all GPLv3 or forks of other repos and they are listed. I’m sanguine because the license allows for study and if it’s good enough for humans I don’t see why it’s not for clankers.

    • Miaou@jlai.lu
      link
      fedilink
      arrow-up
      1
      ·
      20 days ago

      My dotfile repo is there and it doesn’t have a licence. Meaning it’s technically not open source. Didn’t stop them