For her safety, Doe has opted to receive alerts from the US Department of Justice Victim Notification System any time she may be a victim in a new criminal investigation. Although she has received countless alerts, she was shocked when the CCCP notified her that it had identified AI-generated CSAM on xAI that depicted her. This re-traumatized Doe, whose complaint alleged that messages were found on online forums “between offenders chatting about creating AI generated CSAM of Plaintiff and other similarly situated known, legacy, victims of CSAM.”

Now, Doe fears that xAI has not only made it easier to make more violative images of the most distressing time in her life, but also that xAI allegedly has stored the images that Grok generates and uses those outputs to further train Grok. Because of this, she believes that Grok has been trained on both the initial set of images that have haunted her for more than 20 years and the more recent AI-generated ones.

This is the first case to accuse xAI of training on CSAM, and the complaint does not go into great detail on that claim. Previously, Ars reported on a controversial dataset that was later scrubbed after researchers found CSAM in the training data, but there’s no indication xAI trained on that data. In a press release from lawyers representing Doe, it explained that Doe’s images were included in a CSAM Hash List maintained by NCMEC, and “that same material” allegedly “was part of the dataset xAI used to build Grok’s image and video generating capabilities.” The complaint similarly only alleged that “CSAM depicting Plaintiff with its longstanding well-known hash values has been used as a part of the dataset used by xAI.”

  • QueenHawlSera@piefed.social
    link
    fedilink
    English
    arrow-up
    1
    ·
    11 minutes ago

    I knew Elon was a pedophile (he begged to be on the island) but to actually train your AI on CSAM…

    Well he should be arrested for possession of child pornography and Grok should be shut down until all of its CSAM is erased from its databanks

  • NathanRanch@lemmy.zip
    link
    fedilink
    English
    arrow-up
    1
    ·
    1 hour ago

    Look up AI over fitting. Because of the math, every answer can be provided in terms of how much CP data was used in every answer.

  • Waterpumpee@lemmus.org
    link
    fedilink
    English
    arrow-up
    101
    ·
    6 hours ago

    This should make the entire model illegal. Train it on illegal stuff: delete the whole model.

    • Voroxpete@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      24
      ·
      4 hours ago

      In all seriousness, there are some very interesting legal questions that will be raised if this case makes it that far.

      The problem is that there’s no existing law that would effect this on its own. To my knowledge, no country in the world has a law on the books specifically dealing with AI models trained on CSAM. So the question, under existing laws, would turn on whether the data stored within the model itself would constitute CSAM.

      The problem, in no small part, is that we have serious gaps in our public consensus knowledge about how LLMs actually work.

      There’s a case that, AFAIK, is still being argued in Germany pushing the theory that LLMs actually do, in effect, store a copy of all their training data, just in a compressed form. This certainly seems to hold some water given both the tests they relied on, and the situation with this Jane Doe where the model produced images so alike to real images of her that they tripped hash detections.

      The German case argues that this is analogous to the difference between an MP3 and a WAV, or a JPEG and a PNG. That sharing a lossy copy of a work is no less infringing just because it’s imperfect.

      If the underlying claim - that LLMs function as a form of lossy compression - can be substantiated then there would be a real argument that the model itself would constitute CSAM. Since there would be no realistic method that I’m aware of for removing the offending material from the model - and presumably SpaceX would have to somehow prove that they’ve done so - that would make the entire model contraband. They’d have to retrain on a clean dataset.

      Of course I said “if the case makes it that far” at the top because I don’t think it will. SpaceX will do anything and everything to avoid handing over meaningful discovery in this case, including, I suspect, outright destruction of evidence. If there is anything that actually proves that they used CSAM in the training data then they are so far beyond fucked that there’s simply no downside to further illegality in pursuit of concealing their crimes. They have the world’s wealthiest asshole in a position to throw literal billions at making this go away. I genuinely wouldn’t be surprised if people turn up dead off the back of this if that’s what it takes.

      • Lemming6969@lemmy.world
        link
        fedilink
        English
        arrow-up
        1
        arrow-down
        1
        ·
        23 minutes ago

        Models do not store the original, they store the model, which is a huge graph of probabilities.

      • cantstopthesignal@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        8
        ·
        3 hours ago

        It’s illegal to posses. Anyone training the models on that is criminally liable for possession as is the company. And it’s a conspiracy to posses that was directed by someone, which is now RICO. Will they prosecuted? No.

        • Voroxpete@sh.itjust.works
          link
          fedilink
          English
          arrow-up
          1
          ·
          2 hours ago

          Sure. Never argued against that. I was discussing the assertion that the model itself would be illegal as a result. Different thing.

      • Scrubbles@poptalk.scrubbles.tech
        link
        fedilink
        English
        arrow-up
        16
        ·
        4 hours ago

        including, I suspect, outright destruction of evidence

        If it was trained on it, that means they are in possession of it, which that right there is straight to jail. I have a feeling they’re scrubbing everything they can right now as we chat

        • CorneliusTalmadge@lemmy.world
          link
          fedilink
          English
          arrow-up
          6
          ·
          3 hours ago

          Was going to say basically say the same exact same thing. There is no law saying how AI is handled when using stollen materials or other “illegal” content.

          But somehow we have all been brought to believe that somehow “new” technology isn’t subject to existing laws.

          If x or any other company downloaded CSAM everyone in the company should be arrested.

          • Voroxpete@sh.itjust.works
            link
            fedilink
            English
            arrow-up
            3
            ·
            1 hour ago

            If x or any other company downloaded CSAM everyone in the company should be arrested.

            I assume you cannot possibly mean that as written, right?

            I’m absolutely for arresting anyone who was involved with this, or had knowledge that it was happening. But we’re obviously not talking about going after Jane the intern here, right?

            • Scrubbles@poptalk.scrubbles.tech
              link
              fedilink
              English
              arrow-up
              2
              ·
              49 minutes ago

              I won’t say everyone. I’ve actually been at a company who was investigated (not for CSAM, but other things that happened). I had no idea it even happened, and luckily was not involved with any of it. So for me no, I wouldn’t have wanted that. That being said 2 things, say I had been in the position. If our scraper was downloading it and it was my scraper, damn right I would have flagged it to legal, HR, and everyone I could have, along with writing some way to prevent it, and written everything down in a complete log (off company computer). If it wasn’t stopped immediately I would either quit, whistleblowed, or happily talked with anyone raiding and making sure any of the decision makers were hauled off. I don’t blame someone for being lowest level at a shit company, been there. (Although I will say, xAI, come on, no one is “stuck” there, but that doesn’t mean that Dave the brand new intern out of college should be hauled off). Who I blame are the suits who were probably told it was happening and chose to ignore it, and any engineer I do blame if they knew about it and chose not to do one of the above.

              As an engineer I’ve had my fair share of let’s say… challenges that I’ve had to morally grapple with. Things I’ve been asked to do that may not be moral. However, there’s a pretty wide chasm between “Implement this dark pattern so people won’t unsubscribe” and “host this and don’t tell anyone”

            • CorneliusTalmadge@lemmy.world
              link
              fedilink
              English
              arrow-up
              1
              ·
              45 minutes ago

              If you started a company that made CSAM do you think you and every one involved should be let off, because you claim it was a technological oopsie?

              People have been pointing out that xAI is a CSAM machine since basically the first day it came out. And when they basically said they don’t care. All the employees that stayed are all accomplices to the crimes… let alone the people who started working there after.

              The only way to stop these companies is to start holding people accountable. And they will never hold the people at the top to account.

              So make it hurt for everyone involved and then people will think twice before they sign up to do evil deeds for evil people.

      • Bane_Killgrind@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        3
        ·
        2 hours ago

        I think there’s no real change that needs to be made to the law. The images are stenographically embedded into these models. There exists some generic input data that results in the reproduction of the embedded data.

        This is not really any different than a password protected zip file.

  • Blaster M@lemmy.world
    link
    fedilink
    English
    arrow-up
    34
    ·
    7 hours ago

    amazing… I know of smaller models that specifically avoided that sort of thing because even negative training could go wrong so easily… they could have avoided that thing entirely, but nope…

    I could understand if they were training a safety classifier (a model that learns what danger stuff is so it can recognize it when it sees it) but I doubt this is what they were doing here. That and handling such a radioactive dataset makes handing the demon core seem safe.

    • Kaligalis@lemmy.world
      link
      fedilink
      English
      arrow-up
      4
      ·
      5 hours ago

      They probably let it just eat the internet without any humans looking at the training data. And yeah - there probably is some CSAM somewhere on the internet. Maybe, they used an agent swarm to specifically search for stuff not yet in the dataset and forgot to a blocklist. It’s not like the tech bros are genuinely careful in what they do. Recklessness seems to be a common trait.

  • frustrated_phagocytosis@fedia.io
    link
    fedilink
    arrow-up
    19
    ·
    7 hours ago

    If it’s using csam images to produce new ones, is there any way to guarantee that any given nude it makes didn’t source csam? Is the whole thing poisoned at this point?

    • pelespirit@sh.itjust.works
      link
      fedilink
      English
      arrow-up
      2
      ·
      edit-2
      2 hours ago

      I mean, this is the argument for artist’s copywritten works. There is crossover here. The fights over AI are going to be endless.

      To answer your question, yes it’s poisoned.

      • frustrated_phagocytosis@fedia.io
        link
        fedilink
        arrow-up
        1
        ·
        2 hours ago

        That’s why I don’t understand why anyone with money to invest in ai doesn’t have a lawyer or lawyers waiving red flags about these issues. It’s copyrighted materials being illegally used, output being machine derived and not able to be copyrighted, the use of csam for training models that output pornography, and on and on. Yet billions of dollars are dumped into what looks like a black market with no concern for how these issues will be settled.

    • Kaligalis@lemmy.world
      link
      fedilink
      English
      arrow-up
      5
      ·
      5 hours ago

      You can’t ever guarantee that a neural network isn’t dreaming of digital sheep - neither for an artificial nor a natural one.
      What once has been seen can’t be made unseen again. In that, the clankers are like us.

    • Kirp123@lemmy.world
      link
      fedilink
      English
      arrow-up
      9
      ·
      6 hours ago

      Yeah. All the results are tainted, even more than they were from the simple fact it was used by loser to creep on women.