As AI tech companies increasingly buy and destroy books to feed to their AI models, Anna’s Archive is calling for volunteers to help preserve them for the public record.

  • Apathy Tree@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    3
    ·
    edit-2
    1 小时前

    The only book I have that might be worth scanning is Knight’s Modern Seamanship 13th edition from ~1960 (was provided to my mom as part of a record-finding request regarding my grandpa’s WW2 deployments, I have no idea why), but if my time in the Navy is any indication, its one of those things that was given during basic to all going through basic (I have a modern copy from 2007 as well), so idk if its a valuable contribution. Probably already in the archive.

    I also have a copy of the joy of cooking from the 70s, and a Betty Crocker cookbook from the 50s, but I assume that’s also already in there. Plus some language textbooks and the like, some bird and geology books, some limnology books, etc. the sort of thing that’s definitely already there.

    I’m willing to scan them, however a person does that, but I’m not willing to destroy them. The old books came from my grandparents, through my mom, and all of those people are long dead. Everything else is from my own education and I use them for reference.

  • KurtVonnegut [comrade/them]@hexbear.net
    link
    fedilink
    English
    arrow-up
    13
    ·
    3 小时前

    Job Opening: Now Hiring

    Description: Destroying the Library of Alexandria

    Compensation: Minimum Wage ($7.25 per hour) plus free daily lunch (2 pizza slices)

    Experience Required: Masters Degree in English, and at least 5 years experience working at a library or book store

    Cover Letter: Please include a personal essay describing how destroying books gives you a personal sense of fulfillment

    • Scrath@lemmy.dbzer0.com
      link
      fedilink
      English
      arrow-up
      2
      ·
      32 分钟前

      Obviously destroying books is the perfect job for me as it brings me closer to my german ancestors

  • Catoblepas@lemmy.blahaj.zone
    link
    fedilink
    English
    arrow-up
    16
    ·
    4 小时前

    I have access to a book scanner, but I’m not sure what books need to be scanned, are legal to be scanned and uploaded (I don’t want to get expelled for using a school scanner to do something illegal), or whether there is personally identifiable data in the PDFs it generates (I think I can also save in GIF and maybe PNG?). Is there a beginner’s guide anywhere for all this?

    • Derpenheim@lemmy.zip
      link
      fedilink
      English
      arrow-up
      6
      ·
      2 小时前

      Im gonna be real chief. Stop worrying about IP legality. Scan and upload everything you can get your hands on, wherever you can reasonably upload it.

      AI companies are violating all copyright laws at all times of day and then destroying the originals, or even DMCA claiming them. You cannot fight that and be worried about legality.

    • Truscape@lemmy.blahaj.zone
      link
      fedilink
      English
      arrow-up
      17
      ·
      4 小时前

      Anna’s Archive does have guides for all of that (apart from the legal advice) on their site (use wikipedia to find the right domain).

      I don’t think you’d want to do so as a university student unless you want to share the same fate as Aaron Swartz though.

      • Catoblepas@lemmy.blahaj.zone
        link
        fedilink
        English
        arrow-up
        9
        ·
        3 小时前

        Surely at least the stuff out of copyright is kosher to scan and share? Which seems like a good area to focus on if they’re being bought up and destroyed

        • dan@upvote.au
          link
          fedilink
          English
          arrow-up
          11
          ·
          edit-2
          2 小时前

          Stuff that’s out of copyright isn’t an issue though.

          The reason AI companies are destroying books is because current US copyright caselaw (most recently the Anthropic lawsuit) doesn’t allow books to be copied, but considers it fair use if you transform the format of a legally purchased book (like from print to digital) without creating a new copy. By destroying the original book, they haven’t made a copy of it, and so they operate within US law.

          This doesn’t apply to public domain books (books that are no longer copyrighted) as you can do whatever you want with public domain content.

        • Maybe there is a way to do this anonymously? Anna’s Archive deals with books new and old, but maybe that ‘physical-only’ book that has no digital print would be the best place to start?

          I remember no-eyes from #ebooks (IRCHighWay) prioritized recommendations of novels (fiction) to expand their available books, so I think that’s the safest type of book to scan?

    • Fluffy Kitty Cat@slrpnk.net
      link
      fedilink
      English
      arrow-up
      6
      ·
      3 小时前

      Try your beat to cover your tracks while you do it, but if you can scan some rare books, please do. You’d be a hero

  • Evilsandwichman [none/use name]@hexbear.net
    link
    fedilink
    English
    arrow-up
    5
    ·
    3 小时前

    There’s a ton of ancient documents and books that don’t have copies of them (Chinese stuff for example that wasn’t all digitized), please tell me data centers aren’t seriously trying to take those as well and burn those as well

    • Pistachio@lemmy.zip
      link
      fedilink
      English
      arrow-up
      5
      ·
      2 小时前

      I think its a solid bet to say that if there’s a price on it, data centers will pay.

      That being said, China is famously protective of their shit. So if its in their hands, its probably safe.

  • reallykindasorta@slrpnk.net
    link
    fedilink
    English
    arrow-up
    5
    ·
    4 小时前

    We could probably recruit the special collections departments at university libraries to help since they’re already equipped, but how do we identify and procure what needs to be preserved? I kind of assumed most books were already being digitized. As a student I worked on a project digitizing old newspapers and entering basic metadata.

    • Fluffy Kitty Cat@slrpnk.net
      link
      fedilink
      English
      arrow-up
      6
      ·
      3 小时前

      I sometimes try and track down rare books from the 20th century, sometiems less than 75 years old, only to.run into the problem that the only copy is in a rare books collection in a Canadian university or some shit like that