• Em Adespoton@lemmy.ca
    link
    fedilink
    English
    arrow-up
    67
    arrow-down
    1
    ·
    2 天前

    I haven’t used Google as a primary search engine in a decade or more. Nothing I do depends on Gemini (I don’t even use Apple Intelligence that’s a Gemini white label).

    Personally, I think what Gemini is destroying is a market-based financial model that many came to associate with the Internet, information access and lifestyle.

    Most of the websites I visit pre-date Google and some have gone as far as blocking all bots via robots.txt. They survive not via predatory ad networks but by donations and volunteers.

    It’s why I like Lemmy; it follows the same model.

    • buddascrayon@lemmy.world
      link
      fedilink
      English
      arrow-up
      4
      ·
      1 天前

      Just a quick correction. Robots.txt doesn’t block anything, it simply politely asks the search engine to not look here or use what is contained within. A lot of the search engines ignore this and catalogue the contents of the website anyway. And pretty much all of the AI database downloaders ignore it with impunity.

    • floofloof@lemmy.ca
      link
      fedilink
      English
      arrow-up
      42
      ·
      2 天前

      I don’t think today’s bots take any notice of robots.txt. It’s a relic from a more civilized time.

      • Clearwater@lemmy.world
        link
        fedilink
        English
        arrow-up
        13
        ·
        2 天前

        I run a few personal sites with some being available to the wider internet, and I have found that they generally do respect it. All the bots which advertise themselves as OpenAI, Anthropic, or Google do obey. However, I do occasionally see a stealth bot appear which uses a normal browser’s user agent and those just do whatever they want.

        • Ecco the dolphin@lemmy.ml
          link
          fedilink
          English
          arrow-up
          9
          ·
          edit-2
          2 天前

          Meta absolutely does not respect robots.txt. they are scraping what I have at around 200 hits / minute.

          I also notice there is some data center somewhere (or multiple) proxying all its requests through residential proxies so I can’t tell who is scraping. They don’t scrape like Meta though. Far slower. These are the stealthy bois using spoofed user agents you mention. Can’t tell who they are, though. How do they get so many residential IPs?

          I have no organic traffic (I am literally just running a crawler tarpit. My page has nothing) so its really obvious that Im watching AI scrappers. My tarpit generates random links that all resolve to the same place so its super obvious its not human. Its also super obvious when two distinct IPs follow the same random word salad link milliseconds apart.

          • jsproc@lemmy.world
            link
            fedilink
            English
            arrow-up
            6
            ·
            1 天前

            It is a business model. People install free apps on their phone. These apps generate their income by acting as residential proxy for scrapers. They form a botnet of millions of ips, actual phones, without the owners knowing it.

          • Clearwater@lemmy.world
            link
            fedilink
            English
            arrow-up
            3
            ·
            1 天前

            I’ll have to double check my stats later. Haven’t looked at it for a while.

            I certainly don’t recall them ignoring robots.txt, but I also don’t remember seeing Meta in my dash at all, so it’s entirely possible they, for wherever reason, never found my site.

      • Em Adespoton@lemmy.ca
        link
        fedilink
        English
        arrow-up
        1
        ·
        1 天前

        I don’t; I think googlebot is paying attention to it, and AI sites get a LOT of their initial visibility and site ranking from the Google index.

        Case in point: the sites I spend time on that set robots.txt to deny all bots before the LLM wave started still aren’t getting hammered with AI bot requests, while the ones that didn’t are.

        I can see it on my own websites as well; I’ve got one that’s in Google’s index, and it gets a constant low volume traffic from AI bots. The ones that aren’t listed? They just get traffic from exploit scanners.

    • LittleBorat3@lemmy.world
      link
      fedilink
      English
      arrow-up
      7
      ·
      1 天前

      I am back to early web 2.0 in many cases , forums and such. Simply because they needlessly fucked something that worked perfectly before 2015 or so.

    • givesomefucks@lemmy.world
      link
      fedilink
      English
      arrow-up
      16
      ·
      2 天前

      Nah, Gemini is really bad.

      There’s a push to use AI where I work and we have our own Gemini, openAI, and Grok models and they literally check to make people are at least “trying” to use it…

      Gemini will spin “thinking” for maybe 30 seconds, then something like “checking google results” and then just spitting out the same “ai overview” you’d get from searching the same thing.

      All Gemini can do is Google something and summarize the first few results, but at that point a human could just pick a trustworthy site themselves instead of going with what has the best search optimization all mashed up.

      With that one (and likely a lot of others) the big search engines becoming shit from AI slop and search optimization is going to also make AI even shittier, then the top links the next AI will be shittier.

      We’ll get AI rewrites of AI rewrites, repeating itself exponentially faster as more money is sunk into it and hallucinations just keep repeating.

      Like, this is the natural results of monopolies. Corporations get so large they can’t help but ruin the parts that work by integrating parts no one wants to try and make them profitable.

      Because one corporation can’t function and be as large as Google. It’s just too much to keep organized while constantly growing profits. It’s like how the square cube law limits the size of buildings and animals.

      The larger an animal gets, the weaker it gets pound for pound, until it would reach a point where it isn’t even strong enough to breathe…

      That’s where these giant tech companies are headed.

      • UnspecificGravity@piefed.social
        link
        fedilink
        English
        arrow-up
        9
        ·
        2 天前

        This is exactly the kind of circular cascade failure that we are going to see. Once the AI engines all start training on their own data and sourcing each others slop it just gets worse and worse.

        • givesomefucks@lemmy.world
          link
          fedilink
          English
          arrow-up
          7
          ·
          2 天前

          If it wasn’t already an issue, AI companies wouldn’t be destroying old books to guarantee they weren’t training on AI slop…

          The problem is if AI replaces writers, there will never be anything new to train AI on.

          Humans haven’t hit a wall on innovation, were just no longer in the massive boom from early internet.

          Historically the only way to cause those booms is to connect humans together to facilitate the exchange of ideas and collaboration.

          We’re headed so fast to the opposite I don’t know if tech bros are too overconfident to understand they were a product of their time and not the other way around, or if they’re smart and selfish enough that they’re intentionally trying to freeze human innovation so that there won’t be another boom so no other group could clean up and replace them.

    • newbeni@lemmy.world
      link
      fedilink
      English
      arrow-up
      5
      ·
      2 天前

      So, I disabled everything AI the I can think of, use DDG as a search engine, ad block wherever, blah blah blah, my search results SUCK. Is there a better way to fix it?

      And I use Linux at home…no blows crap

      • brb@sh.itjust.works
        link
        fedilink
        English
        arrow-up
        3
        ·
        2 天前

        Kagi is pretty decent but the sad truth is that something like duck.ai works better as a search engine than normal search engines

      • Hudell@lemmy.dbzer0.com
        link
        fedilink
        English
        arrow-up
        2
        ·
        2 天前

        Suck in what way? And what kind of stuff do you usually search?

        I used to think search was kind of an universal thing but after looking at some people’s examples of what is a good or bad result I noticed that is definitely not the case.

        For me a good search result is a list of sites that contain the words I typed on the search bar, sorted by how closely the words match and then perhaps subsorted by some website ranking system. For that, DDG has been quite decent.

        • newbeni@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          18 小时前

          So, I can’t really mention actual results after actual queries, a lot of it would be doxxing myself. Maybe that’s what it is, I’m trying to get way into my personal life instead of a general search. What I usually look for is not like “this specific thing happened to me” type things, but more like a general consensus. Sometimes it’s great, sometimes I’m wondering what the heck happened.

        • FudgyMcTubbs@lemmy.world
          link
          fedilink
          English
          arrow-up
          6
          ·
          1 天前

          Not OP, but I find the DDG results are SEO sites at the top for pages.

          My search: how to fit a dog collar

          Top 50 results from DDG:

          fitting a dog collar is difficult and many people wonder how to fit a dog collar. This article will show you how to fit a dog collar.

          A dog collar is difficult to fit. Perhaps you’ve wondered how to fit a dog collar.

          To be fair, SEO ruined the internet and Internet searches well before AI and it’s not DDG’s fault. But I would love it if blatant SEO was filtered out of my results.

          • Hudell@lemmy.dbzer0.com
            link
            fedilink
            English
            arrow-up
            1
            ·
            1 天前

            True, I think I probably just skip over those results and don’t even register them anymore. They definitely happen a lot but it’s not something I even remembered when I thought about search results.

          • Em Adespoton@lemmy.ca
            link
            fedilink
            English
            arrow-up
            1
            arrow-down
            1
            ·
            1 天前

            Part of it might be how Google trained everyone to use natural language queries in their searches.

            Instead of “how to fit a dog collar” does

            “pet advice” +”dog collar” “guidance for new dog owners”

            give you better results?

            The idea is to search for a cluster of terms that are likely to be on the page you want to find, but aren’t hyper-focused like SEO pages.

        • chunes@lemmy.world
          link
          fedilink
          English
          arrow-up
          2
          ·
          1 天前

          Very curious whether you’re old enough to remember how google used to be in say, 2007.