• Ecco the dolphin@lemmy.ml
    link
    fedilink
    English
    arrow-up
    6
    ·
    edit-2
    9 hours ago

    Meta absolutely does not respect robots.txt. they are scraping what I have at around 200 hits / minute.

    I also notice there is some data center somewhere (or multiple) proxying all its requests through residential proxies so I can’t tell who is scraping. They don’t scrape like Meta though. Far slower. These are the stealthy bois using spoofed user agents you mention. Can’t tell who they are, though. How do they get so many residential IPs?

    I have no organic traffic (I am literally just running a crawler tarpit. My page has nothing) so its really obvious that Im watching AI scrappers. My tarpit generates random links that all resolve to the same place so its super obvious its not human. Its also super obvious when two distinct IPs follow the same random word salad link milliseconds apart.

    • Clearwater@lemmy.world
      link
      fedilink
      English
      arrow-up
      1
      ·
      5 hours ago

      I’ll have to double check my stats later. Haven’t looked at it for a while.

      I certainly don’t recall them ignoring robots.txt, but I also don’t remember seeing Meta in my dash at all, so it’s entirely possible they, for wherever reason, never found my site.