Cartoony fonts, way too much text everywhere, no visual higherarchy, colored rounded boxes, why does it all have seemingly one style rather than having multiple like you see in other ai mediums?

I hate generative ai btw I was just wondering

  • Grandwolf319@sh.itjust.works
    link
    fedilink
    arrow-up
    20
    ·
    7 hours ago

    Generative AI tools are by definition averaging machines. They take lots of data and give an average output based on input.

    Not all training data is as valuable as others. The most valuable are distinct and highly associated with a certain tag or description. So a picture that is just a dog with over exaggerated dog features is better training data than a whole landscape with tons of stuff including a dog. This makes cartoonish stuff a better fit for training data.

    Similarly, people using these tools tend to be less detail oriented and just want something that has an obvious, center stage, flat message. They want something fast and cheap so far less thought goes into it than traditional art.

    Given these two, you end up with a huge bias towards the bland cartoonish style.

    Beyond that, there is also the fact that as time goes on, generated images find their way into training data, so it keeps getting more closer to that bland style.

    Seriously, I feel like the first generation of stable diffusion models were superior to what we have now.

    • dalekcaan@feddit.nl
      link
      fedilink
      arrow-up
      12
      ·
      11 hours ago

      That’s what really gets me about everyone who touts AI as the future of creativity. All it does is give you what it determines to be the most likely response to your prompt. It’s never going to be clever or subversive.

      Even if you wade through all the slop it shits out and only go with the results closest to what you asked for, the absolute best you can hope for is predictable and mediocre. If you’re not interested in making something at least halfway decent, why bother at all?

    • mrmaplebar@fedia.io
      link
      fedilink
      arrow-up
      23
      ·
      12 hours ago

      It’s literally that.

      Trained on everything, biased towards middle-of-the-road, AI is going to produce the most generic “art” possible, unless specifically prompted to do something else.

  • Fallibilist@feddit.uk
    link
    fedilink
    arrow-up
    41
    arrow-down
    7
    ·
    13 hours ago

    It doesn’t - you only notice the low-effort ones that default to the generic style where as the rest fly under your radar and go undetected. It’s called a toupee fallacy: “All toupees look fake, I’ve never seen one that didn’t”

    • Starya67@lemmy.world
      link
      fedilink
      arrow-up
      3
      ·
      11 hours ago

      So is it still possible to give a style prompt? I haven’t dabbled with it in ages. When it started out, you’d tell it you wanted, say, the style of Rembrandt but with a dash of horror or something. Can you still do that?

      • cecilkorik@piefed.ca
        link
        fedilink
        English
        arrow-up
        1
        ·
        6 hours ago

        it’s a language model. You can give it any kind of prompt you want. How effectively it interprets those instructions is a separate discussion, but in general newer models expand on those sort of capabilities, they don’t strip them out.

        • Starya67@lemmy.world
          link
          fedilink
          arrow-up
          1
          ·
          5 hours ago

          Aight. I just tried again for science. I gave it very detailed instructions and it produced a person with three arms. I did not specify three arms.

          • cecilkorik@piefed.ca
            link
            fedilink
            English
            arrow-up
            1
            ·
            3 hours ago

            “Tried again” is pretty vague. There are millions of different text-to-image models and adapters and millions of different ways of configuring and running them and thousands of different services providing these technologies at different service tiers. I’m going to go out on a limb and assume you’re using a free service, which is obviously going to be the worst possible option using the cheapest possible model, there are likely self-hosted free models that can do better. You can’t just assume you’re using a newer model if you don’t even have access to describe what model you’re running.

            • jacksilver@lemmy.world
              link
              fedilink
              arrow-up
              4
              ·
              3 hours ago

              I like how this sounds kinda like a bratty response, but is legitamte advice.

              If you’re not aware, there is the concept of “negative prompts” with image generation where you basically tell it not to draw something ugly or wrong, but also can try to supress certain things too like three arms.

  • PiraHxCx@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    61
    arrow-down
    1
    ·
    14 hours ago

    That’s actually because people doing AI art aren’t artists, they have absolute no artistic vision and most probably lack basic art knowledge, and the very reason they are doing AI art is because they are mediocre. Their prompts are “give me [actor] doing [action] in [location]” and that’s it, they don’t talk about style influence, color palette, scene composition, etc etc, so it all comes out the same.

    • Shimitar@downonthestreet.eu
      link
      fedilink
      English
      arrow-up
      24
      arrow-down
      5
      ·
      13 hours ago

      This. That’s like vibe coding. Vibe arting get you slop as well.

      Now, give an AI to a good coder, and you don’t get slop code. Same with artists

      Do you still get AI stuff? Yes, but a whole better result.

      • Nibodhika@lemmy.world
        link
        fedilink
        arrow-up
        23
        arrow-down
        4
        ·
        13 hours ago

        Except good artists and coders enjoy the creation process, so they won’t use it like that.

        I need to change every instance of this function to another? Sure, let an LLM do it, that’s tedious to do and easy to verify. Implement a new system? Hell no, that’s interesting to do and hard to verify. I imagine artists have similar feelings on the matter, but I don’t know what’s a tedious task that’s easy to verify for an artist.

        • PiraHxCx@lemmy.dbzer0.com
          link
          fedilink
          English
          arrow-up
          5
          arrow-down
          1
          ·
          edit-2
          5 hours ago

          if you are just doing a photo composing or a design, I guess “remove fake png background” “remove watermark” hehe
          (although content aware tool already does a great job)

  • sbird@retrofed.com
    link
    fedilink
    English
    arrow-up
    8
    arrow-down
    1
    ·
    10 hours ago

    I would imagine that it’s the easiest for the model to produce. Bold colours, clear and distinct shapes, that sort.

    It doesn’t help that the companies have scraped the entire internet (which is a can of worms in terms of both legality and ethics) and then some, so it will mimic the advertising strategies employed everywhere (again, bold colours, clear shapes, etc.)

    • Ephera@lemmy.ml
      link
      fedilink
      arrow-up
      2
      ·
      8 hours ago

      I would assume, they even explicitly tell the model to use this default style, if the user doesn’t specify anything else, because it is so easy to generate.

      At least, this started early last year and there were fairly distinctive points when each chatbot’s outputs started looking all the same.

  • usernamesAreTricky@lemmy.ml
    link
    fedilink
    arrow-up
    19
    ·
    14 hours ago

    This was for general diffusion model outputs a few years ago, but a lot of the speculation of why the outputs are similar likely hold here

    There are also only so many good data sets available for people to use to build image models, Phillip Isola, a professor at the MIT Computer Science & Artificial Intelligence Laboratory, told me, meaning the models might overlap in what they’re trained on. (One popular one, CelebA, features 200,000 labeled photos of celebrities. Another, LAION 5B, is an open-source option featuring 5.8 billion pairs of photos and text.)

    […]

    Five years ago, he explained, image generators tended to create really blurry outputs. Researchers realized that it was the result of a mathematical fluke; the models were essentially averaging all the images they were trained on. Averaging, it turns out, “looks like blur.” It’s possible that, today, something similarly technical is happening with this generation of image models that leads them to plop out the same kind of dramatic, highly stylized imagery—but researchers haven’t quite figured it out yet. Additionally, “most models have an ‘aesthetic’ filter on both the input and output that reject images that don’t meet a certain aesthetic criteria,” Hany Farid, a professor at the UC Berkeley School of Information, told me over email. “This type of filtering on the input and output is almost certainly a big part of why AI-generated images all have a certain ethereal quality

    […]

    The third theory revolves around the humans who use these tools. Some of these sophisticated models [this is not as sophisticated as the author is implying] incorporate human feedback; they learn as they go. This could be by taking in a signal, such as which photos are downloaded. Others, Isola explained, have trainers manually rate which photos they like and which ones they don’t. Perhaps this feedback is making its way into the model

    […]

    This could be intentional: If such imagery has a market, maybe companies would begin to converge around it. Or it could be unintentional; companies do lots of manual work in their models to combat bias, for example, and various tweaks favoring one kind of imagery over another could inadvertently result in a particular look

    https://archive.is/TlzVj

    • Sibbo@sopuli.xyz
      link
      fedilink
      arrow-up
      1
      ·
      13 hours ago

      That’s interesting. I always thought it was a conscious choice by model trainers.

  • Zwuzelmaus@feddit.org
    link
    fedilink
    arrow-up
    3
    arrow-down
    7
    ·
    15 hours ago

    You are asking: “why are all AI users the same kind of stupid and want it this way - at least all the ones that I am watching”