惯性聚合 高效追踪和阅读你感兴趣的博客、新闻、科技资讯
阅读原文 在惯性聚合中打开

推荐订阅源

Microsoft Azure Blog
Microsoft Azure Blog
有赞技术团队
有赞技术团队
IT之家
IT之家
博客园 - 聂微东
Jina AI
Jina AI
Hugging Face - Blog
Hugging Face - Blog
Last Week in AI
Last Week in AI
Apple Machine Learning Research
Apple Machine Learning Research
WordPress大学
WordPress大学
小众软件
小众软件
爱范儿
爱范儿
奇客Solidot–传递最新科技情报
奇客Solidot–传递最新科技情报
V
Visual Studio Blog
雷峰网
雷峰网
酷 壳 – CoolShell
酷 壳 – CoolShell
阮一峰的网络日志
阮一峰的网络日志
宝玉的分享
宝玉的分享
博客园 - 三生石上(FineUI控件)
大猫的无限游戏
大猫的无限游戏
博客园 - Franky
量子位
月光博客
月光博客
博客园 - 【当耐特】
博客园 - 叶小钗

sebszyller.com | Blog

LLMs Will Brand Themselves Out of Existence Trustworthy ML Is a Kitchen Sink Best Practices for a New Research Project Repo in 2026 Looking Back at Learning Rust with LLMs: Works 100% of the Time... 50% of the Time Fairness Is Hardly About DEI DeepSeek Drama -- Model Watermarking to the Rescue Open Weights Have Nothing to Do with Open Source Learning Rust with Large Language Models (Part III): Finding a Needle in a Haystack Learning Rust with Large Language Models (Part II): an Outdated Manual Written by a Newbie Food Markets Are a Bad Analogy for Data Marketplaces Learning Rust with Large Language Models (Part I): a Project for 2024 No One Cares About Large Language Models Anymore Content Provenance Needs Critical Mass Data Marketplaces for Individuals (Still) Don't Make Sense The Security Through Obscurity Moment of Large Language Models... and Money Can You Spot a Deepfake? Kosher Data for Your Ethical Needs Why Synthetic Data Is Not Private Data Marketplaces for Individuals Don't Make Sense On the Difficulty of Cross-disciplinary Communication Whose Model Is It Anyway?
Finding the Beauty in the Imperfections of Generative Art
Sebastian Szyller · 2022-09-05 · via sebszyller.com | Blog

Generative art has taken over the internet by storm. It fails in more than one way but do we really care?

A world view centered on the acceptance of transience and imperfection. The aesthetic is sometimes described as one of appreciating beauty that is “imperfect, impermanent, and incomplete” in nature. This is how the Wikipedia describes wabi-sabi — a concept popular in Japanese art.

A city floating in the sky.

Generated with MidJourney using the prompt: "a lost city in the sky 4k artstation". Picture source.

Cool new stuff

In the last couple of weeks, we’ve seen a lot of art generated using text-to-image models, mostly DALL·E, MidJourney and Stable Diffusion. They aren’t new as a concept but there has been a lot of progress in recent years.

If somehow you missed all the hype, in essence, they are deep learning models that take a description of a scene, aka the prompt, as the input and try to produce an image that represents it the best.

They vary a bit in what they’re good at. In my experience, Stable Diffusion is the best at producing a high quality image, even if it ignores or misinterprets some of the prompt. MidJourney is quite similar but more in an artsy sense. As far as I know, its authors used more training data from art-sharing websites such as ArtStation and DeviantArt. Lastly, DALL·E, which is the best at capturing the relationships between the objects in the prompt, even if the image is not as impressive as from the other two models.

Two corgi dogs in a karate fight

Generated with DALL·E using the prompt: "a cinematic photo of two corgis wearing kimonos in a karate fight". Notice how the partial kimono is actually made of dog's hair. Picture source.

Output artifacts

Apart from misunderstanding the prompts, generated images can still have many artifacts — flaws in the output caused by the quirks and rough edges of the model. These can include materials melted into each other, hair growing out of fabrics, colours overflowing from one item into another, unintended psychedelic textures stemming from the internal representations, and many more.

We consider them flaws but is that all they are?

Screenshot from a game with pixel graphics

Pixel based graphics are a creative choice that can give a sense of nostalgia. Picture source.

New art direction

It’s 2022 and we still make games with pixel graphics even though we can render almost life-like images. We use hand drawn illustrations as assets, even though it’s incredibly expensive. We do it because we enjoy the style, what it stands for, and the feelings it invokes. It’s a creative choice.

In the same vein, I think the artifacts of generative models are a style, an art direction to pursue.

They are a snapshot of the current era. Even though as time goes by, we’ll get rid of most those flaws. The fuzzy, texture clipping generative art might be the vibe that we associate with the 20s. The psychedelic AI hallucinations — a perfect match for the global pandemic, war, and looming climate change.

Though it isn’t just about the style. Accessible tools allow more people to participate in the creative world and express themselves. Do you want to make a realistic render of your child’s crayon drawing? Well, now you can. Are you running a Dungeons & Dragons campaign, and need some concept art? Just generate it with a prompt. Do you want an edgy edit of your math’s teacher photo that gets you expelled from school? You got it. A trend with a lot of momentum is what defines an era.

However, this isn’t to say that there aren’t any challenges. Collectively, we need to figure out how to handle the copyright and ownership of generative art (more on that soon). How to compensate the authors whose work ends up used for training models? How to bring this technology to people that cannot afford to pay for the service or beefy computers?

To wrap it up, check out this eerie video posted on Reddit recently; made with Disco Diffusion.