Yesn’t. What LLMs lack and causes these idiotic suggestions is context from mediums other than language. If you get down to it humans talking are also just trying to predict what should come next, but we predict this from more input than just text. Particularly with something like design, we come up with the design and then the text and maybe concept art to describe it. The llm comes up with a description that seems likely and then feeds it to a diffusion model to make an image to go with the text.
LLMs can be multimodal now. You can chuck a screenshot into claude code or whatever you use and the image generation process can also be done by the LLM itself rather than an external tool call as far as I know
Yesn’t. What LLMs lack and causes these idiotic suggestions is context from mediums other than language. If you get down to it humans talking are also just trying to predict what should come next, but we predict this from more input than just text. Particularly with something like design, we come up with the design and then the text and maybe concept art to describe it. The llm comes up with a description that seems likely and then feeds it to a diffusion model to make an image to go with the text.
LLMs can be multimodal now. You can chuck a screenshot into claude code or whatever you use and the image generation process can also be done by the LLM itself rather than an external tool call as far as I know