Important update on the previous first experiment.
In a previous post, I shared an experiment I did some months ago to see if ChatGPT (4o at the time) could draw a lion from scratch, something that even a five-year-old child could do, even if the drawing they produce comes out amateurish. This is different from what most people are accustomed to when AI generates entire images in a flash. I wanted to see if ChatGPT can draw in real time, watching it move the mouse cursor on its own to produce any depiction of a lion.
I told my co-worker the other day about my experiment. I told him that it didn’t produce nearly anything that I imagined it would, for it was merely scribbles, and that this kind of task, and drawing in general, requires a different kind of intelligence. Boy, was I wrong! Now, I have to share this unexpected update of the same experiment done this morning with the latest ChatGPT model: 5.6 Sol at max settings on the app.
I have been experimenting with the latest ChatGPT on Codex to explore different use cases. One of the things I was testing was building different video games for the past two weeks. It started with simply handling how people make Pokémon ROM hacks and trying to learn a bit about that using tools like Porymap, all the way to creating an entire retro 2D Star Wars-esque Jedi game. Needless to say, for the Pokemon ROM Hacks, it was struggling a bit to create what I had in mind, but some of what was due to the restrictions of the software, and me just giving up early.
I learned the hard way that making a video game from scratch isn’t easy at all. I couldn’t just vibe code a video game out of thin air. No matter how much I tinkered and asked Codex to make the game better, etc., it wasn’t up to par with any video game, let alone the handheld games I used to enjoy playing as a kid.
Anyway, in the attempt to understand why or how I could get ChatGPT to draw something that resembles a lion, or why this could be a big impediment, I used my speedy deep research workflow to investigate this inquiry, which I don’t think I shared in my previous post. One of the flaws or limitations of outsourcing research to AI is the missing or overlooked details that can generate tremendous insights. Nonetheless, it is still a useful but imperfect research report to get a general overview of what you’re inquiring about. The final product of that could be seen in the NotebookLM video below:
This morning, I was tinkering with Codex to create a website and remembered what I mentioned to my co-worker that I wanted to test the lion sketch benchmark with the new ChatGPT model when it comes out. He wanted to know the results of it, and he shall in the coming days. Surprisingly, the token limitation didn’t run out when creating a website; that’s when I decided to give it a test run. I opened a new chat on Codex and simply prompted:
Go on this website and sketch a lion from scratch: https://sketch.io/sketchpad/.
As ChatGPT was setting up and starting, I recalled that I forgot to record, and I instantly just got my phone camera on to record my laptop screen in real time, because I normally don’t screen record. Random off-side topic: on the website I was creating through Codex just before the lion experiment, there was a point where ChatGPT/Codex wanted to screen record, but that didn’t feel safe to me, so I denied that request. It was a strange and unsettling moment.
Back to the topic at hand. While I was watching through my phone, in real time, ChatGPT draw, I was confused about all the steps it was doing at first. It took a long time to select the colors. It started with a red sort of sun-like shape. My inner dialogue was like, “Okay, yes, that can be the head.” I was wondering where it was going with this.
Initially, I have to admit I didn’t have high hopes and seemed to think that it was going to eventually fail to draw the lion that I requested.
Then ChatGPT suggested it was making the body next, and I looked at it and was confused about its shape, and it did use a yellow-mustard color. I figured surely it was going to fail now.
But then something spectacular happened.
ChatGPT then said it was going to draw the face/head, and it started drawing eyes in the sun-like shape area. I was amazed at what I was witnessing. It made an entire happy face with black. My jaw dropped. In 8 minutes and 12 seconds, ChatGPT produced what I can definitively say is a drawing of a lion! I witnessed and recorded it in real time because scientists must have their empirical proof.
Here’s a screenshot of it:

Maybe I’m overblowing it, and maybe other LLMs could have done this before, but I was ignorant of this fact. Or maybe this latest model is just simply leaps and bounds better than the previous one. Sketching a lion was a huge benchmark for me towards AGI. Note, I don’t have any AI “benchmark” per se. This is a big deal to me, and my mind is blown. Even if it isn’t perfect and seems like a five-year-old child made this depiction of a lion, it is nonetheless impressive to say the least. ChatGPT passed my benchmark and made my previous post almost irrelevant now. This is a “holy shit!” moment.
“That’s one small step towards AGI, one giant leap for AI.”
// Note: My phone that was recording this process is like 5GB, probably because I had a high setting on and this was literally on the spot, quick on my feet thinking, and that’s why I didn’t upload it here, but it is uploaded in my Google Drive.