Trained a neural net on my cat and regret everything
Recently I had my first-ever experiment with training a neural net to generate images, when I trained StyleGAN2 to generate screenshots from The Great British Bakeoff. Although recognizable as the baking show, it was a highly distorted, nightmarish version - it would have helped had the neural net been able to keep track of how many faces humans have.
The problem was that the task was too hard. StyleGAN2 has good success when it only has to do human faces from the front, but when its task is broader, like generating internet cats, it struggles hilariously. So I decided to give it a more restricted cat dataset: just one cat. My cat. In more or less the same pose.
After a couple of data-collection attempts that were sabotaged by Charâs tendency to put her nose in the camera, I managed to get a full two minutes of video of her sitting still, staring intently at something invisible in the corner of the kitchen. Probably ghosts. It doesnât matter. The point is, I got the video.
I extracted 3,934 frames from the video and used runwayml to train StyleGAN2 on them for 3000 iterations. Training took about 5 hours, and by the end, it was producing this:
Now, they arenât perfect, but they are much much better than the Bakeoff images. The cat spent the most time looking off to the right, so the neural net learned that pose the best.
But my word, to GET to this point the model had to travel through terrible places. You see, I started with a version of StyleGAN2 that had been pretrained on human faces, which saved a lot of time since the AI already knew how to do edges and textures and hair. But it means that the process of switching gears from generating humans to generating cats yielded a stage that looked likeâŚ. this.
This is one of the most horrible things I have ever seen, and I saw CATS (2019) in the theater.
Hereâs a video of the full training process.
What is particularly alarming, perhaps, is the PATH it took from human to cat. Rather than transforming human eyes, nose, and mouth into the feline equivalents, it erases all the human features and then forms the cat features out of featureless fur. Hereâs a snapshot of the training process in which you might be able to see how the cat/people have four eyeballs, or no eyeballs. (I adjusted the contrast to make these more visible, as these steps got really dark for some reason, as if to shield human eyes from what was happening).
What is great/horrifying about this is that if you have a computer, an internet connection, and a two-minute video of something, you can do this too. Got a cat? A bunch of Easter eggs? A lemon? Itâs now seriously easy to generate your own horrible abominations. The simpler and more consistent your subject, the more recognizable your results will be.
Subscribers get bonus content: more amazing pictures that wouldn't fit in this post.
My book on AI is out, and, you can now get it any of these several ways! Amazon - Barnes & Noble - Indiebound - Tattered Cover - Powellâsďťż - Boulder Bookstore (has signed copies!)
Anya is live and ready to show you everything. Watch her strip, dance, and perform exclusive shows just for you. Interact in real-time and make your fantasies come true.
â Live Streamingâ Interactive Chatâ Private Showsâ HD Qualityâ Free Actions
Free to watch ⢠No registration required ⢠HD streaming
Has science gone too far? Inspired by Janelle Shane (aka @lewisandquarkâ), Iâve taught a neural network how to draw my cat Mulder. The results are ... delightfully off-kilter.
Hereâs what the process started with:
Hereâs the ALARMING midpoint:
And here, after 3,000 iterations, is the rather impressive result!
The process is actually easy enough a non-computer-geek like me could figure it out â thereâs no coding involved, nor do you need a ton of computing power. All you need is a video of your pet (or whatever you want to see AI try to recreate), and two pieces of free software: VLC Media Player and RunwayML.
DETAILED TUTORIAL BELOW THE CUT... If you use it, please reblog this post with your results -- Iâd love to see âem! Note: Iâm a Windows user; the process for Mac might be slightly different at some steps.
Step One: Gather your dataset
In order to train an AI, youâll need to feed it a sampling of 500-5,000 relevant images. Rather than take all those pictures individually, weâre going to use VLC Media Player to extract the frames from a video you take.
1. Record a roughly two-minute video of your subject (eg., your own cat, a mushroom, your mug collection, your own face if youâre really brave, whatever). The subject should be centered in the frame. For the best (most coherent) results, youâll want good lighting and contrast, for your subject to remain relatively stationary.
2. Download the video onto your computer from Google Photos or whatever cloud service your phone backs up onto.
3. Download and install VLC Media Player. If you already have VLC Media Player, ensure youâre running the most recent version by clicking âHelpâ and then âCheck for updates.â
4. Run VLC Media Player as administrator. (Right-click the shortcut on your desktop or in the start menu and select âRun as administrator.â)
5. Hit ctrl+P in VLC to open your preferences menu. In the bottom lefthand corner of the menu, youâll see the words âShow Settingsâ and the options âSimpleâ and âAll.â Select the âallâ option. You should see this:
6. In the menu on the left, scroll down to the bottom and click on âFiltersâ (not the â>â sign next to Filters, but the word itself.) Find âScene video filterâ and tick the box, as shown:
7. On your computer, prepare the file where you want those frames to go once VLC extracts them. I just created a new file on my desktop named âMulder frames.â Double-click the file to open it, then in the navigation bar, click the dropdown to show the whole file path. Itâll be something like C:\Users\elven\Desktop\Mulder frames. Ctrl+C it.
8. Back in VLC, in the menu on the left, expand the Video>Filters submenu by clicking the > button next to Filters. Scroll down til you see âScene filterâ and click that to open its settings. In the âimage formatâ field, input âjpgâ or âpngâ.
9. In the âdirectory path prefixâ field, paste in the file path from above.
10. Then, choose your ârecording ratio.â If you put in 1, VLC will extract every single frame from the video. If you input 3, itâll be 1 in 3, and so on. If your videoâs on the short side, youâll probably want to grab every frame. I went with 1 in 3.
11. DONâT CHANGE ANYTHING ELSE in this menu. Youâll especially want to avoid ticking the âalways write to the same fileâ box -- thatâll just overwrite the same frame over and over again. (Yeah, I did that to myself.) Your settings should end up looking like this:
12. HIT SAVE.
13. Hit Ctrl+O and select your video. VLC will extract the frames automatically as it plays. Open the destination file you set up earlier -- it should be filled with hundreds of pictures. (Once thatâs done, if you intend to continue using VLC, make sure to open settings again and UNTICK the âScene video filterâ box you ticked in #6. Otherwise, VLC will continue extracting frames from every single video you watch.)
Step 2: Set up your experiment
1. Download and install RunwayML.
2. Launch RunwayML. Create your account, then dismiss whatever âwelcome to the programâ-type popup it gives you.
3. In the lefthand column, click the button that looks like a lil wiggle. That should open this page:
4. Click âTrain a new image model,â name your experiment, and click âCreate.â
5. Itâll prompt you to select a dataset. Click the first box, with the + symbol, and navigate to the file of frames you made earlier. Select it and wait for it to upload. (If your file is larger than 5GB, youâll have to delete some of the images.)
6. Click next. Thatâll take you to âSetup.â This is where you choose a model someone else has trained as a starting point for your own model -- itâs faster to teach a Neural Network how to turn, say, human faces into cats than it is to teach it to make cats from scratch. I suggest just leaving the settings as-is, though if you want you can click âchangeâ and browse other models.
7. Click âstart training.â
Step 3: Enjoy what thou hath wrought
1. Be patient while RunwayML runs the experiment. It took about 3 hours for mine to wrap up. While you wait, you can watch the progress in Runway. Donât worry about missing something -- youâll be able to go back to this experiment and review the whole training process whenever you want. The FID score on the right should slowly count down -- the closer it gets to 0, the closer the AIâs generated results are to the dataset you gave it.
2. Once your experiment is complete, you can use the slider to move back and forth through the training steps. Click the âsave sample imageâ button to save steps you find particularly entertaining, or hit the âcreate progress videoâ button and Runway will generate a video of the whole process.
3. Click âNext.â Thatâll take you to this screen, with a gif of the final product and a few other options.
4. Click âsave videoâ so you can look back at your freaky AI-rendered pet whenever you want. (Thatâs mine at the top of the page.) From here, you can also generate more sample images using the model youâve just trained.
Congrats! You did it! Now try not to see the results in your nightmares!
A new paper from NVIDIA recently made waves with its photorealistic human portraits. Called StyleGAN, the algorithm had a new training dataset pulled from Flickr, with a wider range of ages and skin tones than in other portrait datasets. (If your pictures are on Flickr with the right license, your picture might have been used to train StyleGAN). Thanks to that big dataset, a new method of generating images, and a staggering amount of computer power, its human faces are indeed impressive. But, to prove that their method doesnât just work for human faces, they also generated bedrooms, and cars... and cats.
The cats are so much fun.
In many ways, generating cats is more of a challenge than generating human faces, since cats can be in so many different poses. It has an easier time generating texture than figuring out where all the legs and tails go.
Iâve noticed that for some reason, whenever it generates kittens, one is normal and the rest are haunted. There must be something difficult about pictures containing multiple subjects.
Its humans are even worse, proving that the difference between it and the version that was impressively good at generating human faces was all in the training data.
Speaking of training data, StyleGANâs training data was something called LSUN Cat. Evidently, LSUN Catâs images were sourced from the internet, because when it generated some cats, it added meme text to them, assuming that white blocky lettering is just part of what âcatâ is.
It also attempted to generate cats with the Shutterstock watermark, although actually spelling âShutterstockâ proved tricky for it.
Another delightful thing about its internet-derived dataset: quite a lot of the cats look an awful lot like Grumpy Cat. Sheâs only one cat, but she had a profound impact on what StyleGAN thinks cats look like.
There are far more amazing cats than would fit in this blog post. In this bonus material, Iâve collected a few more of my favorites, including several with meme text, perfect for your internet communication needs. Support AI Weirdness and get bonus content!
You can look through 100,000 example cats, and generate your own, here.
Anya is live and ready to show you everything. Watch her strip, dance, and perform exclusive shows just for you. Interact in real-time and make your fantasies come true.
â Live Streamingâ Interactive Chatâ Private Showsâ HD Qualityâ Free Actions
Free to watch ⢠No registration required ⢠HD streaming