A semi-frequent blog about random things.




Review: Local AI Generation

With GPU prices going out of control, I decided to do something bold: Buy an used GPU. Yes, it’s a gamble, I know, but there’s no way around it. I bought myself an NVIDIA Geforce RTX 3060 with 12GB GDDR6 Video Memory. Used. Yeah, I’m an idiot.

There’s a reason for this purchase. This is to support a lower spec machine with my upcoming game (More details at some point). I paired it with a i7 4790k from over a decade ago (seriously this thing is a beast!) and got 32GB of system memory instead of the 8GB it had to do things. 32GB for your system is essentially recommended now, but I can remove modules and run with only 8 or 16 if I wanted to.

That brings me to my next point, CUDA is essentially the backbone of the AI Industry. NVIDIA made lots of money with it, so I thought… “Why not try some *local* AI generation instead?” And this post was made.

Now I wanna be very clear with this. This is a review of its intended design, not its morals or lack thereof. If you want my opinion on AI, I’ve made posts in the past:

So I’ll be as objective as possible.

Local AI Image Generation

I took a big step to do this. I’ve dabbled with it back when my old GTX 1080 was alive and well, but a lot changed in recent years. And yet, it’s still pretty terrible.

Installing things is mostly automated but it’s still kinda wacky. At least you don’t need to download a billion things yourself but you still need Python anyway. There are a million Stable Diffusion WebUI forks, and forks of those forks. I ended up getting ‘Stable Diffusion WebUI Forge – Neo’.

The terrible part is like with any tinkering, it takes an incredible amount of effort to make it look good.

This image above took me 2 days, as I had to pick a style I wanted to do (aka ‘steal’ it from artist or combo of artists), then train a LORA (this took a while), to make it so everytime I use it, it has that style, then find a model that can be used for generation, then wait a few seconds and see your GPU turn a bunch of code into magical art. While admittedly cool as heck, it’s so much work for this. Oh, you want a realistic model? Or an anime model? You can pick one that does all kinda badly, or dedicated anime or realistic models. Still kinda wonky.

Oh, you want a different character? You need to train a LORA again, or find someone that did it for you in a Website. Results can still be iffy either way, and it might take several times to get a decent result. Bonus, faces suck so you need to get something *else* that can detect faces and run the process again to fix them. So those 7 seconds that it said it’ll complete? It’ll take more like 15.

Want a three quarter view image? Get a LORA for that. Multiple characters? Same thing… I could go on with this. In the end you’ll have several LORAs, increasing the complexity of the generation, and thus, it’s time to generate.

Creating LORAs? Different application that has the worst controls I’ve seen in my live. People complain about Linux being obtuse but I couldn’t find a way to make anything until I got help from a friend to let me know that a second tab does allows you to list the example files. The official Github page goes over the details of what you can do with the tabs and settings but not the instructions on how to make a LORA yourself. You’ll have to go somewhere else for the tutorial.

Also, another thing. You’ll might need to generate a bunch of images to get one that you go “I like this.” Each image has a small footprint but if you’re doing it every day, the size adds up. You can repeat seeds with images, essentially the RNG that dictates how the image will be like. Add new information and it’ll change.

So if you want a quick and clean solution? Don’t really do this. NovelAI, Gemini, ChatGPT, Grok, they are all better, with the stand out being NovelAI, if you love anime, and less messy. Sure, they cost money, but they are a few words and a few clicks away from this. Granted, the companies are losing money on this (Probably not Anlatan, the ones behind NovelAI), but it’s much faster to deal with.

If you like to tinker with settings, LORAs, and a bunch of other stuff and have an Nvidia GPU with at least 10GB of VRAM, you can give it a shot. Just don’t say I didn’t warn you.

Oh you got an AMD GPU? Or Intel GPU? What, you crazy?

Local AI Chatbot

I’ve mentioned before that I dislike chatbots but what if it was completely at your whims?

A lot of people use ChatGPT or even Gemini to do some Therapy or Roleplaying, which is a funky concept. But why not give it a shot myself, but locally?

Finding what I needed to get was a little confusing. I found that TextGen works best, so I got that. I needed to get the Full Install, not the portable one. Okay, did that. Change settings so I get it on my network. Done. I need a model, looked up only to see good recommendations. Reddit saved my ass for that, so I got one. Found out that the ‘Transformers’ version makes it so it leaks from the Video RAM to System RAM, turning seconds long generations into 10 minutes plus generations. Swapped to a ‘GGUF’ version. It works fine. No hiccups, the GPU is only used whenever needed, while the VRAM is used automatically.

Ironically, this was the most straightforward one…

I didn’t look on how I could make my model for a chatbot, and I assume I’d need enterprise class hardware for that so it wouldn’t take like a literal century to finish (no hyperbole).

Now the writing quality? Dubious. Thankfully, it works in the same idea of seeds for image generation, so if you didn’t like the result, tell it to redo.

It can do really interesting scenarios if you let it take over OR if you control it as shown in the image above. However, there’s a LOT of redoing, and a lot of tell it to redo things that you didn’t say or do. It’s not as perfect as you’d imagine but I can’t help to feel a little curious to see if it’s gonna be absolutely garbage, absolutely funny or somewhat usable to thing I want it to do.

So far, I’m kinda impressed. Easy-ish to set up, easy-ish to get it running, easy-ish to get to get results! Now I know why people use ChatGPT so much.

Local Video Generation

I tried to make it work. Downloaded ComfyUI. Too literally about 5 hours to download on a 600Mb/s connection and 30 minutes to start. The UI looks confusing but I kept at it. I toggle the video gen thing like in the tutorial of its page (which was pretty nice, actually). ComfyUI Manager says that I need to get 2 more files, 20GB each. I start the download. 90% done for both, they fail. “Restart?” I press the button, it started from scratch.

I gave up.

Other Types

I didn’t try any coding stuff. I didn’t try anything else that calls itself AI but it’s just regular automation without any sort of Large Language Model attached to it. I also didn’t try any Voice Cloning/Gen yet, but I think it’s gonna be as strange as the other things I tried.

Closing Words

So far, Local AI Generation has both impressed me and disappointed me. I can see this being a good tool for some stuff. But it’s not the earth shattering game changer that some people want it to be. Installation can be either a complete nightmare or smooth sailing. Using as well.

But I still stand by what I said before, AI isn’t going to replace all jobs and thinking otherwise is a fool’s errand. The technology behind is cool, yes, but just about any technology behind things is cool. I have a SmartWatch that is as big as a HotWheels 1:64 miniature vehicle that has a battery, a screen and a microphone for calls. Dude, that’s actually so cool! So excuse me if I don’t sound excited about this.

I’d not recommend giving a shot to this unless you have some insane need to get more AI in your life. You could use less of it.