[Return] [Bottom]

Posting mode: Reply

BBCode
(for deletion)
  • Allowed file types are: gif, jpg, jpeg, png, webp, webm, mp4
  • Maximum file size allowed is 25000 KB.
  • Images greater than 255 * 255 pixels will be thumbnailed.
  • 115 unique users in the last 30 minutes (including lurkers)




    First
    [1]
    Last

    I'd like to start a running discussion on AI image and video tech, and how to set yourself up to run your own text-to-image (T2I) and text-to-video (T2V) on your own computer.

    I'm a noob at this myself. I've just managed to get SwarmUI up and running with a newly-arrived video card (NVIDIA RTX 3090-based card with 24GB), after an unsuccesful attempt a week ago with my old video card, a card that proved to be utterly inadequate to the task.

    I've managed so far to generate a couple of boring images of houses and a couple of low-quality videos of cars driving down a street. Whoop tee doo. Enough to show my system basically works, way short of pleasing results.

    I found the process of getting Wan 2.2 (the only video model I've installed so far) working very frustrating. I'm fairly sure it's only half working given the quality of my results so far.

    I'd like this thread to be a place where noobs like me can share what we learn as we go along, and where those with more experience can hopefully jump in with advice.

    I used ChatGPT to help get as far as I have, but it was often as misleading as it was helpful. The tutorials I've found online all assume too much previous experience for their advice to be easily followed.

    huggingface.co is apparently the place to go for downloading models, LoRAs, and all the other crap I don't know much about yet. At this point I find huggingface.co more overwhelming than helpful, but I hope to get beyond that point in time.
    >>
    I'm now going through the joy of learning that Wan 2.2 has both a "low-noise" and a "high-noise" model, and that I don't want to select just one of them (as was the only obvious thing to do by default), but that I need a "workflow" that combines both of those models>

    That comes in a JSON file... that I now have to do something with.

    Sigh. This is going to be a long slog, isn't it?
    >>
    My best work so far in this new setup? A three-second long clip at 480p of a car driving down a street.

    "Only" took a little over 10 minutes to generate that shit. Sigh.

    My first rendering on that theme looked terrible, not having yet learned that the separate "low noise" and "high noise" Wan 2.2 models aren't much good on their own, and need to be combined through a "workflow".

    I might not have created any good guro yet, but I am doing a great job of creating torture... for myself.
    >>
    I've made some progress. I can do decent quality stuff with Wan 2.2 now at reasonable speed *although I miss the Wan 2.6 features I'd grown to like).

    I've tried to work with other T2V models, but keep running into issues with them.

    For instance, I'm trying to use a workflow template I downloaded for Hunyuan T2V, but the prompts I feed it are being completely ignored, resulting in lovely videos with, say, a woman quietly sipping a cup of coffee -- no matter what I'm actually asking for. Perhaps I can figure out enough about how these workflows work to fix that.

    I tried a workflow for LTX 2, and but I kept getting some error message about audio metadata I couldn't solve. I tried to delete everything in the workflow related to video, hoping at least to get a video-only workflow going, but my attempts just made the workflow more broken.

    One annoying thing I ran into (even with the Wan 2.2 workflow that works now) is that for some reason SwarmUI demands that the ComfyUI video workflows you use provide a still image output as well as a video output. That wasn't hard to solve for the Wan 2.2 workflow. There was already an image output inside that workflow which could be used to feed another pair of nodes, ImageFromBatch and PreviewImage, to satisfy that requirement.

    I've looked at other workflows that only expose a video output, however, no image output, and have yet to find a tool that will pluck a single frame from a video stream so that this annoying (and senseless, as far as I'm concerned) need for an image output can be sated.

    As for LoRAs...

    I've found a whole bunch which I think will be fun and useful... but putting them to work is going to be a pain. It's hard to know with LoRAs which work with which image and video models. I tried using a futanari LoRA, for instance, and just got abstract splotches of color as a result.

    And the LoRAs aren't easy to employ. I'd hoped I could simply have a list of available LoRAs and check off which ones I wanted to use for any given rendering task. Nope! You've got to build specific workflows and wire in each and every LoRA for each particular mix of LoRAs you might want to use together.

    Apparently I can use some commercial services like Wan 2.6 from within SwarmUI/ComfyUI, but not by running them locally, but as paid services that run remotely. The cost looks to be a lot cheaper than using such things on a commercial web site, but I'll have to figure out some plan for anonymous payment if I'm going to use these things.

    And I might have a tough time solving that stupid image output requirement too.
    >>
    What helped me is copying what others do.
    Most PNG images in the AI board have the workflow embedded in them, you can import them in ComfyUI. (And I think SwarmUI lives on top of ComfyUI)
    Some movie types might also have the workflow embedded in them.

    It;s a good place to learn how everything connects.
    Also. A weird problem I sometimes run into with other people's workflow that I get a black image.
    But when I manually rebuild the workflow and it's settings I get the expected output.

    CivitAI is a good place for tutorials on generic NSFW stuff. Guro, not so much. But there are threads here with guro specific lora's.

    Also, don't mix and max Lora's for different checkpoints. Some might work, some not.
    (ie. Use Pony lora's for Pony checkpoints and Illustrious Lora's with Illustrious Checkpoints.)
    Sometimes they work, sometimes you get artistic splotches.

    Also fiddle around with the Steps, sampler_name and scheduler if you get lots of fireflies.

    Hope it helps.
    >>
    Thanks! I would never have guessed that workflows get embedded into PNG images.

    As for "don't mix and max Lora's for different checkpoints", I haven't figured out how I can know what checkpoints or diffusion models a LoRA is meant to work with. On civitai I haven't seen much info making that clear in most cases.
    >>
    My latest joy on the adventure is trying to set up a NSFW Wan 2.2 workflow that I found, and getting an error about "SageAttention" that turns out to be a Python problem, for which the apparent solution is doing fresh install of Python and some Python extensions...

    It's like pulling teeth every step of the way (um, not in a good way I figure I'd better say, considering where I'm posting). But alas, I kind of expected that.
    >>
    >>15435
    In CivitAI you can filter Lora's on the left. Just select Illustrious or Pony, or any other filter depending on your needs. :)

    Most Lora's will state for which they are build in the description if not the title.

    The SageAttention thing is a bit of a pain on windows, but there are some useful tutorials for that on CivitAI and youtube.

    You can also try to bypass the SageAttention node (ctrl-b, or right click -> bypass) generation will take longer but it might help a bit.
    >>
    File: slash.mp4
    (2.68 MB, 720x720)
    2813105
    This isn't really what I was trying to produce, but I might as well post some of my interesting misses as I go along.
    >>
    985399
    Sometimes the random errors are just... amazing.
    >>
    391443
    The effect of a particular "NSFW" workflow I tried out for Wan 2.2 isn't exactly what you'd want for guro. I guess it's NSFW in that it's been sexed up in a porn-with-a-cheeky-attitude way, rather than just fewer limits and better anatomy.

    This is the best I could do for a hanging so far using that workflow. 😄
    >>
    8752602
    I'm finally getting closer to what I need for a good crucifixion scene. It's a shame that compliance with my prompt is so random. I used this prompt:

    Pam, a naked woman tied down at her wrists and ankles to a wooden cross on the ground. PAM CANNOT MOVE HER HANDS. PAM CANNOT MOVE HER FEET. She can only squirm helplessly.

    A solitary young, pretty black-robed nun enters the scene from above, carrying a mallet in one hand and sharp railroad spike in her other hand. Pam opens her eyes wide and looks terrified.

    Zoom in on the palm of the Pam's hand, keeping Pam's face in view.

    The nun places the tip of her spike in the palm of the Pam's hand, and hammers the spike through the permeable palm and into the wood of the cross. Blood spurts from the wound.

    Pam screams.

    DO NOT EVER LET PAM'S HANDS OR FEET EVER MOVE FROM THEIR ORIGINAL POSITIONS.


    ...and on the forth try I got this. The previous clips had freaky stuff like the woman turning into two women, more nuns than I'd asked for, etc.

    As you might guess from this prompt, it's hard to convince the rendering model that the subject can't move her hands anywhere she wants. At some level the model seems to understand that a human being would try to resist what's happening.
    >>
    I want to do a good crucifixion video eventually, and the first step to get there is a good starter image.

    This image did not come from a single text-to-image prompt. Way, WAY from that. I lost track, but I'd say it probably took 20-30 steps to get here, and each of those steps involved a lot of trial and error with trying to phrase prompts for best results, and often just settling for a baby step as the best that I could do.

    Since explaining these issues requires multiple images, so I'll do that in a series of replies.
    >>
    Here's one example of trying to request an X-shaped wooden cross, sometimes refered to as a Roman cross or a St. Andrew's cross.
    >>
    Eventually I found an image online of something wooden with basically the right shape. This image is the image I found, but after using Photoshop to remove distracting details from the image.

    I used image-to-image conversion until I found a prompt that converted this cross into almost the cross I wanted, and then settled for photoshopping away some extraneous bits.
    >>
    In a few more steps I got to here. No matter how much I used prompts to plead for a direct, looking straight down perspective, a bird's eye view, I always got perspective foreshortening that I didn't want.

    So I gave up on getting an AI to give me the right perspective and took the best image I had up until that point, dragged it into Photoshop, and applied a perspective transform.

    That procedure messed up the shadows around the cross and messed up the grass, so I regenerated the grass and added new shadows using a Photoshop layer effect.
    >>
    My next step was to get a particular woman (already created separately) spread out on that cross with her hands and feet in the correct positions.

    I didn't save the exact wording that I tried to use, but it was all variants on something like this:

    Woman lies down on the cross, arms and legs spread out along the beams of the cross.


    Most of the poses I got back were wildly off. Sometimes the cross itself got modified into a different shape. Even when parts of the prompt request went mostly the right way, the particular model I was using (Gwen Image Rapid Edit) was obstinate about the woman keeping her knees bent.

    What got me closest to where I wanted to go was to put the woman in front of the cross, mark where I wanted I her hands and feet to go, then create a prompt (I no longer have the exact wording) basically asking the rendering model to have the woman use her hands and feet to cover the green spots, with the added request that her hands be palms-up while doing this.

    This finally got the woman spread out the right way (I'll show this image in the next response), with just a big of still-visible green to clean up in Photoshop.
    >>
    So, I was at this point and needed ropes next.

    (The idea is that the woman is at first tied down to the cross to hold her in place and keep her from getting away, THEN she is nailed down to the cross and the ropes are removed, finally to be raised up to hang from the spikes driven through her hands and feet.)

    So, ropes...
    >>
    This is the kind of crazy mess that attempts to request that the woman be tied down to the cross got me. The duplicate arms weren't a common flaw, but all manner of wild and ineffective rope use would happen over and over again.

    What to do?

    I still have access to some online service, and for as little as US $0.40 could have multiple rendering models take a shot at getting the ropes right.

    But not on the whole body and whole cross at at once. The service I was using doesn't even allow simply nudity, forget about anything much spicier.

    And even if not for the nudity restriction, I never got both hands and feet right at the same time in previous tries, so I used just a clip of the woman's feet, and then a clip of just her hands, in separate trials.

    I soon had, as separate pieces, reasonable good rope work for both the hands and feet. I copied those bits as pieces as pasted them over the no-ropes image, using resizing and warping and then a bit of touch-up to have a single image with everything right.

    Oh, and after the woman was lying on her back her tits didn't look as big, so I asked for a boost there from Gwen Image Rapid Edit. It went a little overboard, but I can live with it.
    >>
    The very last thing I wanted to request was that the boards under the woman's feet get lengthened, since that woman's feet were too close to the bottom of the cross for her to hang properly (presuming the cross were merely tilted up and propped up, rather than being raised into the air).

    I kept getting crazy results from that request like this.
    >>
    When random chance finally gave me this, I said good enough, and used Photoshop to get rid of the unwanted triangular piece of wood that popped into unwanted existence.
    >>
    Looking fantastic. Now torture that big tittied bitch to death!!!

    [Top]

    Delete post: []
    First
    [1]
    Last