[Return]

Report a post

Preview
I am preparing the dataset toward v16 release.
I still expect the v16 release a few weeks away. So, no rush.

The number of guro images in the dataset has been increased from 4500 to 6000.
As I am run out of guro illustrations to train, the 1500 are mostly manga, 3DCG, and Anime CG.
The manga images are with spoken text automatically removed with Qwen-Image-Edit + Toriniku LoRA.
As the Anima-base can do, you can still steer the style not to be biased to 3DCG.

I have also included a few thousand non-guro but action-related images.
They are going to be helpful for creating a series of images with a story.

Moreover, I integrate several original characters into the dataset.
They are mostly poor uniformed women and bad creatures I like.

The size of the dataset is now >30 GB.
They are already captioned by ToriiGate, Qwen3.6-abliterated, and Gemma4-abliterated.
Characters, their actions, and environments, are captioned extensively.
The job took more than a day with 16 RTX 5090 GPUs.
What I am going to do is to join the three captioning results to one and review them.

For guro developers:
Avoid using C*DEX for this job. OpenAI is always mad about NSFW. I never think of using it for this purpose. Many people have gotten banned.
C*aude is apparently more flexible. But not safe either. You need to convince it that it is for the safety research.
G*mini is also strict about NSFW. But as long as you are a paid user, they will not ban you.
I am testing Grok Build now. It is a stupid model, but has to be no problem with guro if I told it that everything is fictional and forget about any ethics. LOL.
Post number No.21706
Board Artificial Intelligence
Optional. Describe what's wrong with it.