[Return]

Report a post

Preview
>>26464
>>26465
>>26467
What new venv? You only need to update and add --use-ck-attention to its .bat. Even stable has it now, so no need for nightly. Add H3 related nodes from comfyui-Kjnodes to your workflow as I said earlier. larryvrh is good but try ltx2v and alibaba-pai ones. A new native node sparse attention have been pushed into stable and it may make things even faster paired with ck-attention possibly without tuning them into shit. Use always 8 steps with turbos, even 2 more steps and final quality will improve a lot. Euler/simple is what you should use, other ODE solvers will burn and SDEs are obviously wrong, but er-sde is an exception, it will often outperform the others.
For vllm in comfy use ComfyUI_Simple_Qwen3-VL-gguf node, read the optimal parameters for your model (i guess you already know that since you use llmstudio) and you're good to go, or use LLMstudio as you said. Feed it a system prompt which will contain jailbreak and https://github.com/MiniMax-AI/MiniMax-H3/blob/main/.agents/skills/h3-prompt-writing/references/base-en.txt or https://github.com/MiniMax-AI/MiniMax-H3/blob/main/.agents/skills/h3-prompt-writing/references/ref-en.txt. I would suggest you using Gemma4 instead of Qwen. Gemma 12b-it is similar size as yours, but if you can use 26b: it's MOE, it's much bigger but runs x4 faster and performs 4x better.
Post number No.26497
Board Artificial Intelligence
Optional. Describe what's wrong with it.