


























ROCm 环境多模态开发开源项目汇总
参考开源项目:
ComfyUI:节点式 AI 图像 / 视频生成工作流
https://github.com/Comfy-Org/ComfyUI
Stable Diffusion WebUI:经典 AI 文生图 / 图生图 Web UI
https://github.com/AUTOMATIC1111/stable-diffusion-webui
Diffusers:Hugging Face 扩散模型工具库,适合自定义 AI 生成流程
https://github.com/huggingface/diffusers
Transformers:多模态模型、视觉语言模型、文本 / 视觉 / 多模态模型基础库
https://github.com/huggingface/transformers
InvokeAI:AI 图像生成与创作工作流
https://github.com/invoke-ai/InvokeAI
Fooocus:简化版 AI 图像生成工具
https://github.com/lllyasviel/Fooocus
ControlNet:可控图像生成参考
https://github.com/lllyasviel/ControlNet
InstantID:基于人脸身份保持的 AI 图像生成参考
https://github.com/InstantID/InstantID
Real-ESRGAN:AI 图像 / 视频超分辨率增强
https://github.com/xinntao/Real-ESRGAN
GFPGAN:AI 人脸修复与增强
https://github.com/TencentARC/GFPGAN
AnimateDiff:AI 图像动画 / 视频生成参考
https://github.com/guoyww/AnimateDiff
Stable Video Diffusion:AI 视频生成参考
https://github.com/Stability-AI/generative-models
See-through:单张动漫角色图像自动分层 / 补全 / PSD 输出,可参考 AI 角色素材拆分、Live2D 前处理、2.5D 创作工具方向
https://github.com/shitagaki-lab/see-through
CosyVoice:多语言 AI 语音生成 / TTS / 声音克隆参考,可用于方言绕口令、AI 配音、角色语音、有声内容创作等方向
https://github.com/FunAudioLLM/CosyVoice
此内容由惯性聚合(RSS阅读器)自动聚合整理,仅供阅读参考。 原文来自 — 版权归原作者所有。