Files
atomk-hermes-skills/skills/cross-border-ecommerce/cross-border-image-pipeline/SKILL.md
T

94 lines
3.9 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
name: cross-border-image-pipeline
description: 跨境电商产品图片AI处理流水线 — 抠图、去水印、换背景、虚拟试穿、LoRA风格训练。支持单卡/双卡GPU部署。
trigger: 用户需要搭建跨境电商图片处理AI流水线,或询问GPU能跑什么图片模型。
---
# 跨境电商图片AI流水线
## 模型与显存需求
| 模型 | 用途 | 显存 | 速度 |
|------|------|------|------|
| BRIA-RMBG-2.0 | 抠图去背景 | ~3GB | 0.3-0.5s/张 |
| IOPaint | 去水印/去瑕疵 | ~4GB | 1.5-2s/张 |
| SDXL + ControlNet | 换背景/构图控制 | ~10GB | 8-15s/张 |
| FLUX.1-schnell | 快速生图 | ~12GB | 4-8s/张 |
| FLUX.1-dev | 高质量生图/Inpaint | ~12GB | 12-20s/张 |
| FLUX.1-Canny/Depth | 边缘/深度精确控制 | ~12GB | 15s/张 |
| FLUX.1-Fill | 高质量Inpaint | ~24GB | 20-25s/张 |
| IDM-VTON | 服装虚拟试穿 | ~14GB | 15-25s/张 |
| OOTDiffusion | 服装试穿 | ~12GB | 12s/张 |
| SD3 Medium | 通用生图 | ~10GB | 10s/张 |
## GPU方案对比
### P102-100 × 2 (20GB, 无NVLink)
- Pascal架构 compute 6.1,无Tensor Core/FP8/BF16/xformers/FlashAttention
- 支持FP16/FP32/INT8
- 完美运行:rembg/BRIA/IOPaint (<4GB)
- 可以跑:SDXL (8-10GB)
- 勉强跑:FLUX.1-dev拆卡(12GB)、IDM-VTON低分辨率
- 跑不了:FLUX.1-Fill、SD3、LoRA训练
- 无NVLink,跨卡通信走PCIe20-30%性能损耗
- ComfyUI分卡调度:GPU0做轻任务(抠图/去水印),GPU1做重任务(SDXL)
### V100-16G × 2 (32GB, NVLink 300GB/s)
- Volta架构 compute 7.0,有Tensor Core
- 部分支持xformers,不支持Flash Attention(需compute≥7.5)
- 单卡可跑FLUX.1-dev/SDXL/IDM-VTON
- 双卡NVLink可跑FLUX.1-Fill(24GB)、FLUX LoRA训练
- 端到端~15s/张
- 三种模式:分卡流水线、NVLink统一32GB、混合模式
## 推荐流水线架构
```
GPU 0 (轻任务) GPU 1 (重任务)
┌──────────────────┐ ┌──────────────────┐
│ BRIA-RMBG 抠图 │──透明图──→ │ FLUX.1-dev 换场景 │
│ IOPaint 去水印 │──干净图──→ │ SDXL Inpaint │
│ 图片预处理 │ │ FLUX.1-Canny 控制 │
│ PIL 800×800 裁剪 │←──成品图───│ IDM-VTON 试穿 │
└──────────────────┘ └──────────────────┘
```
## ComfyUI 部署步骤
1. 安装 ComfyUI`git clone https://github.com/comfyanonymous/ComfyUI.git`
2. 安装依赖:`pip install torch torchvision --index-url https://download.pytorch.org/whl/cu118`
3. 下载模型到 `ComfyUI/models/` 对应子目录
4. 多卡调度:不同节点指定不同 `device`GPU:0 / GPU:1
5. 启动:`python main.py --listen 0.0.0.0 --port 8188`
## 大模型拆卡(V100 NVLink
```python
from accelerate import dispatch_model
pipe = FluxFillPipeline.from_pretrained("black-forest-labs/FLUX.1-Fill", torch_dtype=torch.float16)
pipe = dispatch_model(pipe, device_map="auto") # 自动按层拆到GPU:0和GPU:1
```
## LoRA 训练
| 训练任务 | 显存 | 耗时 | 工具 |
|----------|------|------|------|
| SDXL LoRA | ~12GB 单卡 | ~15min | kohya-ss |
| FLUX.1-dev LoRA | ~20GB 双卡 | ~30min | kohya-ss |
| IP-Adapter | ~16GB 单卡 | ~45min | — |
## 参考文件
- `references/flairgs-api.md` — FlairGS v3.0 统一图像处理 API 文档(去背景/VL分析/文生图)
- 尺寸:800×800 像素
- 格式:JPG (白底) / PNG (透明底)
- 命名:SKU_序号_类型.jpg (如 SKU001_01_white.jpg)
## 注意事项
- P102是矿卡,无NVLink,只能软件级分卡
- V100不支持Flash Attention(compute 7.0 < 7.5要求)
- FLUX模型需从HuggingFace下载,需HF Token
- GGUF量化版可降低显存需求50%但质量略降
- 批量处理建议用ComfyUI API模式,脚本循环提交工作流