Add cross-border-ecommerce/cross-border-image-pipeline

This commit is contained in:
2026-07-10 16:12:12 +08:00
parent c7b971e724
commit 539bcff6d0
@@ -0,0 +1,93 @@
---
name: cross-border-image-pipeline
description: 跨境电商产品图片AI处理流水线 — 抠图、去水印、换背景、虚拟试穿、LoRA风格训练。支持单卡/双卡GPU部署。
trigger: 用户需要搭建跨境电商图片处理AI流水线,或询问GPU能跑什么图片模型。
---
# 跨境电商图片AI流水线
## 模型与显存需求
| 模型 | 用途 | 显存 | 速度 |
|------|------|------|------|
| BRIA-RMBG-2.0 | 抠图去背景 | ~3GB | 0.3-0.5s/张 |
| IOPaint | 去水印/去瑕疵 | ~4GB | 1.5-2s/张 |
| SDXL + ControlNet | 换背景/构图控制 | ~10GB | 8-15s/张 |
| FLUX.1-schnell | 快速生图 | ~12GB | 4-8s/张 |
| FLUX.1-dev | 高质量生图/Inpaint | ~12GB | 12-20s/张 |
| FLUX.1-Canny/Depth | 边缘/深度精确控制 | ~12GB | 15s/张 |
| FLUX.1-Fill | 高质量Inpaint | ~24GB | 20-25s/张 |
| IDM-VTON | 服装虚拟试穿 | ~14GB | 15-25s/张 |
| OOTDiffusion | 服装试穿 | ~12GB | 12s/张 |
| SD3 Medium | 通用生图 | ~10GB | 10s/张 |
## GPU方案对比
### P102-100 × 2 (20GB, 无NVLink)
- Pascal架构 compute 6.1,无Tensor Core/FP8/BF16/xformers/FlashAttention
- 支持FP16/FP32/INT8
- 完美运行:rembg/BRIA/IOPaint (<4GB)
- 可以跑:SDXL (8-10GB)
- 勉强跑:FLUX.1-dev拆卡(12GB)、IDM-VTON低分辨率
- 跑不了:FLUX.1-Fill、SD3、LoRA训练
- 无NVLink,跨卡通信走PCIe20-30%性能损耗
- ComfyUI分卡调度:GPU0做轻任务(抠图/去水印),GPU1做重任务(SDXL)
### V100-16G × 2 (32GB, NVLink 300GB/s)
- Volta架构 compute 7.0,有Tensor Core
- 部分支持xformers,不支持Flash Attention(需compute≥7.5)
- 单卡可跑FLUX.1-dev/SDXL/IDM-VTON
- 双卡NVLink可跑FLUX.1-Fill(24GB)、FLUX LoRA训练
- 端到端~15s/张
- 三种模式:分卡流水线、NVLink统一32GB、混合模式
## 推荐流水线架构
```
GPU 0 (轻任务) GPU 1 (重任务)
┌──────────────────┐ ┌──────────────────┐
│ BRIA-RMBG 抠图 │──透明图──→ │ FLUX.1-dev 换场景 │
│ IOPaint 去水印 │──干净图──→ │ SDXL Inpaint │
│ 图片预处理 │ │ FLUX.1-Canny 控制 │
│ PIL 800×800 裁剪 │←──成品图───│ IDM-VTON 试穿 │
└──────────────────┘ └──────────────────┘
```
## ComfyUI 部署步骤
1. 安装 ComfyUI`git clone https://github.com/comfyanonymous/ComfyUI.git`
2. 安装依赖:`pip install torch torchvision --index-url https://download.pytorch.org/whl/cu118`
3. 下载模型到 `ComfyUI/models/` 对应子目录
4. 多卡调度:不同节点指定不同 `device`GPU:0 / GPU:1
5. 启动:`python main.py --listen 0.0.0.0 --port 8188`
## 大模型拆卡(V100 NVLink
```python
from accelerate import dispatch_model
pipe = FluxFillPipeline.from_pretrained("black-forest-labs/FLUX.1-Fill", torch_dtype=torch.float16)
pipe = dispatch_model(pipe, device_map="auto") # 自动按层拆到GPU:0和GPU:1
```
## LoRA 训练
| 训练任务 | 显存 | 耗时 | 工具 |
|----------|------|------|------|
| SDXL LoRA | ~12GB 单卡 | ~15min | kohya-ss |
| FLUX.1-dev LoRA | ~20GB 双卡 | ~30min | kohya-ss |
| IP-Adapter | ~16GB 单卡 | ~45min | — |
## 参考文件
- `references/flairgs-api.md` — FlairGS v3.0 统一图像处理 API 文档(去背景/VL分析/文生图)
- 尺寸:800×800 像素
- 格式:JPG (白底) / PNG (透明底)
- 命名:SKU_序号_类型.jpg (如 SKU001_01_white.jpg)
## 注意事项
- P102是矿卡,无NVLink,只能软件级分卡
- V100不支持Flash Attention(compute 7.0 < 7.5要求)
- FLUX模型需从HuggingFace下载,需HF Token
- GGUF量化版可降低显存需求50%但质量略降
- 批量处理建议用ComfyUI API模式,脚本循环提交工作流