diff --git a/skills/cross-border-ecommerce/1688-cross-border-sourcing/SKILL.md b/skills/cross-border-ecommerce/1688-cross-border-sourcing/SKILL.md new file mode 100644 index 0000000..f30d129 --- /dev/null +++ b/skills/cross-border-ecommerce/1688-cross-border-sourcing/SKILL.md @@ -0,0 +1,209 @@ +--- +name: 1688-cross-border-sourcing +description: Scrape 1688.com for cross-border e-commerce products using Firecrawl API, rank by suitability, and enrich with Temu/Amazon pricing estimates. +version: 1.0 +tags: [1688, sourcing, cross-border, firecrawl, scraping] +--- + +# 1688 Cross-Border E-Commerce Product Sourcing + +Scrape 1688.com for products suitable for cross-border e-commerce (Temu/Amazon/AliExpress), then rank and analyze them. + +## ⚠️ Critical: 1688 Anti-Scraping Reality + +1688.com (owned by Alibaba) has **extremely aggressive anti-bot measures**. Most scraping methods FAIL: + +| Method | Category Pages | Detail Pages (offer/XXX) | Reason | +|--------|:---:|:---:|--------| +| Direct browser_navigate | ❌ | ❌ | Sandbox can't render 1688 JS | +| Jina Reader (r.jina.ai) | ❌ | ❌ | Blocked by Alibaba Cloud WAF | +| Scrapling StealthyFetcher | ❌ | ❌ | IP blocked by Alibaba Cloud | +| CamouFox (headless) | ❌ | ❌ | Server IP blocked (datacenter) | +| Free proxy + CamouFox | ⚠️ | ⚠️ | Page loads but login popup blocks | +| Google/Bing cache | ❌ | ❌ | JS-only challenge page returned | +| AliExpress reference URLs | — | ❌ | JS-only SPA, 2KB bootloader shell | +| **Firecrawl API** | ✅ | ❌ | Category/search listing pages work; **detail pages TIMEOUT** (2026-07-06 verified, 120s+ no response) | +| **Desktop CDP Chrome** | — | ✅ | **ONLY method for individual detail pages** — Chinese residential IP bypasses WAF | + +**For category/search listing pages**: Use Firecrawl API. +**For individual detail pages (detail.1688.com/offer/...)**: Desktop CDP is the ONLY reliable method. Do NOT attempt Firecrawl, Jina, or direct HTTP — they all fail from non-Chinese IPs. + +### Why CamouFox Alone Cannot Access 1688 + +1688's anti-bot runs on two independent layers (see skill:cloakbrowser "反爬两层防御模型"): + +1. Browser fingerprint layer — CamouFox bypasses this naturally +2. IP reputation layer — Alibaba Cloud WAF rejects all known datacenter IPs outright. The response is a redirect to `bixi.alicdn.com/punish/...` with "Access denied" / "unusual traffic" message + +Even with `geoip=True` and `humanize=True`, CamouFox on a server IP will be blocked at layer 2. The only way to browse 1688 interactively is CamouFox + Chinese residential proxy: + +```python +with Camoufox(headless=True, geoip=True, humanize=True, + proxy={"server": "http://residential-proxy:port", "username": "user", "password": "pass"}) as browser: + page = browser.new_page() + page.goto("https://www.1688.com", timeout=30000) +``` + +For scraping (no interactive browsing needed), Firecrawl remains the best option — it routes through its own proxy infrastructure. + +## Step-by-Step Workflow + +### 1. Scrape 1688 Category Pages via Firecrawl + +Use Firecrawl API with `waitFor=5000ms` to extract product data. Focus on **category listing pages**, NOT search pages (search forces login). + +```python +import requests + +FIRECRAWL_API_KEY = "fc-5536880adfa44fb78a3bd9da8dd52b53" + +# Recommended categories for cross-border sourcing: +CATEGORIES = { + "日用百货": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境日用百货", + "宠物用品": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境宠物用品", + "厨房用品": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境厨房小工具", + "美妆工具": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境美妆工具", + "家居收纳": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境家居收纳", + "运动户外": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境运动户外", + "数码配件": "https://s.1688.com/s/company_offer_page/0_-1_-1.html?keyword=跨境数码配件", +} + +resp = requests.post("https://api.firecrawl.dev/v1/scrape", json={ + "url": category_url, + "formats": ["markdown"], + "waitFor": 5000, +}, headers={ + "Authorization": f"Bearer {FIRECRAWL_API_KEY}", + "Content-Type": "application/json", +}) +markdown_content = resp.json()["data"]["markdown"] +``` + +### 2. Parse Product Data from Markdown + +Use regex to extract structured product data from Firecrawl's markdown output: + +```python +import re + +# Price pattern +price_pattern = r'[¥¥]\s*(\d+\.?\d*)' +# Sales pattern: 成交X万+件 or 成交X+件 +sales_pattern = r'成交(\d+\.?\d*)(万?)\+?件' +# Cross-border tags to detect +cross_border_tags = ['跨境', 'temu', '亚马逊', 'amazon', '专供', '爆款'] +``` + +Products appear in markdown as sections with title, price, and sales info. + +### 3. Score Products for Cross-Border Suitability + +Composite scoring (0-100): + +| Factor | Weight | Criteria | +|--------|--------|----------| +| Cross-border tag | 30% | Has "跨境/Temu/亚马逊" in title = full points | +| Sales volume | 25% | 10万+ = 25, 1万+ = 20, 1000+ = 15, <1000 = 10 | +| Price point | 25% | ¥2-15 optimal range (cheap enough for margin) | +| Shipping friendliness | 20% | Light (<200g) + small package = full points | + +### 4. Enrich with Cross-Border Pricing Estimates + +For each selected product, estimate: + +- **TEMU price**: 1688 price × 2.5-4x (USD), capped at typical TEMU range +- **Amazon price**: 1688 price × 4-6x (USD), capped at typical Amazon range +- **Margin**: (Selling price - 1688 cost - shipping) / Selling price +- **Typical shipping cost**: Small packet <500g ≈ ¥15-25 to US/EU + +### 5. Best Product Categories for Cross-Border + +Based on analysis, these categories consistently perform well: + +1. **宠物用品** (Pet supplies) — High demand, emotional buying, good margins +2. **厨房小工具** (Kitchen gadgets) — Lightweight, impulse buy, proven sellers +3. **美妆工具** (Beauty tools) — Tech premium, high Amazon prices +4. **家居收纳** (Home organizers) — Low cost, high utility, evergreen +5. **日用百货** (Daily goods) — Volume play, staple items + +## Desktop CDP Browser — PRIMARY method for individual detail pages + +## ⚠️ SG5 Server-Side Scraping: ALL Methods Blocked (2026-07 Verified) + +SG5 server IP (43.160.244.125) is blocked by Alibaba Cloud WAF at the IP layer. These ALL fail: + +| Method | Result | Detail | +|--------|--------|--------| +| Direct HTTP (curl/urllib) | ❌ | Redirect to `bixi.alicdn.com/punish/...` | +| Jina Reader (r.jina.ai) | ❌ | "Access denied" page | +| Firecrawl API | ⚠️ | Works ~50% but times out on JS-heavy detail pages | +| Google/Bing cache | ❌ | Returns search pages, not cached product data | +| AliExpress pages | ❌ | SPA only (2KB JS bootloader), no server-rendered data | +| AliExpress API endpoints | ❌ | x5sec anti-bot challenge on all endpoints | + +**The ONLY reliable path**: Desktop CDP Chrome (Chinese residential IP) → Bridge → scrape detail pages. + +## Primary: Desktop CDP Chrome Scrape (Chinese IP) + +The user's Desktop runs Chrome on a Chinese residential IP, bypassing Alibaba Cloud's datacenter IP block entirely. + +**Automated scraping script**: `~/.hermes/scripts/1688_scrape_cdp.py` +- Polls Bridge `/health` for Desktop slot with `cdp_browser_running` +- Navigates to each 1688 offer URL via `/cdp/navigate` +- Extracts title, images, price, specs via `/cdp/evaluate` +- Writes enriched data to MongoDB `premiumproducts.products` +- Deploy via cron: `cronjob create no_agent=true script=1688_scrape_cdp.py schedule="every 5m"` + +## Product Detail Enrichment Pipeline (Submit → Full Data) + +For products submitted via Desktop Submit: + +| Step | What | Tool | +|------|------|------| +| 1. Scrape | Extract raw data from 1688 page | CDP browser_console | +| 2. Structure | LLM extracts title/price/images/specs | Qwen 35B (:40006) / qwen3-vl | +| 3. Translate | zh → en, zh → ru | Qwen 35B / API | +| 4. Images | Remove watermark, resize 800×800 | SDXL/FLUX + AtomK toolbox | +| 5. Classify | Auto-categorize product | LLM | +| 6. Write | Update MongoDB document | Server endpoint | + +## Reference Files + +- `references/cdp-detail-extraction.md` — Extract structured product data from 1688 detail pages via Desktop CDP (body.innerText, image URL extraction, expression filter workarounds) +- `references/image-download-conversion.md` — Download alicdn product images and convert WebP→JPG in Cloud Bridge sessions (cron no_agent pattern) +- `references/mongodb-product-update.md` — Update product pool data via direct MongoDB connection when REST API is read-only (credentials, field schema, pitfalls) + +## Scripts + +- `scripts/1688_scrape_cdp.py` — Automated 1688 detail page scraper: health check → find Desktop slot → navigate → extract title/price/images/specs/weight → write MongoDB. Designed for cron `no_agent=true` polling: silently exits when no Desktop connected, runs when Desktop comes online. + +## Pitfalls + +- **Do NOT use search pages** (s.1688.com with keyword search forces login popup) +- **Category pages work better** — products visible without login +- **Product detail links** (offer/XXXX format) use encrypted redirects, hard to extract direct URLs +- **CamouFox on server IP = still blocked** — 1688 blocks at IP layer (datacenter detection), not browser fingerprint. CamouFox alone is insufficient; must pair with residential proxy or use Firecrawl +- **Free proxies are unstable** — may work one moment, blocked the next +- **Firecrawl waitFor=5000** is critical — without it, 1688's lazy-loaded content won't render +- **🔴 Firecrawl detail page TIMEOUT (2026-07-06 verified)**: Firecrawl API times out (120s+) on `detail.1688.com/offer/XXX` URLs. It works for category/search listing pages only. For individual detail pages, use Desktop CDP (Chinese IP). +- **🔴 AliExpress reference URLs are JS-only SPAs**: `aliexpress.com/item/XXX.html` returns a 2KB JS bootloader with no product data. The internal API (`aeglodetailweb/api/header`) also blocks non-browser requests with x5sec protection. AliExpress cannot be used as a fallback data source from server-side code. +- **🔴 Desktop must be connected for CDP scraping**: When Desktop is not running (0 slots on Bridge /health), CDP scraping is unavailable. Use a cron polling script (`no_agent=true`) that silently exits when no Desktop is connected, and runs the scraper when it comes online. See `scripts/1688_scrape_cdp.py` for the canonical pattern. +- **Firecrawl API key can expire** — returns `{"success": false, "error": "Unauthorized: Invalid token"}`. Fall back to Desktop CDP browser (Chinese IP) or renew the key. +- **Server-side CDP browser gets blocked** — `browser_navigate` to 1688 from SG3/SG5 returns `bixi.alicdn.com/punish/...` (Access Denied). Only Desktop Chrome on Chinese residential IP works for CDP-based scraping. +- **MOQ info** not reliably available from category pages; need detail page access for that +- **Prices shown may be tier prices** — the lowest price often requires high MOQ +- **Sales figures** on 1688 can be inflated; treat as directional, not exact + +## Output Format + +Save results as: +- `/tmp/1688_cross_border_top10.json` — Full enriched data +- `/tmp/1688_cross_border_top10.csv` — Spreadsheet-friendly format + +CSV columns: 排名, 品类, 产品名称, 1688价格(元), 1688销量, 预估重量, TEMU参考价($), Amazon参考价($), 预估毛利率, 跨境标签, 推荐指数 + +## Related Skills + +- `web-scraping-toolkit` — General scraping tools reference (Firecrawl is one of 5 tools documented there) +- `atomk-api-export` — Alternative product sourcing from AtomK system +- `temu-pipeline` — Downstream TEMU order processing